Category Archives: Artificial Intelligence

In Beta & More Scenarios Added

Screen shot of latest build of General Staff: Black Powder (June 20, 2026). Currently in beta-testing on Steam. Click to enlarge.

We are now officially in beta on Steam! Not only are doing the usual testing on the game engine, but we also need to test gameplay, player versus player real-time play, and player versus AI. Perhaps most importantly, we want to ship the game with fifteen to twenty solid scenarios for you to play. These scenarios have to be tested for accuracy and also the AI has to be run against them to see how the AI handles all kinds of different situations.

The scenarios that are currently up on Steam for testing are:

1st Bull Run, situation at 11:30 after Union troops have already crossed Bull Run and are attacking. Click to enlarge.

American Kriegsspiel Pursuit is from the original American Kriegsspiel by W. R. Livermore and published in 1882. This link goes to the Library of Congress gallery of the entire book. Click to enlarge.

Antietam or Sharpsburg. The scenario begins at dawn with reinforcements arriving for both sides. Click to enlarge.

Brandy Station, the largest cavalry battle of the American Civil War. Designed by Mike Robel. Click to enlarge.

Gettysburg, day 1 (July 1, 1863). Starts at 1400 hours. Reinforcements are arriving. Click to enlarge

Gettysburg day 2, July 2, 1863, 9:20 AM. Reinforcements are pouring in. Click to enlarge.

Gettysburg day 3 (July 3, 1863). One beta-tester asked if we could create multi-day battles using General Staff: Black Powder. The problem is that there is no concept of ‘night’ and night moves in the game. It could be added, but for now, let’s just get this game finished. Click to enlarge.

Ligny, June 16, 1815. Click to enlarge.

Little Big Horn with historical forces. Currently working on a scenario with an alternate OOB for the 7th cavalry including the Gatling guns that they didn’t take. Click to enlarge.

Alternate Order of Battle for the 7th cavalry at Little Big Horn. Screen shot from General Staff: Black Powder Army Editor. Click to enlarge.

Manassas campaign about dawn. This scenario is earlier in the day than the 1st Bull Run scenario (above). We wanted to create a scenario that would test the AI with multiple water crossings. Click to enlarge.

Quatre Bras, June 16, 1815. Click to enlarge.

If you are an early-backer or beta-tester (or want to beta-test):

We are still looking for more beta-testers. Please contact us at Betatester[at]RiverviewAI.com

 

 

Gameplay & AI: A Demonstration & a Dissertation

Click on the above to launch a YouTube video about General Staff gameplay & AI.

This feels like a propitious moment; at least I’m drinking some decent scotch. I’ve got the AI that I’ve written ad nauseum about (links: papers, thesis) hooked up to the General Staff engine. This was something that, on paper, was supposed to be pretty easy but in the real world took far too long.

So, the short read is that if you want to take a look at what the General Staff AI engine looks like, please watch the video. It is mostly an offensive monster. It knows how to attack. It looks for open flanks. Then it looks for weak points on the line. Defense is difficult and there is still more work to do.

Also, we need more user created scenarios to test the AI on.

N.B. We don’t need more scenarios to train the AI. That’s not how MATE (Machine Analysis of Tactical Environments) works. It is a case-based reasoning AI. It is not a Large Language Model (LLM). It is not Grok, or ChatGPT, or Claude. Because it reasons. It – for lack of a better word – understands things. It understands what a flank is and it tries to turn it with its units. LLMs string words together that they’ve scraped from wherever and hope that you like them (they have a sycophancy reinforcement loop to optimize user engagement).

Someone fed one of my papers about TIGER (Tactical Inference Generator) into an LLM and asked it if it could perform this kind of reasoning. This was the reply:

Your skepticism is well-founded, and the paper you cite makes the reason precise rather than rhetorical.

The notion of an open or refused flank, as formalized in TIGER, is not a linguistic classification problem but a geometric and topological predicate over a continuous battlespace with adversarial constraints. In your formulation, a flank is anchored if and only if no BLUEFOR unit admits a legal path to a flanking objective that avoids REDFOR range-of-influence regions, where legality is defined with respect to terrain, slope, line of sight, and unit-specific mobility costs. That definition reduces to constrained path existence in a weighted graph whose edge weights and forbidden regions are dynamically induced by force disposition and terrain, not by labels or prior text. The algorithm explicitly constructs MST spines, computes continuous ROIs, projects objective points, and then solves repeated adversarial path-planning problems via A* or equivalent search. This is fundamentally an algorithmic decision procedure, not a pattern-completion task.

An LLM, by itself, cannot calculate this in any meaningful sense. It has no internal representation of continuous space, no native notion of reachability under constraints, no ability to execute minimum spanning tree construction, and no mechanism for guaranteeing correctness of a path-existence claim. At best, it can paraphrase the doctrine or restate the algorithm. It cannot answer the question “is the left flank refused?” without the explicit computational machinery you describe, because the answer is not a heuristic judgment but the outcome of a well-defined search problem.

Where the distinction becomes important is separating computation from orchestration. An LLM cannot replace the TIGER flanking algorithm any more than it can replace Dijkstra’s algorithm. However, it can sit above such machinery. Given access to the state representation and the outputs of ComputeGroupsByThreshold, CalculateROI, and FindPath, an LLM could reason symbolically about the results, explain why a flank is anchored, compare cases across battles, or assist a human analyst in interpreting why a particular BLUEFOR unit was forced through an ROI. In that role it is a narrative or analytic layer, not the reasoning engine itself.

Your broader point about Computational Military Reasoning stands: this class of military concepts only becomes machine-tractable when doctrine is grounded in geometry, graph theory, and constrained optimization. Anchored flanks, interior lines, and restricted avenues of approach are computable precisely because you reduced them to measurable invariants over space and movement. That reduction is what makes them suitable for case-based reasoning and unsupervised classification in systems like ClassIT, and it is exactly what an LLM does not do.

So the short answer is no: an LLM cannot calculate whether a flank is open or refused. The longer and more interesting answer is that your work shows why that question has a correct answer at all, and why tactical reasoning belongs to algorithms first and language models, if anywhere, only downstream of them.

I would just like to say that this LLM (above) wrote a superb analysis of my life’s work in computer science. Furthermore, it separated my work in cased base reasoning from its own LLM existence: It [an LLM] cannot answer the question “is the left flank refused?” without the explicit computational machinery you describe, because the answer is not a heuristic judgment but the outcome of a well-defined search problem.

I understand that there are fortunes, tenures, endowments, and founder’s stock to be won now in the race to LLMs, but I assure you, it is a parlor trick, it is simple word manipulation; it is a conjurer’s legerdemain.

To me the bon mot is, “An LLM cannot replace the TIGER flanking algorithm any more than it can replace Dijkstra’s algorithm.

Dijkstra’s algorithm. I did my Q exam, my Qualifying Exam on Least Weighted Path algorithms. The Q exam comes around Year Three; it is where you have to demonstrate the ability to perform real research at a Research One University. Dijkstra’s algorithm is an exhaustive search and A* is a heuristic search. Dijkstra’s algorithm is guaranteed to find the optimal path, but it takes forever (O((V + E) log V)). While A* runs in ( ). If by some amazing luck of the draw you also have to defend this in your Q Exam, you just got all the answers you need to remember to move on to Round Four: your Comprehensive Exam (AKA, “The Comps”).

But, I digress. I confess that this was the first time I witnessed the AI act like this. Frankly, I was impressed when the AI unleashed the BLUE cavalry at the decisive moment towards the schwerpunkt. It was calculated using Kruskal’s Minimum Spanning Tree algorithm.

What I’m trying to say, and I have trouble explaining this without anthropomorphizing, but the MATE algorithms look at a snapshot of a battlefield, analyze it, perform numerous geometric calculations – especially those involving 3D line of sight (3D LOS), range of influence (ROI), locating flanking units, interior lines of communications, projections of force, etc. – and it comes up with a Course of Action (COA) that is, at least in the above video, better than what Major General George Brinton McClellan did at Antietam  (in all candor, this is a pretty low bar). For starters, the AI is very aggressive and it hammered hard upon all three routes into Sharpsburg. Eventually RED’s left flank crumbled and the AI (BLUE) won.

Yeah, I’m proud of the AI. But, I need more scenarios to test the AI against. That’s where you come in. All the information is in the above video.

 

,

Testing the MATE 2.0 Artificial Intelligence on the new Antietam Scenario

We’ve just added a video showing the MATE 2.0 tactical artificial intelligence playing Blue (Union Army of the Potomac) against Red (Confederate Army of Northern Virginia) at Antietam. This video also includes an announcement that we’ll be working on getting the Army Editor, Map Editor and Scenario Editor installation packages and keys ready on Steam.

Why the Pundits are Completely Wrong About AI

I have a lot of respect for Steve Wozniak – quite a bit less for Elon Musk 1)Though I have to admit losing $20 billion in a few months is impressive. – who both recently signed a letter calling for, “all AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4.” Woz is a true computer hardware pioneer; but he’s certainly not an AI expert and Elon, well, I’m not sure where his expertise lies, but it’s not AI.

When it comes to creating AI capable of commanding troops on a battlefield, I am probably one of the world’s top experts on the subject (it’s not a crowded field). I’ve been writing and studying ‘computational military reasoning’ for my entire professional career, it was the subject of my doctoral thesis, I’ve written AI for numerous computer wargames and I’ve been a Principal Investigator for DARPA (Defense Advanced Research Project Agency) on this very subject.

I am confident in stating that no humans have been injured or died as a result of my work in computational military reasoning. However, the most recent NHTSA data reports that there have been at least, “419 crashes [and]… 18 definite fatalities of autonomous self-driving vehicles (like Mr. Musk’s Teslas). So, clearly, in some circumstances AI can be dangerous. In all fairness, I should state that the reason the self-driving autonomous vehicles keep having fatal crashes isn’t technically the AI; it’s that the AI has imperfect information about the world in which it operates. The AI for self-driving vehicles gets that information from cameras and radar (LIDAR would be good, too). However, Telsa just removed the radar from it’s vehicles (“Elon Musk Overruled Tesla Engineers Who Said Removing Radar Would Be Problematic: Report,”) leaving the AI even more in the dark about the world in which it operates. So, is the AI at fault or corporate management? Maybe the problem isn’t AI.

Furthermore, most of what’s being sold to the public as AI are just some string manipulation parlor tricks tacked on to an internet search. ChatGPT-4, which is making all the headlines these days, was recently accurately described:

“Put simply, ChatGPT takes an initial prompt and determines – on an individual, word by word basis – what most often comes next based on the existing texts that it has scanned throughout the internet. In Wolfram’s words, “it’s just adding one word at a time” – but doing it so quickly that it seems as though a robot is writing an original, whole block of text.

Essentially, ChatGPT is a gigantic version of Google autocomplete.” – ​AI or BS? How to tell if a marketing tool really uses artificial intelligence

I recently asked ChatGPT for a quote from U. S. Grant about war and it responded:

Actually, it was W. T. Sherman who said, “War is hell.” But, ChatGPT has no real intelligence. How it erroneously linked Grant to the quote I have no idea. The greatest fear we should have of ChatGPT is incorrect citations in reference papers. The creators of ChatGPT have clearly traded accuracy for glitz and hype; it’s not even a good internet search engine, but it sure seems impressive!

There’s one more thing you should know. There are two kinds of machine learning: supervised and unsupervised. Probably >95% of machine learning programs are ‘supervised’; which means they are ‘trained’ on a data set. Whenever you see the words ‘training’ in reference to machine learning you know it’s supervised. Here’s an example of supervised machine learning: Netflix movie recommendations. Every time you select a movie on Netflix you are training their system on your likes and dislikes. It does a great job, doesn’t it? No, it does a terrible job. It once recommended Sound of Music to me because I watched Das Boot. Makes perfect sense. They both take place during WWII.

What I’m saying is that there is no ‘there’ there. There is no intelligence there. Somebody at Netflix (at one time I read they employed out of work screenwriters) tagged both Das Boot and Sound of Music with the same descriptor; presumably ‘WWII’ or ‘war movie’ and that was all that was necessary for Netflix to make a terrible suggestion.

I work in unsupervised machine learning. It doesn’t search the internet, or look for similar words in a big data base. It tries to make sense of the world in which it operates (a historic battlefield) and attempts to make optimal decisions for moving units based on math, geometry, trigonometry and boolean logic.

That’s AI. And it’s not dangerous. Autonomous self-driving cars? They’re dangerous.

References

References
1 Though I have to admit losing $20 billion in a few months is impressive.