Ask an AI model for a card move, and you’ll get a confident reply: “Move the six onto the seven.” It reads well. It may also be impossible.
That gap between fluent text and a valid action sits behind many AI agent failures. Here, two-suit Spider Solitaire serves as the test case, because one extra suit is enough to show where models slip.
What does an AI game advisor actually need to know?
A chatbot answers questions. A game advisor proposes actions, and the game’s rules decide whether those actions count.
“Move the six onto the seven” leaves out three facts:
- Which six? The suit matters.
- What sits beneath it? Those cards may move too.
- Is the seven exposed? A covered card accepts nothing.
The fix is to split the work. The model suggests and explains, and a separate validator approves. Thore Graepel, who helped build AlphaGo and now chairs machine learning at University College London, makes a related point about why chatbots don’t reason the way AlphaGo did. AlphaGo combined two systems: one proposed promising moves, and a search process tested them before the program committed.
How often do AI models break game rules?
More often than their confident tone suggests. The SPIN-Bench benchmark tested models on tic-tac-toe, Connect Four and chess, and found that illegal moves became more common as the rules grew more complex.
Chess shows the scale of the problem. In a 2024 test of 1,000 chess puzzles, even GPT-4o made an illegal move 12.7% of the time, and most other models made more illegal moves than legal ones. Without those illegal moves, its creator said, GPT-4o would have rated above 2,000 Elo, the level of a national master. The model clearly understood chess strategy. It just couldn’t keep reliable track of the board.
Which rule trips up an AI game advisor most?
Two-suit Spider turns on one distinction. PlaySolitaire’s rules page puts it plainly: a 6 of hearts may go on a 7 of spades because the ranks fit, but the two cards can’t then be lifted together because their suits differ.
| Action | What the rule allows |
|---|---|
| Place a single card | Onto an exposed card one rank higher, either suit |
| Move a group | Only when the descending run shares one suit |
| Use an empty column | Any single card or valid same-suit group |
A model that tracks only descending ranks will eventually suggest a group move the game blocks.
How should an AI game advisor read the board?
Give it structured state, not just a screenshot:
- Each column, in order
- Which cards face up
- Which columns sit empty
- How many stock deals remain
- Face-down cards marked unknown, never guessed
A wrong board description stays wrong, however long the explanation runs.
Next, request the move as clear fields:
{
"source_column": 3,
"first_card": "6♠",
"card_count": 1,
"destination_column": 7,
"explanation": "Uncovers a face-down card in column 3."
}Structured-output documentation shows how to lock a reply into a schema like this. But a neat format doesn’t make a move legal. “Column three to column seven” passes any schema, even when column three holds no six.
How do you validate an AI game advisor’s move?
Take an exposed six of spades with a five of hearts on top of it. The ranks descend, but the suits differ. An advisor that looks only at the six might try to move both cards onto a seven, and the game will refuse. If the six sits alone, it can go onto a seven of hearts, because mixed suits don’t block a single card.
A validator works through each proposed move in five steps:
- Source: the card exists and faces up.
- Group: find every card that moves with it.
- Ranks: they descend by one.
- Suit: if more than one card moves, they all share one suit.
- Destination: an exposed card one rank higher, or an empty column.
Does your validator follow the right rulebook?
Spider isn’t one fixed game, and an advisor that learned one version can confidently suggest a move another version forbids.
| Rule | PlaySolitaire two-suit | Other versions |
|---|---|---|
| Dealing with an empty column | Allowed; the empty column gets a card too | Some require every column filled first |
| Opening layout | 54 cards: six in the first four columns, five in the rest | Bicycle’s rules deal ten piles of five |
| Group movement | Same suit throughout | Some apps allow same-color or mixed groups |
Is a legal move also a smart move?
Not always. After a move passes, ask what it changes:
- Does it reveal a face-down card?
- Does it build a same-suit run?
- Does it fill an empty column worth keeping open?
Good explanations point to these effects. Weak ones promise a win. The advisor can’t see hidden cards, so it shouldn’t call anything “the winning move.” PlaySolitaire’s own hint feature says the same, describing a suggested move as no guarantee of a win. Its data shows why: across 1,566 random two-suit starts, players won just 17.6% of the time.
What if the game rejects the AI’s move?
Freeze the position and ask:
- Did the board description miss a card?
- Did the advisor treat a covered card as free?
- Does the validator follow a different Spider variant?
Fix the input, not the tone. Don’t push the model to sound surer, and don’t let it repeat a rejected move.
The stakes run higher outside games. A study by the Centre for Long-Term Resilience, funded by the UK’s AI Security Institute, collected nearly 700 real-world cases of AI agents ignoring instructions, evading safeguards and deceiving people between October 2025 and March 2026. Over that period, reports of such behavior rose fivefold. As AI coding agents and other tools take real actions, an outside check matters more than a persuasive explanation. A solitaire validator is a small version of that safeguard.
Key takeaways
- Fluent advice can still break the rules.
- Structured state beats screenshots.
- A schema controls a reply’s shape, not whether the move is legal.
- Validate source, group, ranks, suit and destination for every move.
- Know which rule variant your validator follows.
- Keep “allowed” separate from “smart.”
- Treat rejections as clues, and fix the input first.
FAQs
Q. Why does an AI game advisor suggest illegal moves?
Language models predict likely text and don’t reliably track the full game state. Without a validator, nothing tests their moves against the rules.
Q. Does structured output stop invalid AI moves?
No. It gives you predictable fields, but those fields can still point to cards that aren’t available.
Q. What’s the key rule in two-suit Spider Solitaire?
A single card can go on any card one rank higher, regardless of suit. A group moves together only if every card in it shares one suit.
Related: AI Chess Coach vs Engine: Which One Actually Helps You Improve?
| Disclaimer: This article was contributed by a guest author. The views, opinions, and recommendations expressed are those of the contributor and do not necessarily reflect the views of AIInsightsNews or its editorial team. Readers should independently verify technical details and claims before relying on them. |
