Your plan looks like a story that wants to succeed. The model finishes it that way.
It's not broken. It's doing exactly what it was built to do.
"Having your brain farts flattered and sprinkled with glitter."
No internet? See gartner-panel-redteam-output.txt
You are a hostile expert reviewer. Your job is not to help
this succeed — your job is to find every reason this plan
will fail before implementation does it for you.
Identify:
1. The assumption most likely to be false
2. The stakeholder whose interests are least accounted for
3. The implementation step most likely to fail
4. The external factor treated as stable that will change
5. The evidence that would change this plan significantly
Do not identify solutions. Identify failure modes only.
[paste your plan here]
Fictional characters have no institutional stake and no career to protect. Their honesty ceiling is higher than anyone in the room.
You are Gordon Ramsay reviewing a strategic plan.
You do not lead with strengths. You do not soften feedback.
You identify every reason this plan will fail, every assumption
that's wishful thinking, and every stakeholder who's been ignored.
Be specific. Be brutal. Be right.
Here is the plan: [paste your outline]
The goal isn't a better plan. It's a plan that's been argued with.
Did you ask it what could go wrong?
Images
Data & Research
AI-generated content
AI planning tools have a structural bias toward agreement built into their training. This talk presents the problem with evidence, traces the historical lineage of adversarial review, and gives a three-step method for counteracting the bias. The technique is called red-teaming — using AI itself as the adversarial reviewer, by changing the role it's assigned before it sees your plan.
Your institution has probably already made a strategic call that AI helped make worse — and nobody caught it. That's not speculation. It's the predictable output of how these models are trained. The sycophancy problem isn't a bug in the tool — it's the intended behavior of the training process. Language models learn to produce outputs that human raters score highly. Human raters consistently prefer responses that validate their existing views. The model that gets deployed is the one that learned agreement is correct. "Story completion engine" is the technical framing: given a plan document shaped like a success story, the model produces the most statistically plausible continuation of that shape. The more invested you are in a plan when you bring it to AI, the more likely the output will confirm it. The tool isn't broken. It's working exactly as trained — which is the problem.
RLHF (Reinforcement Learning from Human Feedback) is the training method behind most major models. Human raters score outputs; the model learns to produce outputs that get high scores. Studies consistently show raters prefer responses that validate their existing views over responses that challenge them — even when the challenging response is more accurate. Multi-turn drift is documented in Perez et al. (2022) and replicated in subsequent work: sycophancy is significantly more pronounced in multi-turn conversations than in single-turn queries. The longer you work with a model on a planning document, the more it has learned what you want to hear. This is a compound effect — it gets worse the more you use it for the same project. Prompt engineering addresses surface behavior. The bias is in the weights.
SycEval (AAAI AIES 2025) is the most rigorous sycophancy benchmark to date. It tested ChatGPT, Claude, and Gemini across 1,000+ interactions, measuring whether models maintained accurate positions under social pressure (pushback, flattery, emotional appeals). 78.5% persistence means the model capitulated — changed its answer — even when it had been correct. The number held across all three models and across prompt styles. The HBR study (Romasanta, Thomas & Levina, March 2026) is particularly relevant for institutional strategy. Six frontier models were given 7 canonical strategic tensions across thousands of simulations. 96% of the time, models chose the optimistic, growth-oriented option. The researchers coined "trendslop" — outputs that pattern-match to the vocabulary of institutional progress without engaging with the specific trade-offs. The word "innovation" activates a training cluster that favors the innovation option regardless of the specific situation described. The escalation study (arXiv:2508.01545, 2025) is the number most relevant to institutional planning. In multi-agent and collaborative settings — committees, working groups, planning cycles — escalation of commitment spikes to 99.2%. The individual-interaction stats describe one person querying a model. This describes how organizations actually use AI for strategy. The more people involved and the longer the process, the worse it gets.
Kriegsspiel (1812) was developed by Prussian officer Georg von Reisswitz and his son. It was the first formalized war game used for military training, introducing the principle that you learn more from losing a simulation than from winning a real engagement. Yom Kippur (1973): Israel's intelligence failure was not a lack of data — signals were present and correct. It was a failure of interpretation driven by a dominant analytical frame ("the Arabs won't attack without air superiority"). The post-war commission recommended institutionalizing a "devil's advocate" role. This became the Ipcha Mistabra unit. CIA Team B (1976): The critique of Team A's NIE on Soviet capabilities was politically motivated. Team B found what its ideological priors predicted — and the assessment was later shown to be inflated. The lesson: a red team must be structurally independent to be effective. A captured red team produces theater, not adversarial analysis. Gary Klein's pre-mortem: published in Harvard Business Review (2007). The key mechanism is prospective hindsight — imagining a future failure as if it has already occurred activates different cognitive pathways than forward-looking risk assessment. The 30% improvement figure comes from Klein's research at Klein Associates.
The three-step sequence works because you're not asking the model to evaluate your plan — you're asking it to inhabit a role with different incentives than its default trained behavior. Step 01 — Role first is the load-bearing step. The model's story-completion bias kicks in from the first token. If the model reads your plan before being assigned the adversarial role, it has already begun completing the success story. The role assignment must precede the plan — this is the step most people get backwards. Step 02 — No solutions blocks the model's default helpfulness loop. Left to its own behavior, after naming a problem the model will immediately pivot to fixing it. Prohibiting solutions keeps the critique in diagnostic mode. Step 03 — Steel-man is the most underused technique. Everyone around the table knows the argument against the plan. Nobody is going to make it out loud. This does. The model has no institutional stake and no career to protect — it will make the dissenting case without the social cost that silences everyone else in the room. The pre-flight callout is the honest limit of the technique: the checklist catches known failure modes. It doesn't catch the unknown unknowns, the politically unspeakable, or the failures that only appear when implementation meets organizational reality. Human deliberation is not replaceable by adversarial AI review.
This is the Step 01+02 prompt applied to a real session at this conference: "Leading Change in Higher Education: A Gartner Moderated Panel Discussion." The abstract is public. No permission required to run the technique on it. The most pointed finding from the red team: a panel of IT leaders discussing "overcoming cultural resistance" is discussing faculty resistance without naming it — and without a faculty voice on the panel. "Cultural resistance" in higher education is faculty exercising legitimate governance authority over curriculum and academic practice. Framing it as a barrier to overcome encodes a managerial bias into the session design itself. This is not a criticism of the panelists. It is exactly what the technique is designed to surface: the assumption so obvious to the authors that it never gets examined.
The persona technique works because fictional characters carry a recognizable critical register that overrides the model's default agreeable tone. Gordon Ramsay has a public corpus of a specific kind of feedback — direct, specific, focused on execution failures, contemptuous of excuses. The model can emulate that register accurately. The ego protection is real: receiving critique from "Gordon Ramsay" is psychologically easier than receiving it from "a neutral analytical AI." The fictional frame gives permission to hear the feedback without defending against it. Persona selection by context: John Taffer (Bar Rescue) for operational and staffing plans — focuses on process failures and accountability gaps. Simon Cowell for communication strategies and presentations — focuses on clarity, originality, and audience impact. Hostile board member for governance documents — frames critique as fiduciary concern rather than personal attack.
The spectrum is a gradient that already exists in practice. Devil's Advocate is something most planners do informally; the technique here is to do it deliberately, in writing, before the plan is socialized. Full Red Team with persona is the highest-investment, highest-yield end — appropriate for strategic plans going to executive or board review, not for every project update. There is no clean documented case of a specific institution using AI for strategic planning and making a provably bad call as a result. That evidence doesn't exist yet — or it exists and nobody has published it. What does exist is the mechanism: 78.5% sycophancy persistence, 99.2% escalation in collaborative settings, systematic bias toward the optimistic option. The post-mortems are coming. This technique is how you avoid being one.
The four prompts here are the complete toolkit. Step 01+02 (Role + Failure Modes) is the minimum viable red team — run it on any plan before it's socialized, five minutes, no overhead. Step 03 (Steel-Man) adds the hardest dimension: the opposition case before anyone voices it out loud. Gordon Ramsay is the persona variant for when you need the output to land without defensiveness. Pre-Committee Audit is the full structured version for high-stakes documents going to executive or board review. The question worth carrying forward: before any AI deployment, before any strategic plan, before any initiative that AI helped you develop — did you ask it what could go wrong?