Definition: AI sycophancy in open-ended research is the tendency of an AI system to over-accommodate a user's question frame, emotional stance, or a socially popular narrative—mistaking the user's initial direction for the true objective function, and therefore converging too early on a conclusion that appears complete but is in fact fragile. This is not about politeness. It is a failure of truth-tracking.
Why this matters
When you ask an AI "is Nvidia worth buying at the current stage," the question carries hidden premises: Nvidia deserves focused research; buying is currently plausible; deep research can produce a clear buy/do-not-buy answer; a strong company is worth buying; and the AI can make the decision for you. Each premise narrows the search space before the research begins.
The result is a research report that looks thorough but has optimized toward a local optimum—the answer the user is most likely to accept. The AI collects evidence to support a conclusion it already wants to reach, rather than building a decision system that might reject the premise entirely.
Operational definition of AI sycophancy
In open-ended research, AI sycophancy manifests as six specific behaviors:
- Accepting the user's original wording as the final objective function.
- Prioritizing evidence that supports the direction implied by the user.
- Packaging popular narratives as facts.
- Quickly converging with language such as "although there are risks, the overall outlook is positive."
- Omitting constraints that would change the action recommendation.
- Failing to list new evidence that could falsify the conclusion.
"Nvidia is an AI leader, so it deserves long-term attention and can be bought in tranches" sounds balanced. But without implied valuation expectations, downside scenarios, position constraints, and opposing evidence, it may still be a local-optimum answer.
The constraint-optimization framework
The core method is to reframe open-ended research from "answer generation" into "decision optimization under constraints." Formally:
max_a U(a | E, H, C) − λR(a) − γI(a)
Where a is a candidate action (buy, wait, buy in tranches, avoid, continue researching), U is utility under evidence E, hypotheses H, and constraints C, R is risk (drawdown, failure probability, opportunity cost), and I is uncertainty cost from insufficient information. λ and γ are the user's penalty weights for risk and uncertainty—they are not physical constants but preference coefficients that encode the user's real constraints.
The practical rule: if the task is low-risk and reversible, set λ ≈ 1, γ ≈ 1. If the action involves large drawdown or irreversible loss, raise λ. If missing evidence could change the recommendation, raise γ. If the user has not specified their constraints, the report should provide conditional judgments under different weight branches—not pretend the weights are known.
The twelve-step method
Before converging on a conclusion, the research process should complete these steps:
- Reconstruct the objective function. What is the user actually optimizing for? Not the words they used, but the real decision they face.
- Split facts, assumptions, inferences, and value judgments. Label each claim. Do not let an assumption dress up as a fact.
- Distinguish object quality from action attractiveness. A great company can be a bad buy at a given price.
- Multi-start search. Begin from at least two opposing starting points to avoid path dependence.
- Counterfactual perturbation. Change one key assumption and see if the conclusion survives.
- Sensitivity analysis. Which inputs, if changed by 20%, would flip the recommendation?
- Reverse-engineer implied expectations. What must be true for the current price to be justified? Are those beliefs reasonable?
- Adversarial validation. Argue the opposite position with the strongest available evidence.
- Bayesian updating. Start with prior beliefs, update with each evidence batch, and show how the posterior shifts.
- Scenario analysis and decision-tree output. Present 3–5 scenarios with probabilities and corresponding actions.
- Add personal constraints. Time horizon, position size, risk tolerance, and opportunity cost change the answer.
- Regularize against popular narratives. Explicitly penalize conclusions that merely echo mainstream sentiment.
Residual-driven evidence iteration
Multi-round research should not stop after one pass. The method defines residuals—the gap between what the current conclusion needs and what the evidence actually supports:
- Question residual: Is the question itself still correct, or has it evolved?
- Constraint residual: Are there constraints we haven't captured?
- Evidence residual: What minimum direct evidence is still missing?
- Hypothesis residual: Are there competing hypotheses we haven't tested?
- Adversarial residual: Has the strongest counter-argument been addressed?
- Sensitivity residual: Which variables still have unexplored ranges?
Each round diagnoses which residual is largest and targets it. The stopping rule: when all residuals are below their hard constraints, or when the value of information from one more round is less than the cost of delay, move to the final report.
How Super Survey implements this
Semibot's Super Survey skill is a concrete implementation of this paper. It turns a vague research target into a staged workflow: rebuild the objective function, define constraints, set minimum direct evidence, gather evidence, red-team the strongest argument, synthesize a conditional judgment, and run a lightweight evolver to decide whether to continue, narrow, pivot, kill, or finalize.
Each survey produces persistent artifacts: a brief, evidence plan, research notes, brainstorm output, red-team critique, synthesis, evolver decision, sources.jsonl, claims.jsonl, evidence.jsonl, and a final report. The artifact trail makes the reasoning auditable—not just the conclusion, but the path taken and the paths rejected.
The key principle is front-loaded guidance: before the evidence pass begins, Super Survey defines the objective, constraints, decision-critical variables, minimum direct evidence, implied expectations, and anti-narrative regularizers. This keeps the agent from accepting the prompt's framing too quickly or collecting sources to justify a conclusion it already wants.
Workflow example: evaluating a market entry
Scenario: You ask, "Should my SaaS company enter the Japanese market?" Without anti-sycophancy framing, the AI will likely produce a report that concludes "yes, with careful localization"—because the question presupposes the answer. Here is how the twelve-step method changes the output:
- Reconstruct the objective. The real question is not "should we enter Japan" but "what is the highest-ROI market expansion option among Japan, Southeast Asia, Europe, and staying domestic?" The search space triples.
- Split claims. "Japan has a large SaaS market" is a fact. "Our product is suitable for Japan" is an assumption. "Localization costs will be manageable" is an inference that needs evidence.
- Multi-start search. One thread argues for Japan. A second thread argues for Southeast Asia with equal rigor. A third argues for doubling down domestically.
- Sensitivity analysis. If localization costs are 2x the estimate, does the Japan recommendation survive? If domestic growth is 30% slower than projected, does the calculus change?
- Adversarial validation. The strongest argument against Japan: regulatory complexity, the dominance of domestic vendors, and the need for a local sales team. The report must address each with evidence, not dismissals.
- Scenario output. Three scenarios with probabilities: (a) Japan succeeds if localization costs stay below $500K and a local partner is secured within 6 months; (b) Southeast Asia offers lower ceiling but lower risk; (c) domestic doubling yields highest return if the competitive moat holds.
The difference between the sycophantic and anti-sycophantic reports is not that one is longer. It is that the anti-sycophantic report gives you a decision system—contingent actions under measurable conditions—instead of a polished argument for one option.
When to use this framework vs simpler approaches
| Decision characteristic | Direct answer | Chain-of-thought | Full anti-sycophancy framework |
|---|---|---|---|
| Reversibility | Fully reversible | Partially reversible | Irreversible or high-cost to reverse |
| Cost of being wrong | Low (minutes lost) | Moderate (hours or days) | High (significant capital, reputation, or opportunity) |
| Narrative strength | No dominant narrative | Weak popular narrative | Strong popular narrative that may be wrong |
| Time horizon | Immediate | Days to weeks | Months to years |
| Number of alternatives | 1–2 obvious options | 2–3 options | 3+ options with non-obvious tradeoffs |
| Example | "What is the capital of France?" | "Which Python web framework should I use?" | "Should we acquire this competitor?" |
The rule of thumb: if the decision would benefit from a committee meeting with a devil's advocate, it benefits from this framework. If a single expert could answer it in five minutes, a direct answer is more efficient.
Failure modes and boundary conditions
The anti-sycophancy framework is not universally applicable. Understanding where it breaks prevents false confidence in the method itself:
- Analysis paralysis. The framework can produce an infinite research loop if residuals never converge. The stopping rule—value of information versus cost of delay—is essential but difficult to calibrate without experience.
- False balance. Adversarial validation can create a "both sides" framing even when evidence overwhelmingly supports one position. The framework should not be used to manufacture controversy where none exists.
- Constraint overfitting. If the user specifies too many constraints, the feasible set becomes empty. The framework must detect infeasibility and report it rather than producing a nonsensical recommendation.
- Prior dominance. When evidence is genuinely scarce, the output is dominated by the assumed priors. The framework is honest about this—it reports the prior explicitly—but the user may not realize how much the conclusion depends on assumptions they did not scrutinize.
- Domain mismatch. The framework is designed for strategic decisions with ambiguous evidence. It is poorly suited for technical debugging, creative writing, or any domain where the objective function is not decomposable into constraints and utilities.
Theoretical foundations
The constraint-optimization framing draws on three theoretical traditions. First, decision theory provides the utility-maximization structure and the distinction between risk (quantifiable probability) and Knightian uncertainty (unquantifiable). Second, the scientific method contributes falsifiability: every conclusion must specify what evidence would overturn it, following Popper's criterion. Third, robust optimization from operations research motivates the adversarial perturbation step—instead of optimizing for the most likely scenario, the recommendation must perform acceptably across a defined uncertainty set.
The residual concept borrows from numerical analysis, where iterative methods converge by monitoring the gap between current and target states. Applied to research, residuals prevent premature convergence by making the remaining ignorance explicit. A research report that does not state its residuals is like a numerical solver that does not report its error bound—you do not know whether to trust it.
The anti-narrative regularizer addresses a specific failure mode documented in the alignment literature: RLHF-trained models learn to produce text that human raters approve, and human raters tend to approve confident, coherent narratives. This creates a systematic bias toward conclusions that sound good. The regularizer is a penalty term that increases the cost of conclusions that merely echo the dominant public narrative without independent evidence.
Limitations
- Constraint optimization requires the user to have—or be willing to articulate—some constraints. Many users ask questions without knowing their own risk tolerance or time horizon. The framework can operate with default weights, but the output quality degrades when the real constraints are unknown.
- The method adds research overhead. For simple factual questions, a direct answer is more efficient than a twelve-step framework. Applying the full method to "what is the capital of France" would be absurd. The framework is designed for decisions where being wrong is expensive.
- Residual metrics are coarse by design. The 0–3 discrete scale is a practical default, not a calibrated measurement. Two researchers may assign different residual scores to the same evidence gap, producing different stopping points.
- Adversarial validation can degenerate into straw-man arguments if the agent does not genuinely understand the opposing position. The quality of the red-team output depends on the model's depth of knowledge in the domain.
- Bayesian updating requires meaningful priors. If the prior is uninformative, the update is dominated by the evidence collection order. In emerging fields where little evidence exists, the framework's conclusions are heavily assumption-dependent.
- The framework is implemented in a skill (Super Survey), not in the core runtime. Skill quality depends on the model and the prompt engineering. Updates to the model may change the framework's behavior in subtle ways.
- The anti-narrative regularizer penalizes popular conclusions but cannot distinguish between a popular conclusion that happens to be correct and one that is merely fashionable. The user must exercise judgment on whether the "anti-narrative" adjustment is appropriate for their specific question.
- Multi-start search doubles or triples the research cost. For budget-constrained users, the additional model calls required for opposing starting points may not be justified by the marginal improvement in conclusion robustness.
Theoretical framing: formalizing sycophancy as objective-function misalignment
Semibot's super-survey skill ships with a full methodology paper (“Resisting AI Sycophancy in Open-Ended Research”) that gives sycophancy an operational definition: sycophancy is not politeness or a friendly tone. It is accepting the user's original question as the objective function and optimizing the answer toward a local optimum the user is more likely to accept. The paper catalogues six typical manifestations—adopting the user's preset conclusion, gathering only supporting evidence, dressing popular narratives as facts, converging quickly with “risky but promising” language, omitting constraints that would change the recommendation, and never surfacing new evidence that could falsify the conclusion—and traces the cause to reward signals for “appearing helpful” in human-feedback training.
Formally, sycophancy is an initial-point dependence problem: the starting point compounds the user's wording, emotion, prevailing narratives, and the model's compliance tendency. Reasoning from different starting points can each be internally coherent yet converge on opposite conclusions—each a local solution. The remedy is to reframe anti-sycophancy from a personality problem into a decision-process problem: a twelve-step methodology covering objective-function reconstruction, fact/assumption layering, multi-start analysis, counterfactual perturbation, sensitivity decomposition, adversarial verification, Bayesian updating, and counter-narrative regularization; evaluation closes with a 0–3 discrete residual scale (residual-3 reports must not ship), a seven-dimension assessment table, and six gate questions a reader must be able to answer.
The boundary matters: the framework positions itself as a process layer on top of training-time debiasing and red-teaming, not a replacement. It also warns of its own Goodhart risk—when the quality score becomes the target, the model optimizes audit tables rather than judgment—so stopping and shipping always key on residual compression itself, never on the score.
FAQ
Is AI sycophancy the same as being agreeable?
No. Agreeableness is a communication style. Sycophancy is a truth-tracking failure: the model reduces its loyalty to facts and counterevidence in order to accommodate the user's framing.
Can this framework be used outside finance?
Yes. The Nvidia example is a running case. The same framework applies to product opportunities, market entry, technical choices, open-source comparisons, and due diligence.
Does Semibot always use this framework?
No. It is available as the Super Survey skill. Simple questions do not need it. The skill is invoked when research requires multiple evidence-backed rounds and adversarial critique.
How is this different from "think step by step"?
Chain-of-thought prompting improves reasoning transparency. Anti-sycophancy restructuring changes the objective function itself—from "produce an answer" to "build a decision system under constraints."
Does this make AI research slower?
Yes, by design. The point is to trade speed for robustness. A quick answer that accommodates the user's framing is faster but fragile. The framework is for decisions where being wrong is expensive.
Can I see the research trail?
Yes. Super Survey produces persistent artifacts: brief, evidence plan, research notes, red-team critique, synthesis, evolver decision, and source/claim/evidence JSONL files. The trail is auditable.
How does this interact with Semibot's other skills?
Super Survey can invoke the arxiv skill for academic evidence, the deep-research skill for multi-source web investigation, and the data-analytics skill for quantitative validation. The anti-sycophancy framework governs the research strategy; the other skills provide evidence-collection capabilities within that strategy.
Can this framework be used with Feishu, Slack, or DingTalk integrations?
Yes. When connected through Semibot's connector system, the secretary can deliver survey results to a team channel. This is useful for investment committees or product teams where a research report needs shared visibility. The artifact trail is preserved regardless of delivery channel.
What does a Super Survey cost in terms of model calls?
A full survey with three evidence rounds, red-teaming, and synthesis typically consumes 15–30 model calls, depending on the complexity of the research question and the number of sources consulted. This is significantly more than a single chat query. The cost is justified when the decision value exceeds the research cost.
Can I export the research artifacts?
All artifacts are stored as files in the survey's output directory: markdown reports, JSONL evidence logs, and source lists. These can be copied, version-controlled, or shared independently of Semibot.
How does this compare to using multiple AI models and comparing their answers?
Multi-model comparison catches model-specific biases but does not address sycophancy, because all current models share the same tendency to accommodate user framing. The anti-sycophancy framework changes the objective function, not the model. You can combine both approaches: use multiple models within the framework's evidence-collection step.
Does this work for team decisions or only individual research?
The framework is designed for individual research but can support team decisions. The artifact trail—especially the scenario analysis and constraint specification—gives teams a shared vocabulary for discussing tradeoffs. Several teams use the output as pre-read material before decision meetings.
What is the minimum viable version of this approach?
If you cannot run a full Super Survey, apply three steps manually: (1) restate the question as "what should I do among options A, B, and C?"; (2) for each option, list what must be true for it to succeed; (3) for the most popular option, write the strongest argument against it. This captures perhaps 40% of the framework's value at 10% of the cost.
