SemibotSemibot - AI Desktop
All guides

Constraint Optimization for Research: From Answer Generation to Decision Systems

The shift from 'what is the answer?' to 'what decision should I make given my constraints?' reframes research as constrained decision optimization—formalized with sensitivity analysis, Bayesian updating, and decision trees.

Definition: Constraint-optimized research reframes open-ended inquiry from "find the answer" to "build a decision system under constraints." The research output is not a verdict—it is a structured set of conditional judgments, each qualified by evidence strength, risk tolerance, and user-specific constraints. This approach formalizes methods from robust optimization, sensitivity analysis, Bayesian updating, and decision analysis.

The problem with answer generation

"Is Nvidia worth buying?" sounds like a question with an answer. It is not. It is a decision problem with hidden variables: who is buying, for how long, with how much capital, with what drawdown tolerance, compared to what alternatives, and with what information would change the judgment. Without these constraints, the word "worth" has no stable meaning.

Answer generation optimizes for plausibility: produce a response that sounds thorough and balanced. Constraint optimization optimizes for decision utility: produce a system that helps the user make a better decision given their actual situation.

The formal framework

The decision is formalized as:

max_a U(a | E, H, C) − λR(a) − γI(a)

Where a is a candidate action, U is utility under evidence E, hypotheses H, and constraints C, R is risk, I is uncertainty cost, and λ, γ are the user's penalty weights. These weights are not physical constants—they are normalized preference coefficients that encode the user's real situation.

Practical defaults: low-risk reversible tasks get λ ≈ 1, γ ≈ 1. High-drawdown actions raise λ. Missing evidence that could change the recommendation raises γ. If the user has not specified constraints, the report provides conditional judgments under different weight branches.

The twelve-step method

Before converging:

  1. Reconstruct the objective function. What is the user actually optimizing for?
  2. Split facts, assumptions, inferences, and value judgments.
  3. Distinguish object quality from action attractiveness.
  4. Multi-start search. Begin from opposing starting points.
  5. Counterfactual perturbation. Change one key assumption.
  6. Sensitivity analysis. Which inputs flip the recommendation?
  7. Reverse-engineer implied expectations.
  8. Adversarial validation. Argue the opposite.
  9. Bayesian updating. Show how the posterior shifts with evidence.
  10. Scenario analysis and decision trees.
  11. Add personal constraints.
  12. Regularize against popular narratives.

Residual-driven iteration

Research does not stop after one pass. Residuals—the gap between what the conclusion needs and what evidence supports—drive the next round:

  • Question residual: Is the question still correct?
  • Constraint residual: Uncaptured constraints?
  • Evidence residual: Missing minimum direct evidence?
  • Hypothesis residual: Untested competing hypotheses?
  • Adversarial residual: Unaddressed counter-arguments?
  • Sensitivity residual: Unexplored variable ranges?

Stopping rule: all residuals below hard constraints, or value of information from another round is less than delay cost.

Implementation in Super Survey

Semibot's Super Survey skill implements this framework. Each survey round follows a staged loop: frame the user's wording as a starting point, route through the Adaptive Research Framework, write the evidence plan before searching, gather evidence with source/claim/evidence JSONL tracking, red-team the strongest argument, synthesize a conditional judgment, and run the evolver to decide continue/narrow/pivot/kill/finalize.

The key innovation is front-loaded guidance: before any evidence is collected, the skill defines the objective, constraints, decision-critical variables, minimum direct evidence, implied expectations, and anti-narrative regularizers. This prevents the agent from accepting the question's framing as the objective function.

Scenario: evaluating a market entry decision

A startup founder asks: "Should we enter the Southeast Asian market?" Traditional research would produce a market overview with statistics. Constraint-optimized research reframes this as a decision system:

  1. Reconstruct the objective. The actual question is not "is Southeast Asia a good market?" but "given our current burn rate, product-market fit in North America, and limited sales team, does entering Southeast Asia within 6 months maximize expected revenue per dollar spent compared to deepening our North American presence?"
  2. Identify constraints. Burn rate (12 months runway), team size (2 salespeople), product localization needs (3 languages), regulatory requirements (data residency laws), and competitive landscape (3 established local players).
  3. Define minimum direct evidence. Before recommending, the research must include: at least 3 comparable market entry case studies, regulatory analysis from primary sources, and customer interviews or surveys from the target market.
  4. Sensitivity analysis. The recommendation flips if: customer acquisition cost exceeds $X, or if localization takes more than Y months, or if a competitor launches a similar product first. These thresholds become the decision-critical variables.
  5. Conditional judgment. The output is not "enter Southeast Asia" or "do not enter." It is: "If localization can be completed in under 3 months AND customer acquisition cost is below $45, enter. If localization takes 4-6 months, enter only if competitor activity is low. If either condition fails, deepen North American presence instead."

The founder gets a decision system they can act on as information arrives, not a static recommendation that becomes outdated.

Scenario: technical architecture selection

An engineering team asks: "Should we use microservices or a monolith?" This sounds like a technical question, but it is a constrained decision problem:

  1. Reconstruct the objective. The real question is: "Given our team of 8 engineers, our need to ship features weekly, our operational maturity level, and our 18-month growth target, which architecture minimizes total cost of ownership while maintaining deployment velocity?"
  2. Multi-start search. Begin from the pro-microservices position (scalability, team autonomy) and the pro-monolith position (simplicity, faster iteration). Do not start from a neutral position that averages the two.
  3. Counterfactual perturbation. What if the team grows to 20 engineers in 12 months? What if traffic grows 10x? What if the lead DevOps engineer leaves? Each perturbation may shift the recommendation.
  4. Decision tree. Build a tree: if team stays under 12 engineers for 18 months → monolith. If team exceeds 15 AND operational maturity is high → microservices. If team exceeds 15 BUT operational maturity is low → modular monolith as a bridge.
  5. Adversarial validation. The strongest argument for microservices is team scaling. The strongest argument against is operational complexity. The red-team pass tests whether the team actually has the operational maturity to run distributed systems.

The output is a conditional decision tree that the team can update as their situation evolves, not a one-time architectural recommendation.

Scenario: competitive response analysis

A product leader asks: "Our competitor just launched feature X. Should we build a response?" Traditional research would compare features. Constraint-optimized research asks a different question:

  1. Reconstruct the objective. The real question is: "Does building a response to feature X, given our current roadmap, engineering capacity, and customer retention data, produce more expected value than the next highest-priority item on our roadmap?"
  2. Distinguish object quality from action attractiveness. Feature X may be objectively good (high user satisfaction, strong press coverage). But that does not make copying it the right action. The action's attractiveness depends on your specific constraints, not the feature's quality.
  3. Reverse-engineer implied expectations. If the team is panicking because of competitor press coverage, the real driver is fear, not analysis. The research must separate "our customers are actually churning because of feature X" from "we feel threatened by the announcement."
  4. Bayesian updating. Start with a prior: "Our customers value feature X" (low confidence). Update with evidence: customer support tickets mentioning X (moderate signal), churn data correlated with X availability (strong signal), win/loss analysis mentioning X (strong signal).
  5. Scenario analysis. If churn is below 2% attributable to X → do not build, invest in roadmap. If churn is 2-5% → build a minimal response. If churn exceeds 5% → reprioritize roadmap.

The product leader gets a data-driven response plan calibrated to their actual situation, not a reactive "competitor launched X, we must match it" directive.

When to use constraint optimization vs. simpler approaches

Constraint optimization adds research overhead. It is not appropriate for every question:

FactorUse Constraint OptimizationUse Simple Research
Question typeDecision with hidden variablesFactual lookup
StakesHigh (irreversible, costly)Low (reversible, cheap)
Time horizonDecision plays out over monthsImmediate action
User constraintsUser-specific (risk tolerance, budget)Universal (objective facts)
Narrative riskHigh (popular opinion may mislead)Low (consensus is reliable)
Information valueMore research could change the answerCurrent information is sufficient

A factual question like "What is the capital of France?" needs no constraint optimization. A decision question like "Should we expand to France?" does, because the answer depends on the user's specific constraints, risk tolerance, and alternatives.

Comparison: constraint optimization vs. other research frameworks

DimensionConstraint OptimizationChain-of-ThoughtStandard SurveyExpert Consultation
Objective functionDecision utility under constraintsAnswer plausibilityData completenessExpert confidence
Handles sycophancyExplicitly (anti-narrative regularizers)Partially (reasoning transparency)Not addressedVaries by expert
Adversarial validationBuilt-in (step 8)Not built-inNot built-inDepends on process
Sensitivity analysisBuilt-in (step 6)Not built-inRarely includedExpert intuition
Residual trackingSix residual typesNot trackedNot trackedNot formalized
Iteration controlEvolver (continue/narrow/pivot/kill)Single passSingle passSession-based
Output formatConditional judgments with thresholdsReasoned conclusionData summaryRecommendation

Theoretical depth: the formal framework explained

The formal optimization framework (max_a U(a | E, H, C) − λR(a) − γI(a)) deserves deeper unpacking. Each component encodes a specific research discipline:

U(a | E, H, C) is the utility of action a given evidence E, hypotheses H, and constraints C. This is not a single number—it is a distribution over possible outcomes. The evidence may be incomplete, the hypotheses may be uncertain, and the constraints may be imprecise. The utility function captures this by producing a range (expected utility with confidence interval) rather than a point estimate.

λR(a) is the risk penalty. λ is the user's risk aversion coefficient—a high λ means the user penalizes risky actions heavily. R(a) is a measure of downside risk (variance, worst-case loss, drawdown). For a conservative investor, λ might be 3-5x the default. For a startup with nothing to lose, λ might be close to zero. The same evidence produces different recommendations for different λ values.

γI(a) is the uncertainty cost. γ penalizes actions where the evidence is weak. If the user has high γ (they hate making decisions without sufficient information), the framework favors "continue researching" over premature commitment. If γ is low (the user is comfortable with uncertainty), the framework recommends based on available evidence.

Practical defaults help users who have not thought about their risk tolerance. For reversible, low-stakes decisions (trying a new tool), λ ≈ 1 and γ ≈ 1 are reasonable defaults. For irreversible, high-stakes decisions (market entry, architectural choices), raising λ to 3-5 and γ to 2-3 produces more conservative recommendations that emphasize evidence strength.

The twelve-step method in depth

The twelve steps are not a checklist to complete mechanically. They are a structured reasoning process where each step catches a specific failure mode that the previous steps miss:

Steps 1-3 (objective reconstruction, fact/assumption splitting, quality vs. attractiveness) prevent the most common research failure: answering the wrong question. Most AI-generated research accepts the question's framing uncritically. "Is X good?" becomes a search for evidence that X is good, rather than an analysis of whether X is the right choice given the user's constraints.

Steps 4-5 (multi-start search, counterfactual perturbation) prevent anchoring bias. By starting from opposing positions and perturbing key assumptions, the research explores the solution space instead of converging on the first plausible answer.

Steps 6-8 (sensitivity analysis, implied expectations, adversarial validation) stress-test the emerging conclusion. Sensitivity analysis identifies which inputs matter most. Reverse-engineering implied expectations surfaces hidden assumptions. Adversarial validation argues the opposite position with full force.

Steps 9-12 (Bayesian updating, scenario analysis, personal constraints, anti-narrative regularizers) produce the final output. Bayesian updating shows how evidence shifts the posterior. Scenario analysis maps different futures. Personal constraints tailor the recommendation. Anti-narrative regularizers resist the pull of popular stories.

Residual-driven iteration: the stopping problem

The six residual types address a fundamental problem in research: when to stop. Without residuals, researchers either stop too early (missing critical evidence) or too late (diminishing returns on additional research). Residuals provide a principled stopping rule.

Question residual measures whether the question itself has been validated. If subsequent research reveals that the original question was misframed, the question residual stays high even if the evidence for the original question is strong. This triggers a pivot.

Evidence residual measures the gap between the minimum direct evidence defined in the evidence plan and the evidence actually collected. If the plan requires 3 case studies and only 1 has been found, the evidence residual is high regardless of how convincing the 1 case study is.

Adversarial residual measures whether the strongest counter-arguments have been addressed. If the red-team pass found a compelling objection that the synthesis does not adequately counter, the adversarial residual stays high.

Stopping rule: iteration continues until all residuals are below hard constraints, or until the value of information from another round is less than the delay cost of continuing. This prevents both premature convergence and indefinite research loops.

Boundary conditions: when constraint optimization fails

The framework has specific failure modes:

  • Unarticulated constraints. If the user does not know their own constraints (common for risk tolerance and time horizon), the framework produces conditional judgments that the user cannot evaluate. The research says "if your risk tolerance is X, do Y" but the user does not know their X.
  • Insufficient evidence base. For questions where public evidence is scarce (proprietary markets, emerging technologies), the framework cannot overcome the evidence gap. Bayesian updating with uninformative priors produces order-dependent results that vary with the sequence of evidence.
  • Straw-man adversarial validation. If the red-team pass argues against a weak version of the strongest argument instead of the strongest version, the adversarial validation is theater. The model's tendency toward balanced-sounding prose can produce straw-man critiques that do not genuinely challenge the conclusion.
  • Overhead on simple questions. Applying all twelve steps to "What is the best Python web framework for a blog?" is wasteful. The overhead of objective reconstruction, sensitivity analysis, and adversarial validation exceeds the value of the improved recommendation.
  • False precision in utility estimation. The formal framework suggests mathematical precision, but the utility function and penalty weights are rough estimates. Presenting a conditional judgment with specific thresholds ($45 CAC, 3-month localization) creates an illusion of precision that may not match the underlying evidence quality.
  • Narrative regularizer calibration. Anti-narrative regularizers resist popular stories, but the line between "popular narrative" and "well-supported consensus" is blurry. Over-correcting against narratives can lead to contrarian conclusions that are wrong for interesting reasons.

Limitations

  • Requires the user to articulate—or be willing to discover—their constraints. Many users do not know their own risk tolerance.
  • Adds research overhead. Simple factual questions do not need this framework.
  • Residual metrics use coarse discrete scales (0–3), not calibrated measurements.
  • Adversarial validation can degrade into straw-man arguments.
  • Bayesian updating requires meaningful priors; uninformative priors produce order-dependent results.
  • The framework is a skill, not core runtime. Quality depends on model capability and prompt engineering.

Engineering the theory: from a formula to executable gates

super-survey compiles the theory into a machine-checkable research pipeline. The utility function takes the form max U(a|E,H,C) − λR(a) − γI(a): maximize action utility under evidence E, hypotheses H, and constraints C, minus a risk penalty (weight λ) and an information-deficit penalty (weight γ). The action variable is not “bullish or bearish” but six concrete decisions: whether to open a position, how large, lump sum or staged, whether existing holders should stay, what conditions trigger re-evaluation, and whether a better alternative exists. When the user has not supplied a horizon or risk tolerance, the framework requires conditional judgments across λ/γ branches instead of pretending the weights are known.

Iteration is formalized as generalized descent on a seven-component residual vector (problem, constraints, evidence, hypotheses, adversarial, sensitivity, action): every round must compress a clear decision residual—piling up material does not count as descent. The operator order (problem projection → evidence observation → multi-path expansion → adversarial attack → synthesis update → direction choice) forms a non-commutative semigroup: reaching a conclusion first and gathering evidence afterwards is an after-action audit, not iteration. Stopping is doubly conditioned: total residual below threshold, and the information value of the next research action below its cost.

The key engineering decision is making gates hard constraints rather than advice. Three modes (quick / standard / deep) require 80 / 90 / 95 points with per-dimension minimums for sources, claims, and evidence; the 100-point final gate scores six dimensions (anti-sycophancy integrity 20, source and method 15, evidence completeness 20, analysis and red-team 20, actionability 15, structure 10), and any missing sub-score fails the report outright—so a high total cannot hide a report that simply accepted the user's original framing. Evidence lives in JSONL registries (sources / evidence / claims, linked by stable IDs), and the helper rejects weak pairs where a claim cites a number or entity that its linked evidence does not contain. Stage order is enforced by the CLI: the brief precedes evidence gathering, and skipping stages sends the agent back to the correct node.

FAQ

Can this be used outside finance?

Yes. Product opportunities, market entry, technical choices, open-source comparisons, due diligence—all are constrained decision problems.

Is this the same as "think step by step"?

Chain-of-thought improves reasoning transparency. Constraint optimization changes the objective function—from "produce an answer" to "build a decision system."

Does it always produce a decision?

No. If the value of information from further research exceeds the delay cost, the correct output is "continue researching," not a premature recommendation.

Can I see the research trail?

Yes. Super Survey produces persistent artifacts: brief, evidence plan, research, brainstorm, red-team, synthesis, evolver, and JSONL source/claim/evidence files.

How is this different from a survey tool?

Survey tools collect data. Constraint-optimized research builds decision systems. The output is a set of conditional actions, not a data summary.

What if I do not know my risk tolerance?

The framework provides conditional judgments under different weight branches. You can see "if you are risk-averse (λ=3), choose A. If you are risk-neutral (λ=1), choose B." Comparing the branches often clarifies your actual preference.

How does the evolver decide to continue vs. finalize?

The evolver evaluates residuals against hard constraints and estimates the value of information from another research round. If residuals are below thresholds or the expected improvement from more research is less than the delay cost, it finalizes.

Can multiple research rounds contradict each other?

Yes, and this is expected. A pivot (the evolver changing the research direction) produces a new framing that may contradict the previous round's conclusion. The artifact trail preserves both rounds so the user can see how the thinking evolved.

What are anti-narrative regularizers?

Prompts that instruct the agent to resist popular stories and conventional wisdom. For example: "What would the recommendation be if this were not a trending topic?" or "If this company were not well-known, would the evidence still support this conclusion?" They counteract the model's tendency to agree with widely repeated claims.

How does this relate to the Super Survey skill?

Super Survey is the implementation of constraint-optimized research as a Semibot skill. The twelve-step method, residual-driven iteration, and evolver are all implemented as staged workflow steps within the skill. The skill adds operational details (JSONL tracking, artifact persistence, companion routing) that the abstract framework does not specify.

Is the formal optimization formula actually computed?

The formula is a conceptual framework, not a numerical optimizer. The agent uses it to structure reasoning—"have I considered the risk penalty? the uncertainty cost?"—rather than computing exact values. The practical output is conditional judgments with qualitative confidence levels, not numerical utility scores.

How does this connect to the attention policy?

The attention policy determines when research results surface to the user. If the evolver decides to continue researching, the intermediate results are available but do not trigger notifications. Only when the evolver finalizes does the result cross the attention threshold for a push notification. This prevents research-in-progress from interrupting the user.

Related