SemibotSemibot - AI Desktop
All guides

Specialists vs General Chat: Why Domain AI Agents Remember How You Work

A specialist is a named, persistent configuration that remembers how a recurring job should be done—its methods, tools, and quality standards. You create it through conversation, not configuration files.

Definition: A specialist (internally called a Bot) is a named, persistent AI configuration that remembers how a recurring job should be done. Unlike a general chat session where you re-explain context every time, a specialist stores its method, required skills, and behavioral constraints. You create one through conversation—"I need a research assistant that follows this methodology"—and then hand it work by name.

The problem with general chat

General chat is the right interface for one-off questions. It is the wrong interface for recurring work. If you ask an AI to summarize your weekly progress every Friday, you will re-explain the format, the sources, and the audience every time—or hope it remembers from context. Specialists solve this by separating the "how" from the "what": the specialist knows how to do the work; you only tell it what to do this time.

The five-part Bot definition

A specialist stores five things: identity (name and description), behavioral constraints (what it should and should not do), required skills and tools, method (how it approaches the work), and default workspace bindings. This definition is created through conversation—you do not write configuration files. The Bot Studio session compiles these five parts into a Harness that is injected into each Runtime session.

Harness compilation

The Harness is a pure TypeScript function—no model calls, no text parsing. At the start of each Runtime run, the compiler takes the published Bot definition and produces a context pack that includes the specialist's method, constraints, required tools, and workspace access. The run then behaves as if the specialist's knowledge is part of the system prompt. This is deterministic: the same definition always produces the same Harness.

Skill dependency checking

Before a specialist sends a message, it checks whether its required skills are installed and enabled. If a skill is missing, the system offers a one-click install rather than failing silently. This prevents the common failure mode where a specialist is configured to use a capability that is not actually available.

Work Item lifecycle

Each piece of work assigned to a specialist creates a Work Item. The Work Item tracks five states: active, awaiting user, waiting external, completed, and stopped. Completion is not determined by the model saying "done"—it requires evidence: successful tool results, ready artifacts, approval completion, or external side-effect confirmation. This prevents premature completion claims.

Why specialists are not a separate agent kind

A deliberate design choice: specialists reuse the Main Runtime, not a separate agent engine. This means they have access to the same tools, file system, approval model, and conversation interface as any other session. The trade-off is that specialists are not fully isolated—they share the same runtime capabilities. The benefit is simplicity: no separate agent infrastructure, no separate tool registry, no separate approval path.

Workflow examples

Specialists are most valuable when a recurring job has a stable method but variable inputs. Here are three scenarios that illustrate the lifecycle from creation to ongoing use.

Scenario 1: Creating a weekly-report specialist. Step 1: Open Bot Studio and describe the work: "I need a specialist that compiles a weekly progress report. The sources are git commits, the project task board, and my calendar. The format is: accomplishments, blockers, next-week priorities. Audience: my manager." Step 2: The Bot Studio session asks clarifying questions—which repositories, which task board, what tone. Step 3: It compiles the five-part definition (identity, constraints, skills, method, workspace bindings). Step 4: You publish the specialist. Step 5: Every Friday, you assign the specialist a Work Item: "Compile this week's report." Step 6: The specialist checks its required skills (git log, task board access, calendar read), executes its method, and produces the report. Step 7: You review, make minor edits, and send.

Scenario 2: A/B test analysis specialist. Step 1: You describe in Bot Studio: "I need a specialist for analyzing A/B test results. It should calculate statistical significance, check for sample ratio mismatch, identify novelty effects, and produce a recommendation with confidence intervals." Step 2: The specialist is compiled with the data-analytics skill as a required dependency. Step 3: When you have new test results, you assign a Work Item with the dataset. Step 4: The specialist runs its method—loading data, computing metrics, checking assumptions, producing output. Step 5: If the data-analytics skill is not installed, the system offers one-click install before proceeding. Step 6: The result includes the statistical analysis, a plain-language recommendation, and a note on any assumption violations.

Scenario 3: Combining specialist and secretary. Step 1: You create a "competitor watcher" specialist that knows how to analyze competitive moves—pricing changes, feature launches, hiring patterns. Step 2: You delegate to the secretary: "Use the competitor watcher to monitor Competitor X weekly." Step 3: The secretary handles the scheduling and change detection. The specialist handles the analysis method. Step 4: Each week, the secretary assigns a Work Item to the specialist, collects the result, and delivers it through the secretary page. Step 5: If the specialist detects a significant competitive move, the attention score crosses the notification threshold. This combination separates concerns: the secretary manages "when" and "how often"; the specialist manages "how."

Decision framework: general chat vs specialist vs secretary

QuestionUse general chatCreate a specialistUse the secretary
Is this a one-off task?YesNoNo
Is there a stable method?Not neededYes, method is reusableMethod may be simple or delegated
Does it need change detection?NoNot necessarilyYes
Should results accumulate?NoOptionalYes
Do you need specific skills?Ad hocYes, required skills are declaredSkills come from the assigned specialist
Example"Explain this error message""Analyze A/B test results using our standard methodology""Monitor competitor X and summarize changes weekly"

The three models are not mutually exclusive. A secretary task can use a specialist. A general chat session can spawn a one-shot specialist for a complex task. The design goal is to match the interface to the task's temporal and methodological structure.

Comparison with other approaches to persistent AI configuration

DimensionSystem prompt / custom instructionsLangChain agentGPTs / custom chatbotsSemibot specialists
Creation methodEdit text configWrite codeWeb form + instructionsNatural language conversation
Deterministic compilationText injection onlyCode is deterministicNot compiledHarness is a pure TypeScript function
Skill dependency checkingNoneManualPlugin marketplaceAutomatic with one-click install
Workspace accessNone (API only)Code-definedLimited (file upload)Full file system and tool access
Approval modelNoneCustom-codedNoneIntegrated with Semibot's approval gates
Work Item trackingNoneCustom implementationNoneBuilt-in with evidence-based completion

Failure modes and boundary conditions

Specialists are powerful but not universally applicable. Understanding where they break prevents misapplication:

  • Vague method specification. The most common failure: "I need a specialist for marketing." Without a specific method—what sources to check, what format to produce, what quality standards to apply—the specialist is a general chatbot with a name. The Bot Studio session helps by asking clarifying questions, but the user must have a clear enough workflow to articulate.
  • Over-specialization. A specialist that is too narrowly defined becomes brittle. If the A/B test analysis specialist assumes a specific data format, it fails when the format changes. The method description should be specific about approach but flexible about inputs.
  • Harness limitations. The Harness compiler is intentionally simple—it does not call models or discover tools. Complex routing logic (if the data is in format A, use approach X; if format B, use approach Y) must be encoded in the method description as natural language instructions, not as programmatic branching.
  • No version history. Publishing a new definition overwrites the previous one. If a published change introduces a regression, you must manually reconstruct the previous version. In-progress Work Items retain their frozen revision, but future Work Items use the new definition immediately.
  • Shared runtime. Specialists share the same runtime as general chat. This means they have the same file system access, the same tool registry, and the same potential for errors. There is no sandbox isolation between a specialist and the rest of the system.
  • Cold start problem. A new specialist has no feedback history. Its first few runs may produce results that do not match your expectations. The feedback loop (Work Item results, user corrections) takes several iterations to calibrate the specialist's behavior.

Theoretical depth: the separation of method from task

The specialist model is grounded in a principle from software engineering: the separation of mechanism from policy. In operating systems, this means the kernel provides primitives (memory allocation, scheduling) while user-space programs decide how to use them. In AI agents, the specialist provides the mechanism (method, tools, constraints) while the user provides the policy (what to do this time, with what data).

This separation has a specific practical benefit: it reduces the cognitive cost of delegation. When you assign a Work Item to a specialist, you do not need to explain the method—you only specify the input. This is analogous to how you tell a skilled employee "run the weekly report" without re-explaining the report format, data sources, or quality checks each time. The specialist's stored method acts as a persistent instruction set that survives across sessions, unlike a system prompt that must be re-injected.

The Harness compilation model draws on ahead-of-time compilation from programming language theory. By compiling the Bot definition into a deterministic context pack before each run, the system ensures that the specialist's behavior is reproducible—the same definition with the same input produces the same contextual framing. This is different from prompt-engineering approaches where the model's interpretation of instructions can drift across sessions.

Limitations

  • Specialists are user-level definitions, not workspace-level. They do not bind to a specific project until assigned a Work Item.
  • No version history for Bot definitions. The published snapshot is overwritten on each publish. Old runs retain their frozen revision.
  • The Harness compiler is intentionally simple—it does not call models or discover tools. Complex routing logic must live in the method description, not in the compiler.
  • Creating a good specialist requires clear communication of method and constraints. Vague instructions produce vague specialists.

Completion determination: a five-condition evidence gate

A specialist's definition is itself a constrained data structure: five parts (identity, responsibility boundary, operating guide, capability references, completion standard), all natural-language fields. Saving code, scripts, prompts, tool schemas, model configuration, or any credential is prohibited by design—a specialist remembers how you work, not a permission set. Field lengths and item counts have hard limits; an over-limit snapshot is rejected whole, keeping the previous version rather than being silently truncated.

The biggest mechanical difference from general chat is completion determination. Marking a specialist's work item complete automatically requires five evidence conditions simultaneously: the latest goal is judged achieved with a usable result; the completion standard was frozen into the goal before the run started (no moving the goalposts after the fact); at least one traceable result exists (a readable artifact, an external action confirmed by the tool gateway, or user confirmation); no pending approvals, failed required tools, delivery-unknowns, or unverified side effects remain; and file deliveries pass integrity and readability checks—reporting a filename or a Markdown link does not count. A model saying “done” cannot change the item's state. These five conditions turn completion from a linguistic phenomenon into an evidentiary claim.

Execution emphasizes determinism in the same spirit: each run's instruction context is emitted by a pure-function compiler that calls no model, parses no chat text, enforces a hard length cap (over-limit fails closed instead of truncating), and targets p95 compile time under 5 ms. Publishing uses an optimistic-lock atomic swap with explicit conflict errors; work-item deduplication keys on a stable hash of the source object—“similar natural-language titles” never produce a dedupe key. The system keeps no progress percentages or state-machine nodes: an item has five user-facing states, and everything else is ordinary work happening in a normal conversation.

FAQ

Can I create a specialist through conversation?

Yes. Describe the work, its method, and constraints in the Bot Studio session. The system compiles the definition from the conversation.

Can a specialist use the same tools as my regular chat?

Yes. Specialists use the same tool registry and approval model. No separate tool configuration is needed.

What happens if a required skill is missing?

The system detects the missing skill before sending and offers a one-click install. The message and attachments are preserved.

Can I modify a specialist after creating it?

Yes. Open the Bot Studio session, describe the changes, and republish. In-progress runs keep their frozen revision; new runs use the updated definition.

How many specialists can I create?

There is no hard limit. Each specialist is a stored definition—a lightweight object that does not consume resources until assigned a Work Item. You can have dozens of specialists without performance impact. The practical limit is your ability to maintain and remember them.

Can I share a specialist with my team?

Specialists are currently user-level definitions. They are not workspace-level or team-level. Sharing would require exporting the Bot definition and importing it in another account. This is a known gap and an area of active consideration.

What does it cost to run a specialist?

Each Work Item execution consumes model calls, the same as a general chat session. The cost depends on the task complexity and the number of tool calls. A weekly report specialist might use 3–5 model calls per execution. A complex research specialist might use 15–20. The Harness compilation itself is deterministic and costs nothing.

How does the specialist interact with the secretary?

The secretary can assign Work Items to specialists on a schedule or in response to changes. The specialist handles the method; the secretary handles the scheduling, change detection, and delivery. This separation of concerns is the recommended pattern for recurring analytical work.

Can a specialist use the browser or arxiv skills?

Yes. Any skill available in Semibot can be declared as a required skill in the Bot definition. The skill dependency checker verifies availability before each run. If the skill is not installed, the system offers a one-click install.

What happens if the specialist's method is outdated?

The specialist executes whatever method is in its published definition. If the method references outdated tools, data formats, or workflows, the results will be incorrect. Periodic review of specialist definitions—especially after process changes—is the user's responsibility.

Can I create a specialist for a task I have never done manually?

You can, but the result will be weaker. A specialist works best when its method is derived from a workflow you have already validated. Creating a specialist for an untested workflow means the method description is based on your assumptions rather than experience. Start with manual execution, validate the approach, then formalize it as a specialist.

Related