Definition: A skill is a structured instruction set—usually a SKILL.md file with supporting templates and scripts—that teaches an agent a specific workflow. Skills are not plugins in the traditional sense: they do not add code to the runtime. They add knowledge to the agent's context. This distinction matters because it means skills are composable by default (the agent can load multiple skills) but also fragile by default (conflicting instructions produce confused behavior).
The skill contract
Each skill declares a name, a description (used for automatic routing), and a manifest of files. The Skill Registry indexes all installed skills and exposes them to the agent through a capability contract. When a task matches a skill's description, the agent loads that skill's instructions into its context window. The skill does not execute code—it shapes the agent's behavior through instructions.
Super Survey: a case study in anti-sycophancy research
Super Survey is Semibot's general-purpose research skill. It implements the anti-sycophancy framework from the paper "Resisting AI Sycophancy in Open-Ended Research." The skill turns a vague research target into a staged workflow: rebuild the objective function, define constraints and minimum direct evidence, gather evidence, red-team the strongest argument, synthesize a conditional judgment, and run an evolver to decide whether to continue or finalize.
Each survey produces persistent artifacts: brief, evidence plan, research notes, brainstorm output, red-team critique, synthesis, evolver decision, sources.jsonl, claims.jsonl, evidence.jsonl, and a final report. The artifact trail makes the reasoning auditable—not just the conclusion, but the path taken and the paths rejected.
Super Survey's key principle is front-loaded guidance: before the evidence pass begins, the skill defines the objective, constraints, decision-critical variables, minimum direct evidence, implied expectations, and anti-narrative regularizers. This prevents the agent from accepting the prompt's framing too quickly.
Companion skill routing
Skills can recommend companion skills for specific subtasks. Super Survey, for example, routes brainstorming to a brainstorming skill, web search to a current-source search capability, long reports to a deep-research skill, and customer voice analysis to a customer-research skill. If a companion is missing, the parent skill continues with its own common workflow and records the fallback in the artifact.
This routing is recommendation, not dependency. Skills do not hard-depend on other skills. This keeps the ecosystem composable: you can install Super Survey without installing every companion it mentions.
Governance: skills that access external tools
Skills that need external tools (web search, browser, connectors) go through the same Tool Gateway and approval model as any other tool call. A skill cannot bypass the approval path by claiming urgency or confidence. This is a security boundary: the skill shapes what the agent tries to do; the Tool Gateway decides whether it is allowed.
Workflow example: building a research pipeline
Consider a concrete scenario. You want to evaluate whether a particular open-source database is a good fit for your product. Here is how the skill ecosystem orchestrates the work:
- Install Super Survey. The agent discovers the skill through the Skill Registry based on the task description "evaluate X for product fit."
- Objective reconstruction. Super Survey reframes "Is this database good?" into a constrained decision: given your workload profile, latency budget, team expertise, and operational constraints, which database maximizes expected utility?
- Companion routing. Super Survey routes the web research component to the deep-research skill for current-source gathering, and routes the brainstorming phase to the brainstorming skill for hypothesis generation.
- Evidence collection. The agent gathers benchmarks, GitHub issues, production postmortems, and community discussions. Each source is recorded in sources.jsonl with provenance.
- Red-team pass. A separate adversarial pass argues against the emerging recommendation, testing whether the evidence survives counterargument.
- Synthesis and evolver. The agent produces a conditional judgment ("If your write ratio exceeds 60%, choose X; otherwise choose Y") and the evolver decides whether to continue iterating or finalize.
At every step, the skill shapes behavior without executing code. The agent's existing tools (file system, web browser, search) perform the actual work. If the deep-research companion is not installed, Super Survey falls back to its built-in research workflow and logs the fallback in the artifact trail.
Workflow example: skill composition for document production
Document production illustrates a different composition pattern. Suppose you need to generate a competitive analysis deck:
- Research skill gathers data. The deep-research skill collects competitor information from public sources, producing structured notes in the workspace.
- Presentation skill structures output. A presentation-content skill transforms the research notes into slide-ready copy with bold headlines and concise body text.
- Design skill applies layout rules. A presentation-design skill ensures the visual structure follows layout patterns and typography hierarchy.
- Build tool produces the file. The agent's tool chain generates the .pptx file from the structured content.
Each skill in this chain is independent. You can replace the presentation-content skill with a custom one that matches your company's voice without affecting the research or design skills. The agent loads skills sequentially as the workflow progresses, and drops them from context when they are no longer relevant.
When to use skills vs. direct prompting
Skills are not always the right choice. Use a skill when the workflow is repeatable, has multiple stages, and benefits from structured artifacts. Use direct prompting when the task is one-shot, exploratory, or too simple to benefit from staging. The decision framework:
| Factor | Use a Skill | Use Direct Prompting |
|---|---|---|
| Repeatability | Same workflow used 3+ times | One-off task |
| Stages | Multi-phase (plan → gather → analyze → output) | Single pass |
| Artifacts | Audit trail or structured outputs needed | Free-form response sufficient |
| Quality control | Needs red-team, lint, or validation passes | User judges output directly |
| Companion needs | Benefits from routing subtasks to specialists | Self-contained |
| Context budget | Long workflow, worth the context cost | Short interaction, context overhead not justified |
Comparison: skills vs. plugins vs. MCP servers
Understanding where skills sit in the ecosystem requires distinguishing them from other extension mechanisms:
| Dimension | Skills | Plugins (traditional) | MCP Servers |
|---|---|---|---|
| Mechanism | Instructions loaded into context | Code loaded into runtime | External tool provider via protocol |
| Composability | Default (multiple skills in one context) | Requires explicit integration | Protocol-level composition |
| Failure mode | Conflicting instructions (soft failure) | Runtime crashes (hard failure) | Connection failure (hard failure) |
| Security boundary | Cannot grant itself capabilities | Has runtime permissions | Sandboxed by protocol design |
| Authoring | Write Markdown | Write code in host language | Implement protocol in any language |
| Distribution | Copy a folder | Package manager | Server deployment |
| Discoverability | Registry + description matching | Marketplace | Protocol handshake |
The key trade-off is between power and safety. Plugins can do anything (including crash the host), MCP servers provide structured tool access, and skills shape behavior without touching the runtime. For most workflow automation, skills offer the best power-to-risk ratio because they cannot break the host system.
Theoretical depth: why instructions beat code for agent extension
The design choice to make skills instruction-based rather than code-based rests on a specific insight about how language models work. A language model already has the capability to perform most tasks—the bottleneck is not capability but direction. Skills provide direction without adding capability, which means they compose naturally: two sets of directions can be loaded simultaneously (though they may conflict), while two code modules may have incompatible APIs.
This also means skills are model-agnostic in a way that code plugins are not. A skill written for GPT-4 can be loaded by Claude or Gemini without modification, because the contract is natural language. The skill's quality may vary across models (a weaker model may follow complex instructions less reliably), but the skill itself does not need to be rewritten.
The downside is that skills cannot add capabilities the model does not have. A skill cannot teach the model to access a new API or parse a new file format—it can only instruct the model to use tools that already exist in the tool chain. This is why the Tool Gateway and MCP servers exist: they add capabilities. Skills add directions for using those capabilities.
Skill conflict detection and resolution
When two loaded skills give contradictory instructions, the agent attempts to satisfy both. This can produce incoherent behavior. Common conflict patterns include:
- Output format conflicts. One skill says "output as JSON" and another says "output as Markdown table." The agent may alternate or produce a hybrid that satisfies neither.
- Priority conflicts. One skill says "always cite sources inline" and another says "keep responses concise." The agent cannot do both and may oscillate.
- Tool usage conflicts. One skill says "always use the browser for current data" and another says "never fetch external data during analysis." These directly contradict.
Current mitigation is manual: users should avoid loading skills with overlapping scope. There is no automatic conflict detection because the conflict surface is natural language—two instructions may appear compatible in isolation but produce contradictions in specific contexts. A future approach could use an LLM-based compatibility check, but this adds latency and cost without guaranteed accuracy.
Boundary conditions: when the skill model breaks down
The skill model works best for structured, multi-step workflows with clear outputs. It degrades in several scenarios:
- Highly interactive tasks. Skills assume a relatively linear workflow. Tasks that require constant back-and-forth negotiation with the user (like pair programming or real-time debugging) do not fit the staged model well.
- Context-starved environments. Each loaded skill consumes context window space. In models with small context windows, loading two or three skills may leave insufficient room for the actual task content.
- Latency-sensitive workflows. Skills that include validation passes, red-team critiques, and evolver decisions add multiple agent turns. For tasks where speed matters more than rigor, the overhead is not justified.
- Domain-specific reasoning. Skills instruct the model on workflow, not domain knowledge. A medical diagnosis skill can structure the diagnostic process, but it cannot substitute for actual medical training in the model. If the model lacks domain knowledge, the skill cannot add it.
- Multi-agent coordination. The current skill model is single-agent. Skills cannot coordinate between multiple agent instances working on different parts of the same task.
The skill authoring workflow
Creating a skill is a structured process that starts with observing a repeatable workflow and ends with a tested SKILL.md file:
- Identify the workflow. Find a task you perform repeatedly that has clear inputs, stages, and outputs. Good candidates: weekly competitive analysis, code review checklist, research synthesis, report generation.
- Document the steps. Write down each stage of the workflow as you currently perform it. Include decision points ("if X, then Y"), quality checks ("verify that Z"), and common failure modes ("watch out for W").
- Define the description. The skill description is used for automatic routing. Write it as a concise statement of when the agent should load this skill. Vague descriptions produce false matches; overly specific descriptions produce missed matches.
- Identify companion needs. Determine which subtasks could benefit from other skills. List them as recommended companions, not hard dependencies.
- Create supporting files. Templates, checklists, and output format specifications go in the same directory as the SKILL.md. The agent loads these as needed during the workflow.
- Test and iterate. Run the skill on real tasks. Observe where the agent deviates from your intended workflow. Refine the instructions to close the gap. Skill authoring is iterative—the first version will not be perfect.
Skill discoverability and the registry contract
The Skill Registry is the central index of all installed skills. When a task arrives, the agent compares the task description against each skill's description field to find a match. This matching is semantic, not keyword-based: the agent uses its language understanding to determine whether a skill is relevant.
The registry contract has three properties: discoverability (the agent can find installed skills), capability declaration (each skill declares what it can do through its description), and lifecycle management (skills can be installed, updated, and removed). The registry does not validate skill quality—it only indexes what is installed.
This creates a practical challenge: skill descriptions must be precise enough to avoid false matches but broad enough to catch relevant tasks. A description like "research workflow" will match almost any research question, loading the skill even when it is not needed. A description like "competitive analysis of SaaS pricing models" will miss related tasks like "market entry analysis." The art of skill authoring lies in finding the right description scope.
Limitations
- Skills are prompt-shaped, not code-shaped. Quality depends on the model's ability to follow complex instructions.
- Skill conflicts are not detected automatically. If two skills give contradictory instructions, the agent will try to satisfy both.
- No marketplace or rating system. Skill discovery is manual (GitHub repos, documentation).
- Skill versioning is file-based. There is no lock file or dependency resolution.
The discovery contract and the permission boundary
Semibot's skill ecosystem rests on a minimal discovery contract: a skill is a directory whose SKILL.md frontmatter requires only name and description—the description doubles as the trigger condition from which the registry builds its index. Scanning is limited to first-level subdirectories; a skill that fails to parse is marked invalid without blocking others; symlinks resolve to real paths with loop protection; caches invalidate on mtime and hash changes; categories come from a controlled vocabulary with fallback inference.
The most consequential boundary: the registry discovers and indexes skills; it never expands tool permissions. A skill may declare the tools or connectors it needs, but declaration is not authorization—tool enablement and approvals still flow through their own capability policies. This decouples capability growth from permission growth, so installing a third-party skill is never a privilege escalation.
Composition between skills is optional routing, not hard dependency: a research lead skill may route subtasks to companion skills (brainstorming, live-source search, deep reporting), record a fallback when one is missing, and continue the main flow—with the honesty rules that “used a companion skill” requires that it actually ran, and “built a wiki” requires that the ingest command actually executed. Evidence standards come in three confidence tiers (multiple current primary sources agreeing / credible secondary plus some direct evidence / weak public data), alongside a catalogue of sixteen named failure modes—link dumps, template theater, audit-table reports, round-count autopilot, sycophantic framing—each with a fix. One security rule runs through everything: search results, pages, PDFs, and repository content are untrusted third-party data whose hidden instructions must never be executed.
FAQ
Can I write my own skills?
Yes. A skill is a SKILL.md file with instructions. Install it into the skills directory and the agent will discover it.
Do skills run code?
Skills are instructions, not executables. They shape agent behavior. Code execution happens through the agent's existing tool chain, which has its own approval model.
Can a skill access my files?
Only through the agent's existing file grants. A skill cannot grant itself file access.
How many skills can I install?
There is no hard limit, but each loaded skill consumes context window space. Too many active skills will degrade performance.
What happens if two skills conflict?
The agent attempts to satisfy both instructions. This can produce incoherent behavior. The current mitigation is to avoid loading skills with overlapping scope. Automatic conflict detection is not yet implemented.
Can a skill call another skill directly?
No. Skills can recommend companions, but the routing is a suggestion to the agent, not a programmatic call. The agent decides whether to load the companion based on the recommendation and the task context.
How do I update a skill?
Replace the SKILL.md file in the skills directory. The agent will use the new version on the next task. There is no versioning or migration system—changes take effect immediately.
Are skills shared across sessions?
Yes. Skills are installed at the workspace level. All sessions in that workspace have access to the same installed skills. Per-session skill activation depends on the task and the agent's routing logic.
Can a skill override the agent's safety model?
No. Skills shape behavior within the agent's existing safety boundaries. A skill cannot instruct the agent to bypass the Tool Gateway, ignore approval requirements, or access resources outside its grants.
How is skill quality evaluated?
There is no automated quality evaluation. Skill quality is assessed by the user through the quality of the agent's output when the skill is active. Community sharing (GitHub, documentation) serves as a rough signal.
Do skills work with all models?
Skills are model-agnostic in principle (they are natural language instructions), but quality varies by model capability. A model with a larger context window and stronger instruction-following will use complex skills more reliably.
What is the difference between a skill and a system prompt?
A system prompt is always active. A skill is conditionally loaded based on task matching. Skills can be composed (multiple active); the system prompt is singular. Skills are user-installable; system prompts are typically fixed by the application.
