Definition: Semibot's knowledge library is built on the Karpathy LLM Wiki protocol: raw source materials are stored in a raw/ directory, and the agent maintains organized topic articles in a wiki/ directory. This is not a one-shot "generate a report" workflow—it is continuous knowledge maintenance. The agent can ingest new materials, update existing articles, and keep a revision log, all while the user watches knowledge take shape.
Three modes in one workspace
The knowledge library has three capabilities in the same workspace:
- Browse: Real folder tree with file preview. Preserves the user's mental model of directories and files.
- Search: Hybrid semantic search across one or all workspaces. Finds relevant content from raw materials and wiki articles.
- Wiki: Continuously organized topic articles with traceable sources, maintained by the agent using the Ingest/Query/Lint lifecycle.
The raw/wiki protocol
The protocol is simple:
- raw/: Source materials fetched and saved by the skill, organized by topic.
- wiki/: Knowledge articles maintained by the skill.
- wiki/index.md: A human-readable global index.
- wiki/log.md: The skill's maintenance history and breakpoint notes.
Ingest fetches new materials. Triage decides what to do with them. Compile produces or updates articles. Query answers questions by referencing wiki content. Lint checks and repairs existing articles. The skill operates directly on these directories—there is no staging area, no generation pipeline, no automatic rollback.
User mental model: watching knowledge take shape
Users may briefly see half-finished states during Wiki updates. This is intentional: it reflects the real working state of knowledge maintenance. The alternative—hiding all intermediate work and presenting only finished articles—creates a generate-and-forget cycle where the user never builds a mental model of what the knowledge base contains or how it evolved.
The search projection
The search index is a derived projection: it can be rebuilt from raw/ and wiki/ at any time. If the search projection fails to update, the wiki update is not affected. This separation means knowledge maintenance and search indexing can fail independently without corrupting each other.
Local-first knowledge
The wiki lives in the workspace's .semibot/knowledge/ directory. The agent's working directory is fixed to this location. Raw materials, articles, index, and logs are all local files. If the user later switches to another tool, the raw files are still on the machine.
Scenario: onboarding 50 research papers
Consider a concrete workflow. You have 50 research papers on transformer architectures and want an organized knowledge base:
- Step 1 — Ingest. You place the 50 PDFs into the raw/ directory, organized by subtopic (raw/attention-mechanisms/, raw/efficiency/, raw/scaling-laws/). The agent reads each paper and extracts key claims, methods, and results.
- Step 2 — Triage. The agent classifies each paper by topic relevance. Some papers span multiple topics and are tagged accordingly. Papers that are too tangential are noted but not forced into articles.
- Step 3 — Compile. The agent creates wiki articles for each major topic: wiki/attention-mechanisms.md, wiki/efficiency-techniques.md, wiki/scaling-laws.md. Each article synthesizes findings from multiple papers, with inline citations linking back to the raw/ source files.
- Step 4 — Index. The agent updates wiki/index.md with a human-readable table of contents and updates wiki/log.md with the maintenance history.
- Step 5 — Query. You ask "What is the relationship between sparse attention and inference cost?" The agent answers by referencing wiki/attention-mechanisms.md and wiki/efficiency-techniques.md, citing specific papers from raw/.
- Step 6 — Lint. When you add 5 new papers two weeks later, the agent runs a lint pass on existing articles, checking whether the new papers introduce contradictions, fill gaps, or require expanding existing sections.
The critical property is that this is not a one-shot "generate a report" workflow. The knowledge base grows over time. Each ingest cycle builds on the existing wiki structure rather than starting from scratch.
Scenario: maintaining a competitive intelligence wiki
A product team wants to track 10 competitors. Each week, the secretary monitors competitor websites, blogs, and changelogs. Here is how the wiki model handles this:
- Initial ingest. The agent creates raw/ directories per competitor with existing public information: pricing pages, feature lists, blog posts, and press releases.
- Wiki compilation. The agent produces wiki articles per competitor and cross-cutting articles on themes like wiki/pricing-comparison.md and wiki/feature-gap-analysis.md.
- Weekly update cycle. Each week, the secretary detects changes on competitor sites. New content is added to raw/. The agent runs a lint pass that updates the relevant wiki articles, flagging what changed and how it affects the existing analysis.
- Query-driven insights. When the product team asks "Which competitor has the strongest API offering?" the agent answers from the wiki, citing specific features and sources.
Over months, the wiki becomes a living document that tracks the competitive landscape. Each article has a revision history in log.md that shows when information was added and what changed. The team can see not just the current state but the trajectory.
Scenario: personal learning wiki for a new technology
An individual developer wants to learn Rust. Instead of generating a static learning plan, the wiki model creates a growing knowledge base:
- Ingest learning materials. Place blog posts, documentation pages, and code examples into raw/. The agent organizes them by topic: raw/ownership-model/, raw/concurrency/, raw/error-handling/.
- Compile learning wiki. The agent produces articles that synthesize the raw materials into coherent explanations: wiki/ownership-and-borrowing.md, wiki/fearless-concurrency.md, wiki/result-and-option-patterns.md.
- Iterative deepening. As you learn more and add new materials, the agent updates existing articles. Early articles may be shallow; over time, they deepen as more sources are ingested.
- Cross-reference building. The agent automatically creates cross-references between articles. The concurrency article links to the ownership article where relevant. These links emerge from the content, not from manual tagging.
The result is a personalized reference that reflects your actual learning path, not a generic curriculum. The raw/ directory preserves every source you consulted; the wiki/ distills them into organized knowledge.
When to use the Wiki model vs. alternatives
The Wiki model is one approach to knowledge management. It excels in some scenarios and is the wrong choice in others:
| Factor | Use Wiki Model | Use Alternative |
|---|---|---|
| Source volume | 10+ sources that benefit from synthesis | Fewer than 10 sources — direct reading is faster |
| Temporal span | Weeks to months of ongoing research | One-time analysis — use a report instead |
| Cross-referencing | Topics overlap and interconnect | Topics are independent — flat file structure suffices |
| Source traceability | Must link claims back to sources | Sources are ephemeral — summaries suffice |
| Team sharing | Multiple people need to read and extend | Solo consumption — chat history is fine |
| Revision tracking | Need to see how knowledge evolved | Only current state matters |
Comparison: Wiki model vs. other knowledge management approaches
| Dimension | Wiki Model (Semibot) | RAG-only | Chat History | Static Report |
|---|---|---|---|---|
| Knowledge structure | Organized topic articles | Chunked documents | Chronological messages | Linear narrative |
| Maintenance | Continuous (agent-driven lint) | None (re-index on change) | None (append-only) | Manual rewrite |
| Source linking | Inline citations to raw/ | Retrieved chunks (opaque) | Lost after session | Footnotes (if authored) |
| Browsability | Folder tree + search + wiki | Search only | Scroll only | Table of contents |
| Revision history | log.md (agent-maintained) | None | Implicit (message order) | Manual changelog |
| Staleness handling | Lint pass detects and resolves | No detection | N/A (snapshot in time) | Manual update |
| Data location | Local workspace (.semibot/knowledge/) | Usually cloud-hosted | App database | User-chosen location |
Theoretical depth: why continuous maintenance beats generate-and-forget
The generate-and-forget model (produce a report, never update it) has a specific failure mode: the report's accuracy degrades monotonically as the underlying information changes. Within weeks, the report is a historical artifact, not a knowledge resource. Users learn not to trust it and stop consulting it.
The Wiki model addresses this by treating knowledge as a living system. The Ingest/Query/Lint lifecycle creates a maintenance loop where new information continuously flows into the knowledge base. The lint pass is the key differentiator: it does not just add new articles—it checks whether existing articles are still accurate in light of new sources.
This mirrors how human-curated wikis (like Wikipedia) maintain quality. The difference is that the agent performs the maintenance autonomously, with the user watching the knowledge take shape rather than performing the edits manually. The trade-off is that the agent may occasionally make organizing decisions the user disagrees with—hence the ability to manually edit files and the expectation that the agent will see those edits on the next cycle.
The Ingest/Query/Lint lifecycle in detail
The lifecycle is the operational core of the Wiki model. Each phase has specific responsibilities and failure modes:
Ingest fetches and saves new materials. It can be triggered by the user (manual upload), by the secretary (detected change in a monitored source), or by the agent itself (follow a citation chain). Ingest is passive in the sense that it does not organize—it only adds to raw/. If ingest fails (network error, parsing failure), the raw/ directory simply lacks that source, and the wiki continues with what it has.
Triage decides what to do with new materials. The agent classifies each item by topic relevance and decides whether to create a new article, update an existing article, or note the material as tangential. Triage is where the agent's judgment matters most—poor triage leads to fragmented articles (too many narrow topics) or monolithic articles (everything dumped into one page).
Compile produces or updates wiki articles. The agent synthesizes information from multiple raw sources into coherent topic articles with inline citations. Compile does not delete—it adds, reorganizes, and refines. If the agent determines that a source contradicts existing articles, it flags the contradiction rather than silently overwriting.
Query answers questions by referencing wiki content. The agent searches both raw/ and wiki/ to find relevant information, prioritizing wiki/ for synthesized knowledge and raw/ for source details. Query does not modify the wiki—it is a read-only operation.
Lint checks and repairs existing articles. Lint is triggered when new sources are ingested or periodically by the agent. It checks for outdated claims, missing citations, broken cross-references, and contradictions between articles. Lint is what keeps the wiki alive—without it, the wiki degrades into the same staleness as a static report.
Boundary conditions: when the Wiki model does not work
The Wiki model is powerful but has clear boundaries:
- Thin sources. If the raw/ directory contains only a few low-quality sources, the agent cannot produce meaningful wiki articles. The "No material" result is a valid and correct outcome—the skill does not hallucinate content to fill gaps.
- Highly dynamic information. For information that changes minute-by-minute (stock prices, live sports scores), the wiki model's maintenance cycle is too slow. Real-time data needs dashboards, not articles.
- Deep domain expertise required. If the source material requires expert-level interpretation (medical imaging, legal analysis), the agent's synthesis may miss nuances that a domain expert would catch. The wiki is a useful starting point but not a substitute for expert review.
- Adversarial or conflicting sources. When sources directly contradict each other (competing studies with opposite conclusions), the agent flags the contradiction but may not resolve it correctly. Human judgment is needed to weigh conflicting evidence.
- Scale limits. For very large knowledge bases (thousands of articles), the lint pass becomes expensive and the index may not capture all cross-references. The model works best for knowledge bases of 10 to 500 articles.
Source quality and the garbage-in problem
The Wiki model inherits the quality of its sources. If raw/ contains low-quality, outdated, or contradictory materials, the wiki articles will reflect those problems. The agent cannot fabricate knowledge from thin sources—it can only organize and synthesize what is available.
This creates a practical discipline: the quality of the knowledge base depends more on what you put into raw/ than on the agent's organizing ability. Curating sources before ingest produces better results than ingesting everything and hoping the agent sorts it out. The triage step helps, but it cannot compensate for fundamentally poor source material.
The agent's lint pass partially addresses this by flagging contradictions between sources. If two raw sources make conflicting claims, the lint pass can note the contradiction in the wiki article and leave it for human resolution. This is more honest than silently choosing one source over another.
Multi-workspace knowledge management
Semibot supports multiple knowledge workspaces, each with its own raw/wiki structure. This is useful when knowledge domains are distinct and should not be mixed:
- Domain separation. A product team's competitive intelligence wiki should not be mixed with a personal learning wiki on machine learning. Separate workspaces keep the agent's context focused.
- Access control. Workspaces can have different sharing settings. A team wiki might be shared via version control; a personal wiki stays local.
- Lifecycle independence. Different workspaces can have different ingest schedules. A competitive intelligence wiki might be updated weekly; a personal learning wiki might be updated daily during active study.
- Cross-workspace search. Hybrid semantic search can query across all workspaces when needed, while still maintaining the organizational boundaries of individual wikis.
Limitations
- Wiki quality depends on the model's ability to organize and summarize. Poor source material produces poor articles.
- No human editor. The agent maintains articles autonomously. Users can manually edit files, but there is no collaborative editing interface.
- The "No material" result is a valid outcome—the skill does not force article creation from thin sources.
- Revision history is log-based, not version-controlled. There is no diff view or rollback UI.
Direct maintenance and the separation of truth from derivatives
The core architectural decision in Semibot's knowledge wiki is “direct maintenance”: the raw sources (raw/) and the knowledge articles (wiki/) are the only source of truth—no staging area, no generate-then-publish, no automatic rollback. The organizing process is visible to the user in real time, and a half-finished state mid-update is treated as the wiki's honest working state, not a defect. An interrupted run keeps what finished and resumes in place. The bet is that transparency beats tidiness: what you see is what the system is actually doing.
The organizing contract borrows from Karpathy's LLM-wiki practice: two anchor files (a global index and a maintenance log), material flowing through “fetch → triage → compile on demand” into articles, each article maintaining source citations and cross-links, and exactly four legal conclusions per source—new knowledge, update, disputed, or no material (the last being a legitimate outcome, not a failure). Provenance is first-class in the UI: articles show sources and update times, and clicking a source opens the original document.
Two implementation details deserve separate mention. First, SQLite stores only run state and per-source revision cursors; the search index and embeddings are derivable projections, rebuildable from the files at any time—the truth is always the Markdown. Second, incremental updates create an independent processing session per source, with prompts containing only the change type and relative paths, never embedded document bodies—a direct defense against the “stuff hundreds of filenames into one prompt and get fake holistic processing” failure mode. On retrieval, wiki articles act as trusted system sources and are searched alongside raw materials in hybrid retrieval (semantic plus keyword, rank-fused); the wiki never replaces the originals—both recall paths coexist, and conclusions remain traceable to either.
FAQ
Is this the same as a RAG system?
RAG retrieves and passes context to the model. The Wiki model maintains organized articles as a persistent knowledge structure. They can work together: the Wiki provides structured knowledge, RAG provides raw search.
Can I edit the wiki articles?
Yes. They are regular Markdown files on disk. The agent will see your edits on the next ingest cycle.
What happens to old articles when new information arrives?
The Lint cycle checks existing articles against new sources and can update, expand, or flag contradictions. Articles are not frozen.
Does the wiki stay local?
Yes. The wiki is in the workspace's local .semibot/knowledge/ directory. Only model API calls (if using cloud models) involve network.
How does the agent decide what becomes a separate article vs. a section in an existing article?
During triage, the agent considers topic coherence. If new material extends an existing topic, it becomes a section. If it introduces a distinct topic with its own evidence base, it becomes a new article. Cross-cutting material may trigger the creation of a synthesis article that references multiple existing articles.
Can multiple agents maintain the same wiki?
The current model assumes a single agent per workspace. Multiple agents could produce conflicting edits. For team use, the wiki files can be shared via version control, but the maintenance lifecycle is designed for single-agent operation.
What format are wiki articles in?
Standard Markdown. The agent uses consistent formatting conventions (headers, lists, inline citations) but the format is human-readable and editable with any text editor.
How does the agent handle sources in different languages?
The agent can ingest sources in multiple languages and produce wiki articles in a single target language (typically the user's preferred language). Citations preserve the original source language.
Is there a size limit for raw/ files?
Individual files are limited by the model's context window for processing. Very large files (hundreds of pages) may need to be chunked. The agent handles chunking automatically for common formats like PDF.
How does this relate to local-first AI?
The wiki model is inherently local-first. All data lives in the workspace directory. The agent processes everything locally (or via API calls for model inference). If you switch tools, your knowledge base is a directory of Markdown files you can take with you.
Can the wiki export to other formats?
Since wiki articles are Markdown files, they can be converted to any format that supports Markdown (HTML, PDF, DOCX) using standard tools. The agent can perform these conversions if requested.
What happens if the agent makes a bad organizing decision?
You can manually edit or reorganize the files. The agent will respect your changes on the next cycle. The log.md records the agent's decisions, so you can see what was changed and when.
