SemibotSemibot - AI Desktop
All guides

Native Code Graph: A TypeScript Code Knowledge Graph Engine

The code graph is a navigation aid: it shows what is connected to what, so the agent can follow relationships instead of grepping through every file. It is not a correctness oracle.

Definition: Semibot's native code graph is a TypeScript code-knowledge-graph engine that indexes calls, references, type relationships, and structural connections across a codebase. It is built into the core runtime—no Python sidecar, no external MCP dependency. The agent uses it to navigate unfamiliar projects faster, read fewer irrelevant files, and follow the blast radius of a change during code review. The graph is a navigational aid, not an oracle for correctness.

Why keyword search is not enough

Keyword search (grep, find, semantic search) answers "which files contain this term?" The code graph answers "what is connected to this function, and how?" This distinction matters for three tasks: onboarding to an unfamiliar project (understanding structure, not finding files), planning a change (knowing what else is affected), and reviewing code (tracing the impact of modifications).

Schema v4 and the four-layer tool contract

The code graph uses schema v4: unified structure nodes, capability-level readiness, and derived indexes for community detection, data flow, topology, and code quality. The engine exposes four layers of tools (G1–G5) to the agent:

  • G1: Basic queries—find definitions, references, and call sites.
  • G2: Structural queries—module boundaries, dependency directions, layer violations.
  • G3: Impact analysis—given a change at node X, which other nodes are affected?
  • G4/G5: Quality and topology metrics—complexity, coupling, community structure.

The agent does not choose which layer to use—the Tool Gateway routes based on the task and the agent's request. This prevents the agent from accidentally running expensive graph-wide analyses when a simple reference lookup would suffice.

Navigation vs correctness

The graph shows relationships, not whether code is correct. It can tell you that function A calls function B, and that B has three callers. It cannot tell you that B has a logic error. High-risk conclusions—bugs, security issues, performance problems—must go back to source code or reproducible commands. The graph accelerates navigation; it does not replace reading.

This boundary is enforced by design. The code graph is exposed as navigational tools in the Tool Gateway. It does not produce "analysis reports" or "quality scores" that could be mistaken for verdicts. It produces data that the agent interprets.

How the agent uses it in practice

When reviewing a code change, the agent can ask the graph: "what else calls this function?" and "what modules depend on this file?" This produces a blast radius—the set of code that might be affected by the change. The agent then reads the relevant files and reports the impact in the ChangeSet. Without the graph, the agent would need to grep and guess; with the graph, it follows explicit relationships.

Scenario: onboarding to an unfamiliar codebase

A developer joins a team and faces a 200-file TypeScript project. Without the code graph, the agent would need to read files sequentially or grep for keywords, producing a fragmented picture. With the code graph:

  1. Entry point identification. The agent queries G2 for module boundaries and identifies the top-level entry points: the main server file, the router configuration, and the data layer initialization.
  2. Dependency mapping. For each entry point, the agent queries G1 for call sites and references, building a tree of dependencies. The graph reveals that the router calls 12 handler modules, each of which calls a shared data access layer.
  3. Boundary discovery. G2 identifies layer violations—places where the data layer directly imports from the handler layer instead of going through the service layer. These architectural issues are invisible to keyword search.
  4. Community detection. G4/G5 reveals three distinct clusters of tightly coupled code: the auth system, the data pipeline, and the API layer. The agent reports this as the project's natural decomposition.

The result is a structural understanding of the project that would take a human developer days to build through manual exploration. The agent achieves it in minutes by following graph relationships instead of guessing from file names.

Scenario: blast radius for a code review

A pull request modifies the authentication middleware. The agent needs to understand what else might be affected:

  1. Direct callers. G1 identifies all functions that call the modified authentication function. There are 8 direct callers across 4 files.
  2. Transitive dependencies. G3 computes the blast radius: the 8 callers are themselves called by 23 other functions, spanning the API layer, the admin panel, and the webhook handler.
  3. Type propagation. The graph shows that the authentication function returns a type used by 15 other functions. Changing the return type signature would cascade further.
  4. Impact report. The agent reads the relevant files in the blast radius and reports in the ChangeSet: "This change to auth middleware directly affects 8 callers and transitively affects 23 functions. The webhook handler (src/webhooks/auth.ts:45) has a different error handling pattern that may need updating."

Without the graph, the agent would grep for the function name and find direct callers, but miss transitive dependencies and type propagation. The graph provides structural certainty where grep provides textual coincidence.

Scenario: refactoring with confidence

A team wants to extract the notification system into a separate module. Before making changes, the agent uses the graph to plan the refactoring:

  1. Boundary analysis. G2 identifies all files that belong to the notification system based on call patterns and import relationships. The graph finds 12 files with 47 internal connections and 8 external dependencies.
  2. External interface mapping. G1 identifies the 8 functions that are called from outside the notification system. These become the public API of the new module.
  3. Coupling metrics. G4 reports that the notification system has a coupling score of 0.3 (moderate). Two of the 8 external dependencies are to the auth system, which the graph shows is a shared dependency across all modules.
  4. Refactoring plan. The agent recommends extracting the 12 files, creating a clean public interface with the 8 external functions, and noting the 2 auth dependencies that will need to be injected rather than imported.

The graph does not guarantee the refactoring is safe—it cannot verify runtime behavior. But it provides the structural information needed to plan the refactoring with confidence, reducing the risk of missing an important connection.

When to use the code graph vs. alternatives

The code graph is one tool among several for codebase navigation. Use it when structural understanding matters; use alternatives when it does not:

TaskBest ToolWhy
Find a specific string or variableGrep / semantic searchTextual match is faster and sufficient
Understand call relationshipsCode graph (G1)Structural relationships, not text matches
Find architectural boundariesCode graph (G2)Module detection requires dependency analysis
Assess change impactCode graph (G3)Blast radius requires transitive closure
Read implementation detailsFile readThe graph shows structure, not semantics
Verify correctnessTests / runtimeThe graph cannot verify behavior
Measure code qualityCode graph (G4/G5) + lintersGraph provides coupling metrics; linters check style

Comparison: native graph vs. external code analysis tools

DimensionSemibot Native GraphExternal MCP ToolLanguage Server (LSP)
DeploymentBuilt-in, zero setupRequires separate server processRequires language server installation
LatencyIn-process, sub-secondNetwork round-trip + server processingIPC, typically fast
Schemav4 (unified structure nodes)Tool-specific schemaLSP protocol (limited to editor features)
Community detectionBuilt-in (G4/G5)Varies by toolNot available
Impact analysisBuilt-in (G3)Varies by toolFind references only
Language supportTypeScript/JavaScriptDepends on toolPer-language server
Rate limitingBuilt-in (G4/G5 are expensive)User-managedNot typically limited

The native graph's advantage is zero-deployment, in-process access with a structured tool contract. The disadvantage is limited language support. For non-TypeScript projects, the agent falls back to general search and the user's guidance, which is less precise but still functional.

Theoretical depth: why graph structure helps code understanding

Code is a directed graph. Functions call other functions. Modules import other modules. Types reference other types. This graph structure encodes the architectural decisions of the codebase—who depends on whom, which layers are allowed to communicate, where the boundaries are.

Keyword search operates on the textual surface of code. It can find "function authenticate" but cannot tell you that authenticate is called by 47 other functions, that it belongs to the auth module, or that changing its return type would break the API layer. These are graph properties, not text properties.

The code graph makes these properties queryable. Community detection (G4/G5) finds clusters of tightly coupled code that correspond to natural modules—even when the directory structure does not reflect them. Coupling metrics quantify how much one part of the codebase depends on another, providing objective data for refactoring decisions. Topology analysis identifies architectural violations (circular dependencies, layer crossings) that are invisible to text search.

The key insight is that the graph is a derived projection: it can be rebuilt from the source code at any time. It does not store information that is not already in the code—it reorganizes existing information into a structure that is more useful for navigation and analysis. This means the graph can be stale (if the code has changed since the last index) but never wrong in a way that introduces information not present in the source.

Schema v4 in detail

Schema v4 is the fourth iteration of the code graph schema. Each version addressed limitations of the previous one:

  • v1: Basic call graph. Functions and call sites. No module structure or type relationships.
  • v2: Added module boundaries and import relationships. Enabled G2 queries (layer violations, dependency directions).
  • v3: Added type relationships and data flow tracking. Enabled richer G3 queries (type propagation, data dependency chains).
  • v4: Unified structure nodes. A single node type represents functions, classes, modules, and types with consistent relationship edges. Added derived indexes for community detection, topology, and code quality. Enabled G4/G5 queries.

The unified structure node in v4 is the key architectural change. Instead of separate node types for functions, classes, and modules (which required type-specific query logic), v4 uses a single node type with properties. This simplifies the tool contract: the agent does not need to know whether it is querying a function or a module—it asks for "nodes connected to X" and gets all relationship types.

Capability-level readiness means the engine reports which G-tiers are available for a given project. A small project with fewer than 50 files may not have meaningful community structure, so G4/G5 report low confidence. A large project with thousands of files produces high-confidence community detection and coupling metrics.

Boundary conditions: when the graph misleads

The code graph is a navigation aid, but it can mislead in specific scenarios:

  • Dynamic dispatch. In languages with runtime method resolution (Python, Ruby), the graph cannot statically determine which function is actually called. TypeScript's type system helps, but dynamic patterns like eval() or computed property access are partially captured or missed entirely.
  • Stale index. If the codebase changes significantly after the last index, the graph may show relationships that no longer exist or miss new ones. The agent can request a re-index, but the cost of re-indexing a large project may be non-trivial.
  • False precision. Graph metrics like coupling scores are structural measures. A low coupling score does not mean the code is good—it means the module has few dependencies. Quality depends on what those dependencies are and how they are used.
  • Over-reliance on structure. The graph shows what is connected to what, but not why. Two functions may be connected because of a deliberate architectural choice or because of a shortcut taken during development. The graph cannot distinguish between the two.
  • Generated code. Auto-generated files (protobuf outputs, code generators) may inflate the graph with connections that are artifacts of the generation process, not real architectural relationships.

Index management and freshness

The code graph index is stored locally in the same SQLite database that holds conversations and knowledge. The index is a derived projection: it can be rebuilt from the current source files at any time. This means the index can be stale (if the code has changed since the last index) but never contains information that is not present in the source.

The agent monitors for significant codebase changes (new files, renamed functions, restructured modules) and can request a re-index when it detects staleness. For small projects (under 100 files), re-indexing takes seconds. For large projects (thousands of files), re-indexing may take minutes and is performed in the background to avoid blocking the agent.

Incremental indexing is used when possible: only changed files are re-analyzed, and their relationships are updated in the graph. This makes frequent re-indexing practical even for large projects. Full re-indexing (rebuilding the entire graph from scratch) is available as a fallback when incremental indexing would produce inconsistent results.

Limitations

  • The graph indexes static relationships. Dynamic dispatch, reflection, and metaprogramming are partially captured or not captured.
  • Index freshness matters. If the codebase changes significantly after the last index, the graph may be stale. The agent can request a re-index.
  • Language support is currently TypeScript/JavaScript. Other languages use the agent's general search capabilities instead.
  • The graph is for navigation, not verification. It cannot prove code is correct or safe.
  • Graph-wide analyses (G4/G5) are expensive and rate-limited to prevent resource exhaustion.

Benchmarks, and an honest record of failures

Semibot's code graph is a native TypeScript implementation: no Python dependency, no language server, no sidecar process. Syntax comes from Tree-sitter's WASM grammars; module resolution reuses the TypeScript compiler API. The graph model defines 15 node kinds and 10 edge kinds (contains, imports, exports, calls, references, extends, implements, tests, depends_on, resolves_to), and every edge carries a resolution level (syntax-exact / resolver-exact / heuristic / unresolved), a confidence value, and an evidence fingerprint—confidence is produced by resolver rules and is never filled in by a model. Symbol keys are normalized so that moving lines around does not break references.

Queries are organized into four layers totalling 26 tools: symbol search; relationship traversal (find callers, callees, impact); repository-wide quality and topology (community detection via Louvain, cycle detection via Tarjan—both deterministic); and an explicitly change-aware assist layer. The first three layers cannot read the ChangeSet or Git state and reject model-supplied SQL or arbitrary traversal DSLs—query capability is a product definition, not an improvisation. The design document keeps its benchmarks: a full 100k-file index completes in a median of about 10.4 seconds against roughly 43.4 seconds for the third-party Python reference; the index is about 49% of the competitor's size; symbol and impact queries sit at a p95 around one millisecond.

Equally worth citing is the design doc's itemized record of failures: warm-query latency and memory footprint both failed internal gates (including observed I/O stalls under resource pressure), and the release gate accordingly hard-codes warm-query p95 at least 20% faster than the baseline and peak incremental memory within 80% of the third-party implementation. Correctness gates are quantified too: exact symbol retrieval requires top-1 precision ≥99%, depth-2 impact recall ≥98%. Writing the failures into the design document and converting them into acceptance criteria is the most “research” thing about this engine.

FAQ

Do I need Git for the code graph to work?

No. The code graph indexes the file system, not the Git history. Git is optional source control.

Does it support Python/Rust/Go?

Currently TypeScript/JavaScript. Other languages fall back to the agent's general search and the user's guidance.

Can I query the graph directly?

The agent queries it through the Tool Gateway. There is no user-facing graph visualization or query interface.

Is the index stored locally?

Yes. The graph index is in the local SQLite database alongside conversations and knowledge.

How often should the graph be re-indexed?

The agent can request a re-index when it detects stale data. For active development, a re-index after significant changes (new files, renamed functions, restructured modules) keeps the graph accurate. The re-index cost scales with project size.

Can the graph detect circular dependencies?

Yes. G2 identifies dependency cycles at the module level. The agent can report these as architectural issues during code review.

What is the difference between G1 and G3?

G1 answers "what is directly connected to X?" (callers, callees, references). G3 answers "what is transitively affected if X changes?" (blast radius, impact propagation). G3 requires computing the transitive closure, which is more expensive.

Why are G4/G5 rate-limited?

Graph-wide analyses like community detection and topology metrics require traversing the entire graph. For large codebases (thousands of files), these operations are computationally expensive. Rate limiting prevents resource exhaustion and keeps the agent responsive for common queries.

Does the graph replace the need to read code?

No. The graph accelerates navigation by showing what is connected to what. The agent still needs to read the relevant files to understand implementation details, logic, and correctness. The graph tells you where to look; reading tells you what is there.

How does the graph handle monorepos?

The graph indexes the entire workspace. In a monorepo with multiple packages, it captures inter-package dependencies as well as intra-package relationships. Module boundaries are detected based on actual import patterns, not just directory structure.

Can the graph detect dead code?

Partially. G1 can identify functions that have no callers within the indexed codebase. However, this does not account for dynamic calls, external consumers, or test-only usage. The agent treats "no callers" as a signal, not a verdict.

How does this relate to the coding agent checklist?

The coding agent checklist describes what a coding agent should do before, during, and after code changes. The code graph supports the "during review" phase by providing blast radius analysis and structural understanding. It is one of several tools the agent uses to meet the checklist requirements.

Related