SemibotSemibot - AI Desktop
All guides

The native code graph: how an agent actually understands your project

Keyword search finds files; a code graph finds relationships. How Semibot's built-in TypeScript-first symbol graph changes how an agent searches, edits, and reviews code.

The short version: when AI coding goes wrong, it is usually not because the model wrote bad code—it is because it changed the wrong place. The root cause is often the same: the agent's understanding of the project comes from scattered text search. Humans navigate unfamiliar codebases with jump-to-definition and find-references, which is really a mental graph of symbol relationships. Semibot builds that graph as a first-class feature—the native code graph—so the agent navigates along real edges: see the impact before you edit, review along the relationships after.

What it is / what it is not

What it is

  • A symbol-level index of the codebase: call relationships, import/export chains, type references, class and interface hierarchies. Not text matching—syntax-tree understanding.
  • A navigation system for the agent: the agent can query the graph continuously—find definitions, find callers, find affected files—like an engineer hopping through an IDE.
  • Built and stored locally: the index is constructed on your machine when you open a project. Building the graph does not send your code anywhere.

What it is not

  • Not a pretty visualization for humans: the graph does not draw decorative diagrams. It serves the agent's retrieval and workflow.
  • Not runtime analysis: it does not know what the program will do when executed. Dynamic behavior is what tests are for.
  • Not equally deep in every language: TypeScript/JavaScript gets the deepest data; other languages are shallower (see boundaries below).

The idea: why an agent needs a graph

Look at where AI coding tools fail and a pattern emerges: rarely "the model wrote wrong code," usually "the model changed the wrong place." And the root cause is usually the same—its understanding of the project came from fragmented text retrieval.

Grep and guess is a gamble

The traditional loop is keyword search: search a function name, an error string, a comment word. Three structural problems. First, naming drifts from meaning—a function called handleRequest may not handle the request you care about. Second, noise crowds out context—ten files come back, nine are irrelevant, and the relevant one may not come back at all. Third, indirect relationships are invisible—change a field on a type and no keyword finds every consumer.

How humans actually read projects

Engineers rarely read a codebase front to back. They start at an entry point and jump: where is this defined, who calls it, which interface does it implement, where does this type come from. Underneath those moves sits a symbol relationship graph—engineers just maintain it in their head with an IDE.

Semibot's philosophy is to hand that navigation ability to the agent completely. Build the graph first, then work. Every retrieval can follow a definite relationship edge instead of gambling on keywords.

What “native” means, twice

  • Built in, not bolted on: the graph is implemented natively in Semibot (written in TypeScript, no Python sidecar, no language server to configure). Open a project and it works.
  • A first-class primitive, not a demo: the graph is not a visualization for sales decks. Querying it is a basic agent action—same rank as reading files and running commands—and its results flow directly into edit plans and ChangeSet review.

Four scenarios where the graph changes agent behavior

1. Taking over an unfamiliar project

Ask the agent to "fix this error." Without a graph, it searches the error text, wanders directories, and guesses the entry point. With a graph, it starts at the error site: jump to the definition → walk up the call chain → confirm the branch that actually throws → locate the root cause. Fewer files touched, a straighter path, measurably less time and cost.

2. Seeing the blast radius before editing

"Make this field optional." The graph answers first: who references the field, where are the non-null checks, which type definitions change in cascade. The agent can list the affected surface before touching anything and produce an edit that covers every consumer—instead of fixing the main call site and waiting for tests to reveal the misses.

3. Reviewing changes along relationships

A ChangeSet shows the diff; the graph answers whether the diff is complete. While reviewing its own or a colleague's change, the agent walks the edges of every modified symbol: did all callers adapt? Are there parallel implementations that need the same change? Review shifts from "does the diff look right" to "verify along the impact surface."

4. Spending context where it counts

Model context is scarce. Graph retrieval hands the model "the few files that are truly relevant" instead of "a pile of loosely related ones." Cleaner context for the same task means less interference, lower cost, and more stable output quality.

Boundaries and limitations

As with every guide on this site, straight talk: the graph has clear edges to its usefulness, and knowing them is part of using it well.

  • Language depth is uneven: the graph is built for TypeScript/JavaScript and gives those projects the deepest data. For other languages the agent still reads, edits, and runs commands, but impact analysis loses precision.
  • The static-analysis ceiling: dynamic dispatch, reflection, references built from strings, runtime-generated code—the graph cannot see these. Tests are the safety net.
  • Very large repos involve trade-offs: monorepos balance indexing time against relationship coverage, and a cold index needs some patience.
  • The graph does not replace tests: the graph tells you where impact is possible; tests tell you whether it is real. Complementary, not interchangeable.

Who benefits most / least

Where the graph shines

  • Medium and large TypeScript/JavaScript codebases: the denser the relationships, the bigger the gap over grep.
  • Teams that constantly inherit other people's code: relationship navigation shortens the "time to understand."
  • Codebases with a high bar for changes: check the blast radius before editing; verify along it while reviewing.

Where it helps least

  • Scripts and tiny tools: for a few dozen files, plain search is enough.
  • Predominantly dynamic-language projects: with shallower relationship data, treat the graph as a hint machine, not an impact oracle.
  • Runtime-verification-heavy work: performance tuning and concurrency issues live in runtime data, not static graphs.

FAQ

What is the native code graph?

A symbol-level index of the codebase: call relationships, import/export chains, type references, class and interface hierarchies. Semibot builds it locally; the agent navigates along relationships instead of guessing with keywords.

How is it different from IDE “find references”?

Similar underneath, different consumer: IDE search shows a human one list; the graph is an agent primitive for continuous navigation, feeding edit plans and ChangeSet review.

Does it need extra installation or configuration?

No. It is built into Semibot—native TypeScript implementation, no Python sidecar, no language server configuration. Open a project and use it.

Where is the graph data stored?

Locally on your machine, like all of Semibot's work data. Building the graph does not send your code anywhere.

Do I still need tests with a graph?

Yes. The graph is static analysis—it cannot see dynamic dispatch, reflection, or string-built references. It finds where impact is possible; tests verify whether it is real.

How is this different from stuffing the repo into the model?

Repos exceed any context window. The graph wins on context economics: relationship retrieval surfaces the few relevant files first, so the model sees less noise, costs less, and outputs more stable results.

Further reading