Short answer: Do not trust a coding agent on demo videos alone. Verify seven things: it works in a real project, respects folder grants, shows real diffs, supports safe undo, runs commands and tests under evidence, asks before high-risk actions, and does not force Git as a precondition. This guide explains each check, how Semibot handles it, and where Semibot is still limited.
Why a checklist matters
AI coding agents are now fast enough to impress in demos. But a demo on a clean repository and work on a multi-folder project with history are different problems. The risk is not that the agent will try something—it is that you will not see what it did, or you will not be able to undo it cleanly. This checklist is a practical framework for evaluating a tool before you let it touch real work.
The checklist applies to any coding agent, not just Semibot. If you are evaluating Cursor, GitHub Copilot Workspace, or any other AI coding tool, run through the same seven checks. Where a tool falls short, you compensate with your own workflow—extra git commits, manual review, or limiting what you allow the agent to do.
The seven checks
- Real-project access. Can the agent scan and work in your actual repository—or does it need you to copy files into a sandbox? A serious tool works where the work lives. Semibot: scans real projects in the background without freezing the UI; large codebases are handled in background processes.
- Multi-folder grants. Real work often spans more than one directory—a repo, a shared config folder, a documentation tree. A one-task, one-root model is limiting. Semibot: supports one work directory plus several authorized read/write directories in the same task.
- Visible changes (ChangeSet). Before you accept anything, you should see: files touched, real diffs, commands executed, test results. “I changed some files” is not enough. Semibot: delivers a ChangeSet with that evidence for every coding task.
- Checkpoints / safe undo. Agent-controlled edits should be reversible without forcing you into a full git workflow. Later edits by you or your IDE must not be blindly overwritten. Semibot: keeps file checkpoints for agent-edited files with complete evidence; later human/IDE edits are not blindly overwritten.
- Command & test evidence. “It works” is not evidence. Exit codes and test output are. Semibot: lists commands executed and test results for every task.
- High-risk approvals. Send messages, write files, change external data, and git write commands should pass a human gate—with optional per-session automation that stays visibly dangerous. Semibot: uses a unified approval path; auto-execute mode is marked as dangerous and still cannot bypass folder grants.
- Git is optional, not required. Not every project uses Git, and not every file change needs a git commit. Approving a git command must not imply folder access. Semibot: Git is optional source control; approving a git commit never implicitly grants folder access.
Side-by-side
| Check | Typical IDE agent | Pure chat model | Semibot |
|---|---|---|---|
| Multi-folder grants | Often single root | N/A | Yes |
| Safe undo | Git-dependent | N/A | File checkpoints |
| ChangeSet evidence | Diff view | Snippet only | Files + diff + commands + tests |
| Approvals | Varies | N/A | Unified, incl. git writes |
| Requires Git | Often yes | No | No (optional) |
Deep dive: ChangeSet & checkpoints
A ChangeSet is a delivery projection: what changed in every authorized read/write directory. It is the evidence you review before accepting changes.
What Semibot can undo: Files edited through Semibot’s controlled tools with complete checkpoint evidence support safe undo. The checkpoint captures the before-image so the change can be reversed without requiring a full git history.
What Semibot cannot undo: If you or your IDE changed a file after Semibot edited it, Semibot will not blindly overwrite it on undo. The checkpoint only applies to the exact state at edit time. For full command-level rollback, you would need Git worktrees or an isolated execution mode.
Deep dive: the native code graph
Semibot ships a native TypeScript code-knowledge-graph engine—no Python sidecar, no external MCP dependency for this function. It indexes code relationships across a project so the agent reads fewer irrelevant files and can follow the blast radius of a change during review.
What the graph does: navigation and impact analysis. It helps the agent ask “what else is affected by changing this?” without scanning every file in the project.
What the graph is not: an oracle for correctness. The graph shows relationships, not whether code is right. High-risk conclusions should go back to source code or reproducible commands—not be trusted from the graph alone.
Workflow scenarios: applying the checklist
The seven checks are abstract until you see them in a real workflow. Here are three scenarios that exercise different checks.
Scenario: Fixing a bug in a monorepo you did not write
You inherit a monorepo with three packages, a shared config folder, and a documentation tree. A bug report says the authentication module fails silently under certain conditions.
- Check 1 (real-project access): You point Semibot at the monorepo root. The background scanner indexes the code graph—no copying files into a sandbox.
- Check 2 (multi-folder grants): The auth module lives in
packages/auth, but the bug might be in the shared config atconfig/env. You grant both directories as authorized read/write paths in the same task. - Check 3 (ChangeSet): After the agent proposes fixes, you see a ChangeSet: three files touched, each with a diff, two commands run (lint and type-check), and test output with exit codes.
- Check 4 (checkpoints): Before accepting, you notice one edit changes a shared utility. You roll back that single file from the checkpoint and accept the other two.
- Check 6 (approvals): The agent wanted to run a database migration command. The approval gate caught it and asked for confirmation first.
Scenario: Adding a feature to a project without Git
You are working on a project that does not use Git—maybe it is a prototype or a script collection. You want to add a new utility function.
- Check 7 (Git optional): Semibot does not require Git. File checkpoints work independently as a rollback mechanism.
- You describe the feature. The agent proposes which files to create and edit.
- After approval, the ChangeSet shows the new file with its content, the edited files with diffs, and the commands that were run.
- You accept all changes. The checkpoint records the before-image of each edited file so you can undo later if needed.
- No git commit was made, no git history was required, and the approval gate for file writes still applied.
Scenario: Code review with blast-radius analysis
A teammate submits a change to the payment processing module. You want to understand what else is affected before approving.
- Code graph: Semibot's native TypeScript code graph traces the call chain from the changed functions. It identifies three other modules that import from the changed file.
- You ask the agent: "What else is affected by this change?" The graph narrows the analysis to those three modules instead of scanning every file in the project.
- The agent examines the affected modules and reports: one has a test that exercises the changed path, two do not. It suggests adding test coverage before merging.
- Important caveat: The graph shows relationships, not correctness. The agent's analysis is a starting point, not a verdict. You still review the source code for the final decision.
Decision tree: which safety model fits your project?
Different projects need different levels of safety. Use this to decide which checks matter most for your situation:
- Solo prototype, disposable code: You may not need checkpoints or ChangeSet review. A chat model that generates code snippets is sufficient. Speed matters more than safety here.
- Solo production project: You need at minimum: real-project access, ChangeSet evidence, and checkpoints. Approvals for file writes are valuable. Git is helpful but should not be required.
- Team project with CI/CD: All seven checks matter. Multi-folder grants for cross-package work, ChangeSet with test evidence for code review, approvals for anything that touches shared code, and Git as optional (your CI pipeline handles version control).
- Regulated or security-sensitive code: Add audit requirements on top: local-first data storage, visible approval logs, and no data leaving to cloud models without explicit consent. Evaluate whether the tool's data posture meets your compliance needs.
Expanded comparison: coding agent capabilities
| Capability | Typical IDE agent | Pure chat model | Semibot |
|---|---|---|---|
| Background code indexing | Yes (editor-native) | No | Yes (native TypeScript graph) |
| Blast-radius analysis | Limited (editor scope) | No | Yes (code graph relationships) |
| Multi-folder in one task | Usually single root | No file system | Work dir + authorized dirs |
| File-level undo without Git | No (Git-dependent) | No | Yes (file checkpoints) |
| Command + test evidence in output | Terminal panel | None | In ChangeSet (files + diff + cmds + tests) |
| Unified approval for writes | Varies by tool | N/A | Yes (incl. git writes, messages) |
| Office work (docs, research) | Limited | Yes (chat-based) | Built-in workbench |
| Auto-execute mode | Yes (some tools) | Always | Per-session, marked dangerous |
Boundary conditions: when the checklist has limits
- Polyglot projects: The native code graph is TypeScript-based and indexes TypeScript/JavaScript projects natively. For projects in other languages, the graph provides less detailed relationship data. The agent can still read and edit any text file, but impact analysis is less precise.
- Extremely large diffs: If a task touches hundreds of files, the ChangeSet becomes large and harder to review meaningfully. Breaking work into smaller tasks produces more auditable results.
- Commands with side effects: Database migrations, API calls, and external service writes are visible in the ChangeSet, but the agent cannot undo their external effects. The checkpoint only covers file state.
- Concurrent edits: If you and the agent edit the same file simultaneously (e.g., you in your IDE while the agent is working), the checkpoint captures the state at the agent's edit time. Your concurrent changes are preserved and not overwritten—but the agent's checkpoint may not reflect the file's current state.
Limitations to be aware of
- Windows builds are unsigned. The installer may trigger OS warnings. This is a trust barrier worth acknowledging.
- Linux desktop is still on the roadmap. Not shipped yet.
- Cloud model calls need network. Local-first is about data storage, not offline inference.
- Young product, fewer third-party reviews. Less community coverage than established tools. Evaluate it on its own claims, not on hype.
- Code graph is for navigation. Not a correctness verifier. The graph helps the agent understand relationships and blast radius, but high-risk conclusions should be verified against source code and reproducible commands.
- Context limits apply. Very large files or deeply nested codebases may exceed what the agent can hold in a single task context. The code graph reduces what needs to be read, but cannot eliminate context constraints for the largest projects.
- External side effects are not reversible. Checkpoints cover file state only. Commands that call external APIs, modify databases, or trigger CI pipelines have effects that cannot be undone by the checkpoint mechanism.
FAQ
Can it break my project?
Risk is managed by folder grants, checkpoints, ChangeSet review, and approvals—not by hope. But no tool can guarantee zero risk. Review evidence before accepting changes.
Do I need Git?
No. Git is optional source control. Semibot detects Git if present but does not require it. Git write commands still go through approval, and approving git never grants folder access.
What is the code graph for?
Navigation and impact analysis—not as an oracle. Verify high-risk claims in source code or with reproducible commands.
Can I switch to auto-execute?
Yes, per session. It is marked as a dangerous mode. Browser hard gates, system permissions, and folder grants still cannot be bypassed even in auto-execute.
Is it faster than Cursor or Copilot?
We do not publish speed benchmarks. Speed depends on many factors including model, network, and project size. Evaluate it on your own work.
What happens if the agent corrupts a file?
The file checkpoint captures the before-image of every agent-edited file. You can restore individual files from the checkpoint without requiring Git. However, if you or your IDE modified the file after the agent's edit, the checkpoint reflects the state at edit time, not the current state.
Does the code graph work for non-TypeScript projects?
The native code graph engine is built in TypeScript and provides the deepest analysis for TypeScript and JavaScript projects. For other languages, the agent can still read, search, and edit files—it just has less detailed relationship data for impact analysis. The core ChangeSet, checkpoint, and approval features work for any text-based project.
Can the agent run any command I ask?
Within the granted folders and with approval. Commands that require system-level permissions or access outside granted folders will be rejected or require approval. In auto-execute mode, commands run without per-command confirmation, but folder grants and system permissions still apply.
What if my project uses a non-standard build system?
Semibot does not assume a specific build system. It runs whatever commands you or the agent specify. If your project uses Make, Gradle, Bazel, or a custom script, the agent can invoke those commands within the granted folders and approval gates.
How does auto-execute interact with folder grants?
Auto-execute mode skips per-command approval but does not bypass folder grants. The agent still cannot read or write outside authorized directories, and path escapes (via .., symlinks, junctions, UNC paths) are still rejected. System permissions and browser hard gates also remain in effect.
Can I review changes before they hit disk?
In the default approval mode, yes—the agent proposes changes and you approve before files are written. In auto-execute mode, files are written immediately but you can still review the ChangeSet and roll back individual files from the checkpoint.
Does it support test-driven development workflows?
You can instruct the agent to write tests first, then implement the feature. The ChangeSet will show both the test files and the implementation files with their diffs, plus the test execution output. The agent can iterate: write test, run test, fix implementation, run test again—all within one task with full evidence.
