Short answer: A local-first AI assistant keeps your working data—conversations, files, knowledge, memory, and run history—on your own machine by default. It sends data elsewhere only when you explicitly use a feature that requires sign-in, a cloud model call, or a connector. This guide explains what to verify, where Semibot meets the bar, and where the claim has limits.
Why this matters
Cloud-first assistants are convenient, but they shape a trade-off: your working context lives on someone else’s infrastructure. For many everyday tasks that is fine. For work that involves internal documents, client code, or sensitive schedules, some teams prefer to keep that context local unless there is a specific reason to send it out. “Local-first” is a label for that preference—but not every product that uses the label implements the same thing.
A practical checklist
Use this checklist when a tool calls itself local-first. Vague marketing claims are not enough—look for the concrete implementation.
| Claim to verify | What to look for | Why it matters |
|---|---|---|
| Conversations | Stored in a local database you can find on disk | Without this, your history is a cloud object |
| Secrets | Stored via the OS keychain / secret store, not in plaintext config or JSON | Plaintext secrets are a real leak surface |
| Knowledge / memory | Indexed locally; not silently uploaded for “improvement” | Working knowledge often contains confidential details |
| Run history | Commands, tool calls, and results persisted locally | Audit trail should be yours |
| When data leaves | Only for sign-in, model API calls, and connectors you enable | Defines the real boundary of the claim |
| File access model | Folder grants via system file picker; path escapes rejected | Prevents the tool from reading more than intended |
| High-risk actions | Explicit approval; optional per-session automation | Confirms who is in control |
How Semibot implements each item
Semibot’s local-first design means most data lives on the device and data only leaves the machine in specific, visible scenarios. Here is how each item works:
- Local SQLite database. Conversations, knowledge, memory, and run records are written to a SQLite database on your machine. The data is accessible through standard tooling and the application itself.
- Keychain secrets. Credentials are stored via the system keychain and referenced by reference—not dumped into a JSON config file you might accidentally commit or share.
- Folder grants. You authorize directories with the system file picker. The agent reads and writes only within those grants. Escapes via
.., symlinks, junctions, or UNC paths are rejected. - Approval gates. Sending messages, changing external data, writing files, and git write commands go through a unified approval path. You can switch a session to “auto-execute,” which is marked as a dangerous mode—useful for advanced workflows but not the default.
- Git is separate from folder access. Approving a git commit never implicitly grants the agent access to folders it does not already hold.
When data does leave your machine
Local-first is not the same as offline or air-gapped. Real scenarios where data leaves:
- Sign-in. If you use an account (and the built-in trial quota), authentication and quota tracking involve your account.
- Cloud model calls. Prompts and context are sent to a model provider—yours or the built-in one. How much context is sent depends on the task and configuration.
- Connectors. Feishu/Lark, DingTalk, Discord, Telegram, Slack, Gmail, Calendar, local folder watchers, BlueBubbles/iMessage, and MCP tools all involve external communication by definition.
The honest boundary is: the tool is local-first for your working context, not for everything that happens on your network. That is the trade-off worth understanding clearly.
Workflow scenarios: local-first in practice
The local-first label means little without concrete examples of what it changes in daily work. Here are scenarios where the data posture matters.
Scenario: Working with client code under NDA
You are a contractor working on proprietary code under a non-disclosure agreement. The client wants assurance that their code is not being uploaded to third-party servers for model training.
- You open the project in Semibot and grant the project folder via the system file picker.
- The code graph indexes locally—no code is uploaded for this step.
- When you ask the agent to analyze or edit code, prompts and context are sent to the model provider. You can use a provider whose data retention policy the client has approved, or bring your own key to a provider with zero-retention agreements.
- Conversations, ChangeSets, and checkpoints stay in the local SQLite database. If the client asks "where is my code?", you can point to the database file on your machine.
- After the project ends, you can delete the local database. There is no cloud copy to request deletion of (beyond the model provider's prompt handling, which you should verify separately).
Scenario: Auditing past AI interactions
Your team wants to review what the AI agent did over the past month—what files it touched, what commands it ran, what it sent externally.
- Open the run history in Semibot. Every task records commands, tool calls, and results in the local database.
- For coding tasks, the ChangeSet shows every file touched with its diff and the commands that were run, with exit codes.
- For connector interactions, the run log records which connector was used and what data was sent.
- The audit trail is local—your team controls it. It does not depend on a cloud provider's log retention policy.
Scenario: Switching tools or migrating data
You decide to try a different AI assistant and want to take your context with you.
- The local SQLite database contains your conversations, knowledge, and memory. You can export it using standard SQLite tooling.
- The knowledge library files are on your machine—no proprietary cloud format to escape from.
- Conversations can be exported as text. Knowledge articles can be saved as files. The raw data is yours.
- This is a real advantage of local-first: no vendor holds your history hostage. The trade-off is that you manage backups yourself.
Scenario: Air-gapped or restricted network environment
You work in an environment with limited or monitored internet access—common in government, finance, or defense contexts.
- The local data layer (SQLite database, OS keychain) works without network.
- Model inference still requires a network for cloud models. You would need to configure a local model provider accessible within the restricted network.
- Connectors to external services (Feishu/Lark, Slack, Gmail) require network access to those services. In a restricted environment, you would disable those connectors.
- Semibot is not designed as an air-gapped solution. It is local-first for data storage. Full offline operation requires additional infrastructure decisions.
Decision tree: local-first vs cloud-first
The choice between local-first and cloud-first depends on your work, your constraints, and your risk tolerance. Use this to guide the decision:
- If you work with sensitive client data under NDA or compliance requirements,local-first reduces the surface area. Your working context stays on your machine, and you control exactly what leaves. You still need to evaluate the model provider's data handling.
- If you need real-time collaboration with multiple users editing simultaneously,cloud-first is the only practical option. Local-first tools are single-user by nature.
- If you want zero setup and instant access from any device, cloud-first tools win. They run in browsers and sync across devices automatically.
- If you want to control backups, exports, and data retention, local-first gives you that by default. The data is a file on your machine.
- If you work offline frequently, neither approach fully solves the problem for AI tools—model inference typically needs network. Local-first helps with data access but not with model availability.
Comparison: local-first implementations across tools
| Aspect | ChatGPT Desktop | Claude Desktop | Cursor | Semibot |
|---|---|---|---|---|
| Conversation storage | Cloud | Cloud | Mixed | Local SQLite |
| Secret storage | Cloud account | Cloud account | Local config | OS keychain |
| Knowledge / memory | Cloud | Cloud | Project-local | Local SQLite |
| Run history / audit | Cloud | Cloud | Local | Local SQLite |
| File access control | N/A | Connectors | Project model | Folder grants + path escape rejection |
| Data export | Account export | Account export | Git repos | Direct SQLite access |
This table reflects general data posture. Each tool's specifics may vary by version and configuration. Verify current behavior with each vendor.
What local-first does and does not guarantee
Local-first means your data starts on your machine. It does not automatically mean:
- Offline operation. Model inference still needs a network for cloud models. Some tasks could use a local model, but that is a separate decision.
- Perfect privacy. A tool can store data locally and still leak through connectors, logging, or careless sharing. Privacy is a system property, not just a storage decision.
- No data ever leaving. By design, some data leaves when you use features that need it. The claim is about the default and the visibility—not about zero network traffic.
How to audit the claim yourself
- Find the database. A local-first app should have a database file (SQLite, LevelDB, etc.) you can locate on disk.
- Check the config for secrets. Search settings files for tokens, API keys, or passwords in plain text.
- Watch the network. Use a proxy or system monitor and see when the tool makes outbound calls and to where.
- Read the folder-grant model. Is it a real permission system, or just a warning you can dismiss?
- Try a high-risk action. Write to a file or send a message and see if the tool asks before acting.
A local-first claim you cannot verify is a marketing claim. A local-first claim you can audit is a design decision. The audit does not need to be complicated—ten minutes with a file browser, a network monitor, and a test prompt will tell you more than any product page.
Common misconceptions about local-first
- "Local-first means no data ever leaves." False. By design, model inference sends prompts to a cloud provider. Connectors communicate externally. Local-first means your working context (conversations, knowledge, run history) starts on your machine—not that the tool operates in a vacuum.
- "Local-first is always more secure." Not automatically. A local database on an unencrypted disk is less secure than an encrypted cloud service with proper access controls. Security depends on the full system: encryption at rest, access controls, network behavior, and what the tool does with connectors.
- "Local-first means slower." For data access, local is faster—reading from a local SQLite database beats a cloud API call. For model inference, the bottleneck is the model provider's API, and that is the same regardless of where data is stored.
- "Local-first tools cannot collaborate." They cannot do real-time multi-user editing, which is a genuine limitation. But connectors to shared platforms (messaging, email, calendar) allow sharing results and coordinating work through existing communication channels.
- "If it is local-first, I do not need to worry about privacy."Privacy is a system property. Local storage is one component, but you also need to consider what connectors send, what model providers receive, what telemetry exists, and what diagnostic data leaves the machine.
Semibot’s honest limitations
Semibot stores core data locally and uses the OS keychain for secrets. But there are real caveats: cloud model calls involve sending prompts to a provider; connectors involve external services; the product is young and has fewer third-party security reviews than established tools; and Windows builds are currently unsigned. Local-first is a design posture, not a guarantee that nothing leaves your machine.
FAQ
Does local-first mean private?
It is a good start. Privacy depends on the full system: storage, connectors, model calls, and what the tool does with the data. Local storage alone is necessary but not sufficient.
Can I allow only one folder?
Yes. Folder grants are per-directory via the system file picker. High-risk capabilities require approval by default.
Does it work fully offline?
Not for cloud model calls. Local-first is about where data is stored, not about whether the tool can operate without network.
What happens to diagnostics?
Diagnostic logs and diagnostic bundles are sanitized. The scrubbing is a privacy control worth checking on any tool.
Can an external model provider see my data?
Model providers receive the prompts and context you send them. The amount depends on the task. Read the provider's terms and data retention policies.
Is local-first better than cloud-first?
Not universally. It is a trade-off. Local-first gives you control and reduces surface area; cloud-first gives you easier collaboration and often a more polished integration. Pick the trade-off that matches your work.
How large can the local database grow?
SQLite databases can grow to terabytes in theory. In practice, Semibot's database grows with your conversations, knowledge, and run history. Heavy use over months will accumulate data. The knowledge library and run history are the primary contributors to database size. You can manage this by archiving old sessions or pruning knowledge entries.
Can I back up the local database?
Yes. The SQLite database is a regular file on your disk. You can include it in your normal backup workflow—Time Machine, rsync, or any file-level backup tool. There is no proprietary backup format. Just ensure the application is not actively writing when you copy the file.
Does local-first affect performance?
Local storage is fast—reading from a local SQLite database is faster than a cloud API call. However, model inference still depends on your network connection to the model provider. The local-first benefit is in data access speed and availability, not in model speed.
What if I switch from macOS to Windows or vice versa?
Semibot runs on both macOS (Apple Silicon) and Windows x64. The local SQLite database format is the same across platforms. You can copy the database file between machines. The OS keychain entries (secrets) are platform-specific and would need to be re-entered on the new machine.
Are connectors a privacy risk?
Connectors are by definition external communication channels. They send data to external platforms (Feishu/Lark, Slack, Gmail, etc.) as part of their function. Each connector is individually toggleable—you enable only the ones you need. MCP tool connectors do not inherit workspace data or system credentials by default. The risk is proportional to what you connect and what you instruct the agent to send through it.
Does the model provider retain my prompts?
That depends on the provider. Different providers have different data retention and training policies. When using the built-in trial quota, prompts are sent to a hosted model—review the terms of service. When using your own API key, you are subject to that provider's policies. Semibot does not add an additional retention layer on top of the provider.
Is there telemetry or analytics?
Diagnostic logs are sanitized. You should verify the specific telemetry behavior in the application's settings and privacy documentation. Any tool that claims local-first should be transparent about what diagnostic data it collects and where it sends it.
