Definition: A dark-factory AI agent operates without human presence: it executes tasks, monitors outcomes, and escalates only when an approval gate is hit or a meaningful change is detected. The name comes from manufacturing: a "lights-out factory" runs with the lights off because no workers are on the floor. The design challenge is not making the agent intelligent enough to run alone—it is making it restrained enough to not create noise while running alone.
The core design problem
Most AI assistants are designed for interaction: they answer when asked, and the conversation is the product. An autonomous agent inverts this: the product is the work done while nobody is watching. The interaction happens only at handoff—when the agent delivers results, asks for a decision, or reports a failure.
This creates a specific design tension. If the agent is too quiet, the user does not know what it did or whether it failed. If it is too noisy, it becomes another notification stream to ignore. The dark-factory principle resolves this by treating every potential signal through a scoring pipeline before it reaches the user.
The attention score model
Every signal—proposed suggestion, change detection, task completion, failure—enters a scoring pipeline before it can become a notification:
score = ruleScore + recencyScore + relationshipScore + urgencyScore + userPreferenceScore - noisePenalty - cooldownPenalty
The score determines the tier:
- Suggest (≥ 0.70): appear in the agent's suggestion area, no push notification.
- Notify (≥ 0.82): push notification, but only during active hours.
- Auto-draft (≥ 0.90, low risk only): prepare a response or action without asking.
External write operations never auto-execute. The auto-draft tier is reserved for low-risk internal actions like organizing a file or preparing a summary.
The attention budget
An attention budget caps how many signals can reach the user per day. This is not a throttle on the agent's work—it is a throttle on the agent's interruptions. The agent can work continuously; it just cannot continuously demand attention.
The budget works with cooldown: after a notification, a cooldown timer prevents the next one for a configurable period. This prevents the common failure mode where an agent discovers something interesting, notifies the user, then immediately discovers something related and notifies again.
The philosophy: prefer missing an alert over creating noise
The V3 attention policy defaults to a conservative stance:
宁可少提醒,不要频繁打扰。
Prefer missing an alert over creating noise.
This is a philosophical choice, not a technical compromise. The reasoning: a user who misses one relevant notification can recover by checking the agent's summary later. A user who receives too many notifications will disable them entirely—and then miss everything. The asymmetric cost favors restraint.
The secretary as dark-factory instance
Semibot's secretary is the primary dark-factory component. It runs scheduled checks, monitors connected services, compiles digests, and delivers results—all without user presence. The secretary's attention model follows the same pipeline: it only interrupts when it needs input, needs approval, or detects a material change that exceeds the attention threshold.
The secretary does not just run tasks on a timer. It watches for changes: "there is new content since the last check" is a different signal from "the scheduled time arrived." Change-aware delivery means the secretary stays quiet when nothing moved—which is exactly the dark-factory principle.
Approval gates in autonomous mode
Autonomous operation does not mean unlimited operation. The approval model defines hard gates that the agent cannot bypass regardless of its confidence:
- Sending messages to external services.
- Writing files outside granted directories.
- Git write operations (commit, push).
- Any capability flagged as high-risk by the capability policy.
Users can switch a session to "auto-execute" mode, which skips confirmation for most actions. But auto-execute is visually marked as dangerous, and it cannot bypass folder grants, browser hard gates, or system permissions. The design ensures that even a fully-autonomous agent operates within boundaries the user explicitly set.
Dismiss feedback and learning
When a user dismisses a suggestion, the dismissal is feedback. The score model adjusts user preference weights so similar signals score lower in the future. This creates a feedback loop: the agent starts conservative, the user's dismissals teach it what "too noisy" looks like, and over time the signal-to-noise ratio improves.
The alternative—asking the user to configure notification preferences manually—is worse. Most users will not configure thresholds; they will either tolerate noise or turn off notifications entirely. Dismiss-as-feedback is a better interface for preference learning.
Workflow examples
The dark-factory principle is best understood through concrete scenarios. Here are three workflows that demonstrate how autonomous operation, attention scoring, and approval gates interact in practice:
Scenario 1: Weekly project digest via secretary. You tell the secretary, "Summarize project progress every Friday afternoon." Step 1: The secretary parses the instruction and proposes a default rhythm—Friday 4 PM. Step 2: On Friday, it checks the project folder, recent git commits, and any connected task boards. Step 3: It compares current state against the last delivered digest. Step 4: If nothing changed, it stays quiet (change-aware delivery). Step 5: If there are changes, it compiles a summary and places it on the secretary page. Step 6: If the summary reveals a blocker that exceeds the attention threshold, it sends a notification. Otherwise, the digest waits silently for you to read it.
Scenario 2: Competitor monitoring with escalation. You delegate: "Watch Competitor X's product page and pricing page. Tell me if they launch a new feature or change pricing." Step 1: The agent schedules periodic checks using the browser skill. Step 2: Each check captures page state and compares against the stored baseline. Step 3: Minor changes (CSS tweaks, blog posts) are logged but not surfaced. Step 4: A pricing change scores high on urgency and rule-based relevance, crossing the notify threshold. Step 5: The agent sends a notification with the specific change and a link. Step 6: If you dismiss several low-relevance detections, the score model adjusts, reducing future noise from similar signals.
Scenario 3: Overnight code review queue. You assign a specialist to review incoming pull requests. Step 1: The specialist monitors the repository for new PRs. Step 2: For each PR, it reads the diff, checks against the project's conventions, and writes a review. Step 3: The review is auto-drafted (score ≥ 0.90, low risk) and placed in the suggestion area. Step 4: External write operations—posting the review to GitHub—require approval. Step 5: In the morning, you see a queue of drafted reviews with approval buttons. Step 6: You approve the ones you agree with, edit and approve others, and dismiss the rest. Each dismissal feeds back into the score model.
Decision framework: when to use autonomous vs interactive mode
| Factor | Interactive mode | Scheduled / secretary | Full autonomous (auto-execute) |
|---|---|---|---|
| Task frequency | One-off or rare | Recurring (daily, weekly) | Continuous or high-frequency |
| Error cost | Low, user catches immediately | Moderate, reviewed at delivery | Low, internal actions only |
| Required presence | User is present | User is absent | User is absent and unavailable |
| External writes | User approves inline | Queued for approval | Never auto-executed |
| Best for | Exploration, creative work | Monitoring, reporting, digests | File organization, internal summaries |
The progression is deliberate: start interactive, delegate recurring tasks to the secretary, and enable auto-execute only for low-risk internal actions where the cost of approval friction exceeds the cost of occasional errors.
Comparison with other autonomous approaches
| Dimension | Scheduled scripts (cron) | Zapier / IFTTT | ChatGPT plugins | Semibot dark-factory |
|---|---|---|---|---|
| Intelligence | Fixed logic | Rule-based with some AI | AI within each call | Full model with workspace context |
| Attention management | None | Basic filtering | None | Score model with budget and cooldown |
| Change detection | Manual diff logic | Trigger-based | No | Built-in comparison against last state |
| Approval gates | None | None | Inconsistent | Integrated, configurable per capability |
| Context awareness | None | Per-zap context | Per-conversation | Full workspace and file system context |
| Learning from feedback | No | No | Session only | Dismiss feedback adjusts future scoring |
Connector and integration examples
The dark-factory principle extends across Semibot's connector ecosystem. Each integration follows the same pattern: monitor, score, escalate only when the threshold is crossed.
- Feishu / DingTalk. The secretary can monitor group channels for mentions, action items, or document updates. When a message requires your input and scores above the notification threshold, it surfaces as a ping. Routine messages are summarized into a digest you can review at your own pace.
- Gmail. The agent can check for new emails matching specific criteria—sender, subject keyword, or attachment type. High-priority emails (detected by rule score and urgency) trigger notifications. Others accumulate in a digest. The agent can auto-draft replies for routine messages, but sending always requires approval.
- Google Calendar. The secretary reads upcoming events and can prepare briefings—meeting agendas, attendee backgrounds, relevant documents from the workspace. These are delivered before the meeting, not as a notification, but as a prepared briefing on the secretary page.
- Slack. Similar to Feishu: the agent monitors designated channels, scores messages for relevance, and escalates only material items. The attention budget prevents channel noise from becoming notification noise.
- Telegram. For personal workflows, the secretary can deliver digests and pings through Telegram. This is useful when you are away from the desktop app but need to stay aware of high-scoring signals.
Failure modes and boundary conditions
Understanding when the dark-factory principle breaks is as important as understanding when it works:
- Score model degradation. If the model provider is unavailable, the agent falls back to rule-based scoring. Rule-based scoring cannot distinguish "CEO sent an urgent email" from "newsletter arrived." The result is either too many or too few notifications.
- Dismiss spiral. If you dismiss aggressively during a busy week, the agent learns that everything is noise. The next week, when your workload normalizes, the agent has been trained to suppress signals it should surface. The recovery path is to explicitly engage with suggestions for a few days to recalibrate.
- Cross-context noise. The attention budget is global. A week of heavy monitoring on one project can exhaust the budget, suppressing signals from other projects. Per-project budgets are not yet supported.
- Stale baselines. Change detection compares against the last delivered state. If the agent has not checked a source for several days, the delta can be large, producing an overwhelming digest rather than a crisp update. The secretary mitigates this with configurable check intervals, but the user must set them appropriately.
- Approval bottleneck. If the agent produces many auto-drafts that require approval, the approval queue becomes the new notification stream. The dark-factory principle works only when the approval rate is low enough that the queue does not accumulate.
Theoretical depth: information economics of AI notifications
The attention score model is an application of information economics to AI-human interaction. Each notification carries a cost (user attention, context-switching) and a benefit (actionable information). The optimal notification policy maximizes the expected benefit minus the cost, subject to the constraint that the user's attention is a finite resource. This is formally equivalent to a portfolio allocation problem: allocate attention tokens to signals with the highest expected information value.
The cooldown mechanism implements a temporal version of this principle. Immediately after a notification, the marginal value of the next notification is low—the user is already context-switched and engaged. The cooldown allows attention to reset before the next signal, ensuring each notification has maximum impact. The dismiss feedback loop is a Bayesian update on the signal-value prior: each dismissal reduces the estimated value of similar signals, improving the allocation over time.
The preference for false negatives (missing a signal) over false positives (creating noise) reflects an asymmetric loss function. A missed signal costs the user one delayed decision. A noisy notification stream costs the user the entire notification system—they disable it. The asymmetric cost makes conservative thresholds optimal even when they occasionally miss relevant signals.
Limitations
- The attention model requires a model provider for scoring. If the model is unavailable, the agent falls back to rule-based scoring which is less nuanced. Rule-based scoring cannot distinguish urgency levels within the same signal type.
- Active hours are a blunt instrument. The agent does not know whether you are in a meeting, sleeping, or on vacation without explicit configuration. Calendar integration can partially address this, but the agent does not infer availability from context alone.
- Dismiss feedback creates a local optimum: if you dismiss everything, the agent learns to show nothing. This is correct behavior but can feel like the agent stopped working. The recovery path is to explicitly engage with suggestions for several days to recalibrate the score model.
- The dark-factory principle assumes the agent's work is valuable enough to justify running unattended. For ad-hoc tasks, a conversational interface is simpler and cheaper. The overhead of setting up monitoring, scoring, and delivery is not justified for one-off questions.
- The attention budget is global, not per-project or per-source. A busy week on one project can exhaust the budget, suppressing signals from other projects. Per-project budgets are a possible future extension.
- Cooldown timers prevent notification cascades but can also delay genuinely urgent signals. If two critical events happen within the cooldown window, the second is delayed. The system favors preventing noise over guaranteeing immediate delivery of every urgent signal.
- The auto-draft capability is limited to low-risk internal actions. External write operations—sending messages, posting comments, pushing code—always require approval. This is a safety feature but can create an approval bottleneck during high-activity periods.
Four guardrails that make unattended operation viable
Semibot's dark factory (Goal Factory) runs on a single goal kernel—there is no second task system—and the session composer is the only entry point for creating a goal. Unattended operation is viable not because the model got stronger, but because of four explicit gates. First, approvals cannot be pre-authorized: scheduled tasks carry an approval-mode snapshot from creation time, and when an unattended run hits a wait-for-confirmation step it stops and waits; it is never auto-escalated to auto-approval because the user is away. Second, spend must be pre-authorized: every background run holds a points budget fixed at creation before its first managed call; hitting the cap pauses the run and notifies the user—no spending first and asking later. Third, no self-replication: a running task cannot create, delete, or reschedule tasks, preventing the scheduler from propagating itself. Fourth, no unattended credential entry: scheduled browser runs remain subject to the hard login and payment gates.
Just as instructive is a recorded failure. In one release review, creating a goal unconditionally cleared the composer draft, silently discarding the user's materials and intent—unrecoverably. It was ruled a P0 data-continuity gate failure, the release was rejected, and the fix required either persisting the materials into the goal contract or blocking submission with an explicit explanation. The lesson generalizes: the real risk distribution of unattended systems is not wrong model reasoning but silent data loss in engineering seams.
The applicability boundary is written into the design: the dark factory targets complex goals with a clear completion endpoint. Open-ended continued attention belongs to the secretary's continuous-delegation model; the two divide labor architecturally rather than blending together.
FAQ
Does the agent run all the time?
The secretary runs on its schedule. Other tasks run when assigned. The attention system is always active, scoring signals from all running tasks.
What happens when the app is closed?
On desktop, scheduled tasks resume when the app reopens. Pending results that were generated while the app was closed are delivered on next launch.
Can I set different noise levels for different tasks?
Not per-task in the current design. The attention budget and cooldown are global. Per-task thresholds are a possible future extension.
Is this the same as Do Not Disturb?
DND suppresses all notifications. The attention model is selective: high-scoring signals can still reach you during active hours even if lower ones are suppressed. They solve different problems.
How does the agent know what is "material"?
The score model combines rule-based heuristics (new content, status change, error) with model-based relevance scoring. The threshold for "material" is tuned conservatively.
Can I use dark-factory mode with Gmail or calendar integrations?
Yes. Connected services through Semibot's connector system follow the same attention model. The secretary can monitor Gmail, Calendar, Feishu, Slack, DingTalk, and Telegram, applying the same score pipeline to each signal source.
What happens if I ignore the agent for a week?
Results accumulate on the secretary page. Nothing is lost. When you return, you see a chronological feed of everything the agent produced. High-scoring items are visually distinguished. If too many items accumulated, the agent may produce a summary-of-summaries to help you catch up.
How much does autonomous operation cost in model calls?
Idle monitoring costs very little—rule-based scoring does not require model calls. Model-based scoring is triggered only when new signals are detected. A typical day with 5–10 monitored sources might produce 10–20 scoring calls and 2–3 generation calls for summaries or auto-drafts. Cost scales with the number of active monitors and the frequency of changes in monitored sources.
Can I export the notification history or scoring data?
The secretary page retains all delivered results with timestamps and session links. Scoring metadata is used internally for feedback learning but is not currently exposed as a user-facing export. The delivered results themselves are standard Semibot artifacts and can be accessed through the file system.
Is dark-factory mode available on mobile?
Semibot is a desktop application. However, notification delivery through connected channels (Telegram, Feishu, Slack) means you can receive agent signals on any device that supports those channels. The agent itself runs on the desktop; the notifications travel through the connectors.
How is this different from a cron job with email alerts?
A cron job runs at a fixed time regardless of whether anything changed. It has no attention model—every run produces an alert. It cannot learn from your feedback. And it cannot adapt its behavior based on workspace context. The dark-factory agent is change-aware, scored, adaptive, and context-sensitive.
What if two tasks produce conflicting signals?
The attention budget mediates between competing signals. If two high-scoring signals arrive simultaneously, the first gets through; the second waits for the cooldown. If both are critical, the agent can batch them into a single notification. This is imperfect—it means the second signal is delayed—but it prevents the notification cascade that would result from unlimited escalation.
