Definition: Continuous delegation is a contract between a user and an AI assistant: describe a job once, and the AI keeps following up—checking for changes, compiling digests, collecting results in one place—and only pings you when it needs input or an approval. It replaces two unsatisfying alternatives: setting up rigid cron schedules manually, and repeatedly asking the same question in chat.
Why chatbots drop the ball on ongoing tasks
Chatbots are session-bound. When the session ends, context is preserved in history but the active task is not. Reminders can be set, but they fire at a time without understanding what has changed. Results scatter across different threads. If you want a weekly project digest, you need to remember to ask for it each time—or set up a rigid schedule with a separate tool.
The problem is not that chatbots are bad at answering. It is that ongoing tasks need a different contract: stay watching, interrupt only when needed, and collect results in a predictable place.
The secretary contract
Semibot's secretary implements continuous delegation with five design commitments:
- Plain-language delegation. "Keep an eye on the token-reseller scene and send me your read when something moves" is a valid instruction. No cron syntax. If you give a rhythm, the secretary follows it. If not, it proposes a sensible default and tells you what it chose.
- Change-aware delivery. The secretary does not just run on a timer. It compares current state against the last delivered result. If nothing moved, it stays quiet (unless you asked for a scheduled delivery regardless). This is the key difference from a reminder.
- Result collection. All results land on the secretary page with links back to sessions, artifacts, and tasks. You do not need to search through chat history to find what the agent produced last Friday.
- Interrupt only when needed. Missing information or an approval request triggers a ping. Every intermediate step does not.
- Same session continuity. Supplementary requirements go into the same session. The secretary does not create a new conversation for each follow-up—it maintains context.
The theory of change
The secretary model changes the user's relationship with AI from "ask and receive" to "delegate and review." This has three practical effects:
- Cognitive load shifts. The user no longer needs to remember what to check and when. The secretary holds the watch.
- Work accumulates. Instead of one-off answers that disappear into chat history, results collect in a structured feed. Over time, the feed becomes a work log.
- Interruptions become signal. Because the secretary is quiet by default, a ping means something: "I need your input" or "this changed and you should see it."
Continuous delegation vs other models
| Dimension | Chat | Cron / reminders | Secretary |
|---|---|---|---|
| Understands the work | Per session | No | Yes, from workspace context |
| Change detection | No | No | Yes, compares against last delivery |
| Result collection | Scattered in chat | — | Secretary page with session links |
| Needs cron syntax | No | Yes | No, plain language |
| Approvals | — | — | Integrated with existing gates |
What the secretary will not do
- It will not bypass approvals or folder grants. The secretary uses the same security model as the rest of Semibot.
- It will not dump all history into one giant chat. Follow-ups stay bound to the right session.
- It will not claim to be watching when configuration is missing. If there is no model, no connector, or no folder grant, it tells you.
- It will not replace deep creative work. The secretary handles monitoring and summarizing; deep investigation belongs in a focused session.
Workflow examples
Continuous delegation is best understood through scenarios that show the full lifecycle—from initial delegation to ongoing delivery.
Scenario 1: Market intelligence digest. You tell the secretary: "Follow the tokenized real-estate space. Summarize new developments every Monday morning." Step 1: The secretary parses the instruction, identifies the topic domain, and proposes a rhythm—Monday 9 AM. Step 2: It begins monitoring: web searches, arxiv papers, news feeds, and any connected channels you specify. Step 3: Each Monday, it compares new findings against the previous week's digest. Step 4: If nothing significant changed, it stays quiet—unless you specifically asked for a "no news" confirmation. Step 5: If there are developments, it compiles a structured summary with sources and places it on the secretary page. Step 6: If a development is urgent (e.g., a major regulatory change), the attention score crosses the notification threshold and you receive a ping before Monday.
Scenario 2: Pull request triage. You delegate: "When a new PR is opened in our repo, check if it touches the auth module. If it does, prepare a security-focused review." Step 1: The secretary binds to the repository and sets up monitoring. Step 2: When a PR arrives, it checks the diff against the auth module path. Step 3: If the PR does not touch auth, it logs the check and stays quiet. Step 4: If it does, the specialist assigned to this task reads the diff, prepares a review with security annotations, and auto-drafts it. Step 5: The draft appears on the secretary page with an approval button. Step 6: You approve, edit, or dismiss. The dismissal feedback adjusts future review focus.
Scenario 3: Document change tracking. You say: "Watch the Q4 strategy doc. Notify me when someone adds or changes the financial projections section." Step 1: The secretary stores the current state of the document as a baseline. Step 2: It checks periodically (or on document update events if the connector supports it). Step 3: It compares the current version against the baseline, focusing on the specified section. Step 4: If only formatting changed, it updates the baseline and stays quiet. Step 5: If the financial projections section changed, it surfaces the diff with context. Step 6: The baseline updates to the new version, so the next check compares against the latest delivered state.
Decision framework: when to delegate vs do it yourself
| Question | Answer "yes" | Answer "no" |
|---|---|---|
| Is this task recurring? | Delegate to secretary | Do it in chat or use a one-shot task |
| Does the task need change detection? | Delegate to secretary | A cron job or reminder may suffice |
| Should results accumulate in one place? | Delegate to secretary | One-off chat is fine |
| Does the task require deep creative work? | Use a focused session, not the secretary | Secretary can handle it |
| Is real-time (sub-second) response needed? | Use a webhook or event-driven automation | Secretary's polling model works |
| Does the task require the secretary's method knowledge? | Assign a specialist to the secretary task | The secretary can handle it directly |
Connector and integration patterns
The secretary model becomes significantly more powerful when connected to external services through Semibot's connector system. Each connector provides a data source for monitoring and a delivery channel for results.
- Feishu / DingTalk. The secretary can monitor group chats for mentions, action items, or document links. It can compile weekly summaries of team activity and deliver them through the same channel. For Chinese-speaking teams, the secretary supports bilingual operation: it reads and produces content in the language of the source material.
- Gmail. Monitor specific senders, labels, or search queries. The secretary can draft replies for routine emails and queue them for approval. Combined with the attention model, only high-priority emails trigger notifications; the rest accumulate in a digest.
- Google Calendar. The secretary reads upcoming events and prepares briefings: attendee information, relevant documents from the workspace, and notes from previous meetings on the same topic. Briefings are delivered to the secretary page before the meeting, not as push notifications.
- Slack. Channel monitoring follows the same pattern as Feishu. The secretary can track threads, summarize discussions, and flag messages that require your input. The attention budget applies across all monitored channels, preventing Slack noise from becoming notification noise.
- Telegram. Useful for personal workflows and mobile delivery. The secretary can push high-scoring digests and pings through Telegram, keeping you informed when away from the desktop app.
- File system and git. The secretary can monitor directories for file changes, new commits, or branch updates. This is the foundation for the PR triage and document change tracking scenarios above.
Comparison with alternative delegation models
| Dimension | Manual chat (ask each time) | Cron + email | Zapier / automation | Secretary |
|---|---|---|---|---|
| Setup effort | None (re-ask each time) | Medium (write cron + script) | Medium (configure zap) | Low (plain language) |
| Context awareness | Per-session | None | Per-zap | Persistent workspace context |
| Change detection | No | Manual diff logic | Trigger-based | Built-in, compares against last delivery |
| Result collection | Scattered in chat history | Email inbox | Spreadsheet or webhook | Secretary page with session links |
| Intelligence | Full model per query | None (fixed logic) | Limited (rules + some AI) | Full model with specialist support |
| Cost | High (full session per ask) | Free (self-hosted) | Subscription fee | Model calls only (no separate subscription) |
| Noise management | N/A (user-initiated) | None | Basic filtering | Attention score model with budget and cooldown |
Failure modes and boundary conditions
Continuous delegation is powerful but has specific failure modes that users should understand:
- Delegation drift. Over many cycles, the secretary may subtly drift from the original intent. If you delegated "summarize competitor activity" and the competitor shifts strategy, the secretary may continue monitoring the original dimensions without realizing they are no longer relevant. Periodic review of active delegations prevents this.
- Baseline staleness. Change detection depends on comparing current state against the last delivered state. If the secretary has not checked a source for a long time, the delta becomes large and the digest may be overwhelming rather than crisp. Appropriate check intervals mitigate this.
- Cross-delegation conflicts. If you delegate overlapping tasks—"monitor pricing" and "monitor all product changes"—the secretary may produce redundant results. The result collection system shows both, but it is the user's responsibility to avoid semantic overlap in delegations.
- Model availability. The secretary requires a model provider for generation and scoring. If the provider is down, the secretary cannot produce useful output. Pending results are queued and delivered when the model becomes available, but time-sensitive alerts may be delayed.
- Approval queue buildup. If a delegated task frequently requires approvals (e.g., auto-drafted replies to emails), the approval queue can become a new source of notification pressure. The dark-factory attention model mitigates this, but users should set up delegations with appropriate auto-execute thresholds.
- Resource consumption. Each delegated task consumes model calls on its schedule. Ten active delegations checking hourly can produce significant model call volume. Users should balance delegation count against their model quota and adjust check frequencies accordingly.
Theoretical foundations: from push/pull to continuous delegation
Traditional human-computer interaction models describe two paradigms: pull (the user asks, the system responds) and push (the system sends, the user receives). Continuous delegation introduces a third paradigm: the user specifies intent once, and the system maintains a persistent, evolving relationship with that intent. This is closer to how human delegation works—a manager does not re-issue instructions every morning; they describe the role once and expect the assistant to adapt to changing circumstances.
The key theoretical insight is that ongoing tasks have temporal structure that one-shot queries do not. A weekly digest is not the same query repeated 52 times—it is a single evolving task with state (previous digest), a detection function (what changed), a delivery function (how to present), and a feedback loop (user engagement adjusts future behavior). The secretary model encodes this temporal structure explicitly rather than leaving it implicit in the user's memory and manual re-asking behavior.
The change-aware delivery mechanism draws on event-driven architecture principles from distributed systems. Instead of polling at fixed intervals and producing output regardless of state changes, the secretary implements a compare-and-emit pattern: it produces output only when the delta between current state and last-emitted state exceeds a threshold. This is the same principle behind change data capture in databases—emit the diff, not the full state.
Limitations
- Needs a model provider for scoring and generation. Without one, the secretary cannot do useful work.
- Change detection is text-based comparison, not semantic understanding. Subtle changes in meaning may be missed.
- The secretary model is still evolving. Delivery rhythms, result presentation, and history management are areas of active development.
- For tasks that need real-time monitoring (sub-second), the secretary's polling model is too slow. It is designed for human-timescale checks (hourly, daily, weekly).
Implementation choice: the discipline of not building a scheduler
The most notable decisions in the secretary's implementation are the things it does not build: no dedicated task table, no scheduler, no tool gateway, no approval state machine of its own. A delegation is a scheduled task with a delivery policy, reusing the existing scheduling service and runtime: scheduled delivers on the agreed cadence—including the necessary “no changes this period”—and on_change delivers only when something relevant changes. Natural-language delegation does not need the words “daily” or “weekly”: “keep an eye on this” and “tell me when it changes” are continuous delegations. When the cadence cannot be parsed, the system falls back to a daily default and says so explicitly, rather than guessing silently.
The quiet principle is written as decidable rules: no change alerts without important changes; agreed deliveries still report on time; and unavailable sources do not count as “no change.” Together these prevent the system from using silence to feign normality. Change detection compares against the last delivered result, and it is delivery information—not a completion gate for the task.
The interruption-side information architecture collapses into three zones—results (what is worth seeing), in progress (what the AI is doing for you), and needs you (which step requires you)—and “needs you” aggregates only two request types: authorization requests and required-information requests. Ordinary results, monitoring alerts, failures, and verification gaps create no pending obligation; opening or closing the panel is never an approval or an answer. Modifying a delegation (“make it weekly”, “keep it shorter”, “pause that one”) updates the same task definition rather than creating a new one, and resuming restores the original requirements and history. Success metrics are explicit: time to first useful result, delivery adoption rate, two-week continued-delegation retention, and time-to-correct human errors—page views are not a success criterion.
FAQ
Do I need to write cron expressions?
No. Plain language works. If no rhythm is given, a default is proposed and shown to you.
How is this different from a specialist?
A specialist remembers how a job should be done. The secretary keeps an ongoing watch and delivers results over time. A secretary task might use a specialist to do the actual work.
Can I pause a secretary task?
Yes. Tasks can be paused and resumed. Pending items remain until you handle or dismiss them.
Where do results go?
They collect on the secretary page, with links back to sessions, artifacts, and tasks.
Does the secretary cost extra?
No separate subscription. Model calls consume your existing quota or your own provider.
How many active delegations can I have at once?
There is no hard limit. However, each active delegation consumes model calls on its schedule. Ten delegations checking hourly will use significantly more quota than two checking daily. The secretary page shows active delegations and their check frequencies so you can manage cost.
Can the secretary work with Feishu, DingTalk, or Telegram?
Yes. Any connected service can be a data source for monitoring or a delivery channel for results. The secretary can monitor a Feishu group for updates and deliver digests through Telegram, for example. Connector availability depends on which services you have connected in Semibot's settings.
Can I see the history of what the secretary delivered?
Yes. The secretary page maintains a chronological feed of all delivered results, with links to the originating sessions, artifacts, and tasks. Nothing is lost. You can scroll back through weeks or months of deliveries.
What if the secretary misunderstands my delegation?
When you first delegate, the secretary confirms its understanding—what it will monitor, at what rhythm, and how it will deliver. If the first result does not match your expectations, you can adjust the delegation in the same session. The secretary maintains context across follow-ups, so corrections are cumulative rather than starting from scratch.
Can the secretary use specialists for the actual work?
Yes. A secretary task can be assigned to a specialist, which brings the specialist's method, tools, and behavioral constraints to the delegated work. For example, you might have a "code reviewer" specialist and assign it to a secretary task that monitors pull requests. The specialist does the review; the secretary handles the monitoring, change detection, and delivery.
How does the secretary handle errors or failures?
If a check fails (network error, service unavailable, model provider down), the secretary logs the failure and retries on the next scheduled check. Persistent failures are surfaced as low-scoring notifications after a configurable number of retries. The secretary does not silently drop failed checks—it tracks them.
Can I delegate the same task to multiple delivery channels?
A single delegation delivers to one primary location (the secretary page), with optional notification forwarding to connected channels. If you need the same digest delivered to both Slack and email, you would set up the primary delivery on the secretary page and configure notification forwarding to the additional channels.
