SemibotSemibot - AI Desktop
All guides

Attention Policy: Why an AI Assistant Should Be Quiet by Default

The biggest risk in a proactive AI assistant is not that it is not smart enough—it is that it is too noisy. The attention policy resolves this by treating every potential signal through a scoring pipeline before it reaches the user.

Definition: An attention policy is the set of rules that determines when an AI assistant is allowed to interrupt the user. In Semibot V3, the policy defaults to conservative: every signal must pass through deduplication, relevance scoring, attention budget limits, risk policy checks, cooldown timers, and user feedback weighting before it can become a notification. The design philosophy is "prefer missing an alert over creating noise."

Why proactive assistants fail

The failure mode of a proactive assistant is not that it misses things—it is that it notifies too often. A user who receives too many notifications will disable them entirely, then miss everything. The asymmetric cost structure means:

  • One missed notification: user can recover by checking the summary later.
  • Too many notifications: user disables the system permanently.

This asymmetry favors restraint as the default policy.

The score model

Every signal is scored before it can reach the user:

score = ruleScore + recencyScore + relationshipScore + urgencyScore + userPreferenceScore - noisePenalty - cooldownPenalty

Each component serves a different purpose. Rule score captures structural signals (new content, status change, error). Recency score weights newer signals higher. Relationship score considers the connection between the signal and the user's active tasks. Urgency score adds time-sensitivity. User preference score incorporates learned patterns from dismissals. Noise penalty reduces score for signals from noisy sources. Cooldown penalty suppresses signals that arrive too soon after a previous notification.

Threshold tiers

TierThresholdBehavior
Suggest≥ 0.70Appear in suggestion area, no push notification
Notify≥ 0.82Push notification, active hours only
Auto-draft≥ 0.90 + low riskPrepare action without asking; external writes never auto-execute

Attention budget and cooldown

The attention budget caps how many signals can reach the user per day. This is a throttle on interruptions, not on work. The agent can work continuously—it just cannot continuously demand attention. After a notification, a cooldown timer prevents the next one for a configurable period, preventing the common failure mode of related signals arriving in rapid succession.

Dismiss feedback

When a user dismisses a suggestion, the dismissal adjusts user preference weights so similar signals score lower in the future. This creates a feedback loop: the agent starts conservative, the user's dismissals teach it what "too noisy" looks like, and the signal-to-noise ratio improves over time. The alternative—asking users to configure notification preferences manually—is worse, because most users will not configure thresholds.

Scenario: a day in the life of the attention policy

To make the scoring pipeline concrete, consider what happens during a typical working day when the agent is monitoring your codebase, a research project, and a competitor's public feed:

  1. 9:15 AM — Competitor publishes a blog post. The secretary detects new content. Rule score (new content) = 0.4. Recency score = 0.3. Relationship score (you have a competitive analysis skill active) = 0.2. Total = 0.90. This crosses the Notify threshold (≥ 0.82). A push notification appears: "New competitor post about their pricing model."
  2. 9:20 AM — CI build fails on main. Rule score (error status change) = 0.5. Urgency score (blocking team) = 0.3. Relationship score (you committed to main yesterday) = 0.2. Total = 1.0. Crosses Auto-draft threshold. The agent prepares a summary of the failure with the relevant stack trace, but does not auto-fix. External writes never auto-execute.
  3. 9:45 AM — Another CI build fails on a feature branch. Cooldown penalty = −0.3 (only 25 minutes since last notification). Raw score would be 0.85, but after penalty = 0.55. Below Suggest threshold. The failure is recorded in the summary but does not surface until the user asks or the next summary cycle.
  4. 11:00 AM — Research paper update detected. Rule score = 0.3. Recency = 0.2. Relationship score = 0.1 (you have a research session open but it is not the active window). Total = 0.60. Below all thresholds. Silently logged.
  5. 2:00 PM — User dismisses the competitor notification. Dismiss feedback adjusts user preference weight for "competitor blog post" signals downward. Future signals from this category score ~0.1 lower.

By end of day, the user received 2 push notifications instead of the 5+ they would have received without the attention policy. The agent still processed all 5 signals—it just chose which ones were worth interrupting the user for.

Scenario: the noisy failure mode without the policy

Without the attention policy, the same day produces a different experience. The competitor blog post triggers a notification. The CI failure triggers a notification. The second CI failure triggers another notification 25 minutes later. The research paper update triggers a notification. The user receives 4 notifications by lunch, starts ignoring them, and by 3 PM has turned off notifications entirely. Now the user misses the afternoon's critical CI failure on main that blocks the release.

This is the core asymmetry: missing one notification is recoverable (the user can check the summary). Missing all notifications because the system was too noisy is not recoverable (the user has lost trust and disabled the system).

Scenario: adapting to a user who never dismisses

Some users never dismiss notifications—they simply ignore low-value ones without clicking. The attention policy handles this through implicit feedback signals: if a notification is delivered but the user does not interact with the agent within a configurable window, the system treats this as a soft dismissal. Over time, the user preference weights adjust downward for signal categories that are consistently ignored. This is less precise than explicit dismissals but prevents the "user who never clicks" from being permanently flooded.

Decision framework: attention policy vs. alternatives

There are several approaches to managing AI assistant interruptions. The attention policy is one design choice among several:

ApproachMechanismStrengthWeakness
Attention policy (Semibot)Score model + budget + cooldownAdaptive, learns from feedbackRequires model provider for scoring
Rule-based filteringStatic thresholds per signal typePredictable, no model dependencyCannot adapt to individual preferences
User-configured channelsUser picks which event types to receiveFull user controlMost users never configure; config goes stale
Digest/batch modeCollect all signals, deliver periodicallyZero interruptions during workMisses time-sensitive signals
LLM-as-judgeLLM decides each notification individuallyMost nuanced understandingHigh latency and cost per signal
Push everythingNo filteringNever misses anythingGuaranteed noise fatigue

Semibot's attention policy blends the score model with the budget and cooldown, creating a hybrid that is more adaptive than pure rules but more predictable and cheaper than LLM-as-judge for every signal.

Comparison: attention policy design dimensions

Evaluating any attention management system requires examining these dimensions:

DimensionSemibot ApproachNotes
Default postureConservative (prefer missing over noise)Asymmetric cost favors restraint
Adaptation mechanismDismiss feedback adjusts weightsImplicit feedback for non-clickers
GranularityPer-signal scoringNot per-channel or per-task
Rate limitingDaily budget + per-notification cooldownTwo layers of throttling
Escalation pathSuggest → Notify → Auto-draftThree tiers with increasing intrusiveness
Failure mode on provider outageFallback to rule-based scoringLess nuanced but functional
User overrideThreshold adjustment + channel controlScore model itself stays active

Theoretical depth: information foraging and interruption cost

The attention policy draws on information foraging theory, which models how users seek and consume information. In this framework, each notification is a "patch" that may or may not contain valuable information. The user's decision to attend to a notification depends on the expected information value versus the interruption cost.

Research on knowledge worker interruptions shows that the cost of an interruption is not just the time to read the notification—it includes the context-switching cost of returning to the previous task. Estimates range from 10 to 25 minutes to regain full context after an interruption. This asymmetry means that even a 70% accurate filter that blocks 10 interruptions saves more productivity than a 95% accurate filter that lets 2 unnecessary interruptions through.

The attention budget implements a daily cap on these context-switching costs. The cooldown timer implements a minimum interval between context switches. Together, they bound the total interruption cost regardless of how many signals the agent detects.

The secretary integration

The secretary is Semibot's continuous monitoring system. It watches external sources—RSS feeds, web pages, file changes, API endpoints—and produces change signals. Every change signal flows through the same attention pipeline as internal events. The secretary does not decide what to notify you about; it produces signals, and the attention policy decides whether they are worth interrupting you for.

This separation is important. If the secretary had its own notification logic, you would have two systems competing to interrupt you with no coordination. By funneling all signals through a single attention pipeline, the system maintains a coherent interruption budget.

Change-aware delivery means the secretary only produces signals when content actually changes, not on a polling schedule. This reduces the raw signal volume before scoring even begins. A page that has not changed in 24 hours produces zero signals; a page that changes every hour produces signals that the cooldown timer then throttles.

Boundary conditions: when the attention policy fails

The policy is not infallible. It fails in specific scenarios:

  • Cold start. With no dismissal history, the user preference weights are zero. The policy relies entirely on rule-based and relationship scores. Early notifications may be poorly calibrated until the user provides feedback.
  • Category collapse. If the user dismisses enough signals across all categories, the policy learns to suppress everything. The user perceives the assistant as "broken" when it is actually following the learned preference model. This requires explicit re-engagement to reset.
  • Emergent urgency. A signal that is individually low-scoring but collectively important (three medium-severity warnings from different sources that together indicate a system-wide outage) may be missed because the policy evaluates each signal independently.
  • Cross-task relevance. The policy is global, not per-task. A high-priority signal for a background task may be suppressed because the user's active task context lowers its relationship score.
  • Model provider dependency. When the model provider is unavailable, the fallback rule-based scoring cannot capture nuance. A signal that would score 0.85 with model scoring might score 0.50 with rules, causing it to be missed.

The attention budget in detail

The attention budget is a daily cap on the number of push notifications the system can send. It is not a cap on work—the agent processes all signals regardless of the budget. It is a cap on interruptions. When the budget is exhausted, new signals that meet the Suggest threshold still appear in the suggestion area, but no further push notifications are sent until the budget resets.

The budget resets at a configurable time (default: midnight local time). This daily cadence ensures that even a very noisy day does not carry over into the next day. The user starts each day with a full budget and a fresh opportunity to receive relevant notifications.

Budget allocation is first-come, first-served: the highest-scoring signals that arrive first consume the budget. There is no reservation system for "important" signals. If the budget is nearly exhausted and a critical signal arrives, it may not get a push notification. This is by design—if the budget is nearly exhausted, the user has already been interrupted many times today and additional interruptions (even for critical signals) carry diminishing value.

Score component breakdown

Each component of the score model captures a different aspect of signal quality:

  • Rule score (0–0.5): Structural signals like new content, status changes, or errors. These are the baseline signals that are always relevant regardless of context. A CI failure always has a non-zero rule score.
  • Recency score (0–0.3): Newer signals score higher. A signal that is 5 minutes old scores higher than the same signal that is 5 hours old. This prevents stale information from generating notifications.
  • Relationship score (0–0.3): Signals related to the user's active tasks score higher. If the user is actively working on the codebase, a CI failure has a higher relationship score than a competitor blog post.
  • Urgency score (0–0.3): Time-sensitive signals score higher. A blocking CI failure on main scores higher than a non-blocking warning on a feature branch.
  • User preference score (−0.2 to +0.2): Learned from dismissals and interactions. Positive for signal categories the user engages with; negative for categories the user consistently dismisses.
  • Noise penalty (0 to −0.3): Applied to signals from sources that produce frequent low-value output. A chatty log stream has a higher noise penalty than a quiet error monitor.
  • Cooldown penalty (0 to −0.5): Applied when a signal arrives too soon after a previous notification. The penalty decreases linearly as the time since the last notification approaches the cooldown period.

Limitations

  • Scoring requires a model provider. If unavailable, fallback to rule-based scoring is less nuanced.
  • Active hours are a blunt instrument—the agent does not know your calendar.
  • If you dismiss everything, the agent learns to show nothing. Correct behavior but can feel like it stopped working.
  • The policy is global, not per-task. Different tasks cannot have different notification thresholds in the current design.

Implementation contract: an orthogonal decomposition of the decision space

The most consequential engineering decision in Semibot's attention policy is splitting “should this interrupt you” from “should the AI be allowed to do this” into two orthogonal decision types. AttentionDecision (skip / inbox_only / notify / auto_draft) answers only the interruption question; CapabilityDecision (allow / ask_approval / deny) answers only the authorization question. Neither can substitute for the other: a high-scoring external write operation is never executed automatically—it degrades to a notification at most. This is not a pile of rules but a decomposition of the decision space, and it yields a structural guarantee: no matter how aggressive the noise-reduction becomes, it cannot produce an unauthorized action.

Every signal passes through six gates in a fixed order: dedup, relevance scoring, attention budget, risk policy, cooldown, user feedback. The order is itself a design property—dedup runs first so duplicate signals never consume scoring or budget, and feedback runs last so it shapes future signals without touching the current decision. The budget is a hard cap along four dimensions (per day, per source, per risk level, per category): 3 high-priority notifications and 10 medium suggestions per day by default; over-budget signals are merged, demoted to the inbox, or deferred. The budget is a hard gate, not a soft score that a high total can compensate—products of the same insight as the Goodhart warning in measurement theory: when the score becomes the target, the system optimizes the score rather than the judgment.

The feedback loop defines five behaviors—approving raises relevance for similar signals, dismissing lowers it, “fewer like this” strengthens the cooldown, “remind later” only defers without demotion, and disabling a source stops it entirely—and every effect must be reversible and visible in settings. Notifications are whitelisted to three categories (high-value proposals, runs awaiting approval, broken connector authorization) and are barred from carrying email bodies or private calendar details. Every surfaced proposal must answer four questions: why it matters, why now, why the permission is needed, and how to see fewer like it. The policy ships as an AttentionPolicyService whose acceptance tests explicitly cover budget overruns, quiet hours, and the no-auto-execution-of-high-risk rule.

FAQ

Is this the same as Do Not Disturb?

No. DND suppresses all notifications. The attention model is selective: high-scoring signals can still reach you during active hours.

Can I turn it off?

You can adjust thresholds or disable specific notification channels. The score model itself is always active—without it, every signal would reach you.

How does it learn my preferences?

From your dismissals. Each dismissal adjusts the user preference weight for that signal category.

Does the secretary use this?

Yes. The secretary's change-aware delivery feeds through the same attention pipeline. A change that does not meet the threshold does not trigger a notification.

What is the difference between the Suggest and Notify tiers?

Suggest (≥ 0.70) places the signal in a suggestion area that the user can browse at their leisure. Notify (≥ 0.82) sends a push notification during active hours only. The difference is whether the system actively interrupts the user.

Can different tasks have different notification thresholds?

Not in the current design. The policy is global. A future per-task policy would allow critical tasks (like production monitoring) to have lower thresholds than background tasks (like research).

How quickly does dismiss feedback take effect?

Immediately. The next signal in the same category will use the updated weight. There is no batch training or delayed update cycle.

What happens during off-hours?

Signals are still scored and logged, but push notifications are suppressed during configured off-hours. The Suggest tier remains active so signals accumulate in the suggestion area for the user to review when they return.

Can I set a per-source budget?

Not currently. The attention budget is global. The noise penalty partially addresses this by reducing scores for sources that produce frequent low-value signals, but there is no hard per-source cap.

What is Auto-draft and is it safe?

Auto-draft (≥ 0.90 + low risk) prepares an action without asking. The key safety property is that external writes never auto-execute—the agent can prepare a draft email or code change, but the user must approve before anything leaves the system.

How does this relate to the Dark Factory principle?

The Dark Factory principle says agents should work autonomously and only surface results. The attention policy operationalizes this: the agent works continuously, and only interrupts when the signal score justifies the cost. Low-scoring work is completed silently and summarized on demand.

Does the cooldown apply to Auto-draft signals?

Yes. Even high-confidence signals are subject to the cooldown timer. This prevents the scenario where a burst of related high-confidence events causes a rapid sequence of auto-drafts.

Related