Definition: Semibot's built-in browser is an embedded web view that lets users select elements on any web page—buttons, text, images, regions—and attach them as structured references to a conversation. The agent receives the selected content, page context, and a screenshot of the area, without receiving the full DOM, form values, or executable node handles. The design treats web content as data, not as system instructions.
How element selection works
The address bar provides a "select element" entry point. Entering selection mode pauses the browser's automation layer, shows a selection hint, and begins highlighting elements on hover. Clicking captures a reference without triggering the page's own click handlers. Shift-click enables multi-select. Esc exits selection mode and focuses the conversation input.
The captured reference is a BrowserElementReference: a structured object containing the page title, sanitized URL (no username, password, query, or hash), capture time, visible text, element role, bounding rectangle, and a local screenshot. The agent receives this data—never the raw DOM node or CDP handle.
The privacy model
The privacy boundary is the core design constraint:
- URLs are sanitized. Username, password, query parameters, and hash fragments are stripped before storage.
- No form values. The DOM reader script runs in an isolated execution world and explicitly excludes input values, hidden text, and sensitive regions.
- Text and attribute limits. Captured text, attributes, and dimensions have upper bounds to prevent data exfiltration through oversized payloads.
- Screenshots are conservative. Elements containing input controls or sensitive markers are omitted from screenshots. Screenshots are stored in the app's private directory, not in the workspace.
- No system instruction injection. Web content is serialized as data in the conversation input. It cannot act as system instructions or override the agent's behavior.
Control and race conditions
When the user enters selection mode, the browser's automation is paused. The agent cannot navigate, click, or scroll while the user is selecting elements. Control is returned only when the selection session ends, and only if the control version matches—preventing stale responses from a previous selection from affecting the current one.
If the page navigates, the tab closes, or the session switches, the current selection is cancelled. Captured references are historical snapshots: they persist in the conversation history even after navigation, but re-executing an action on the same element requires re-observing the page and confirming the element identity.
Shadow DOM and iframe boundaries
The DOM reader supports open Shadow DOM for text capture. Same-origin iframes are supported with coordinate transformation. Cross-origin iframes, closed Shadow DOM, and Canvas/video elements cannot have their internal DOM claimed—the system shows the reason and preserves the existing draft. These boundaries are enforced, not configured.
Workflow examples
The built-in browser is most useful when you need the agent to see what you see—specific elements on specific pages—without exposing the full DOM or sensitive content. Here are three concrete workflows:
Scenario 1: Understanding an error on a web dashboard. Step 1: You open your monitoring dashboard in Semibot's browser. Step 2: You see an error panel with a cryptic message. Step 3: You enter selection mode and click the error panel. Step 4: The captured reference includes the error text, the surrounding UI context, and a screenshot. Step 5: You ask the agent, "What does this error mean and how do I fix it?" Step 6: The agent receives the element reference—text, role, page title, and screenshot—and provides a diagnosis. It does not receive the full page DOM, your session cookies, or any hidden form data. Step 7: If the fix requires navigating to a settings page, the agent can use browser automation tools (with approval) to do so.
Scenario 2: Comparing competitor pricing pages. Step 1: You open Competitor A's pricing page in the browser. Step 2: You enter selection mode and multi-select (shift-click) the pricing tiers. Step 3: You open Competitor B's pricing page and repeat. Step 4: You ask the agent: "Compare these pricing structures and identify where we are over- or under-priced." Step 5: The agent receives both sets of element references—pricing text, plan names, feature lists—and produces a structured comparison table. Step 6: Because the references are historical snapshots, you can revisit them later even if the pricing pages have changed.
Scenario 3: Extracting data from a research article. Step 1: You navigate to an academic paper's results page. Step 2: You select the results table and the key figure. Step 3: You ask, "Summarize these results and compare them with the approach described in our last internal report." Step 4: The agent receives the table text (structured as element references), the figure screenshot, and the page context. Step 5: It cross-references with documents in the workspace—your internal report—and produces a comparison. Step 6: If the table spans below the fold, you scroll down, select the continuation, and add it to the same conversation. The agent accumulates references across multiple selections.
Decision framework: when to use element selection vs other approaches
| Question | Use element selection | Use browser automation | Use web search / fetch |
|---|---|---|---|
| Is the content behind authentication? | Yes—browser has your session | Yes—same session | No—public URLs only |
| Do you need to act on the page? | No—capture only | Yes—click, fill, navigate | No—read only |
| Is the content dynamically rendered? | Yes—browser renders JS | Yes—browser renders JS | Maybe not—fetch gets raw HTML |
| Do you need privacy protection? | Yes—sanitized references | Moderate—full page access | High—no browser state involved |
| Is speed critical? | Fast—captures in place | Slower—requires scripting | Fastest—direct HTTP |
The general rule: use element selection when you need to show the agent something specific on a page you are already viewing. Use browser automation when you need the agent to interact with the page (click, fill forms, navigate). Use web search and fetch when the content is public and does not require JavaScript rendering.
Comparison with other browser-for-AI approaches
| Dimension | Playwright / Puppeteer (headless) | Browser extensions | Screenshot + OCR | Semibot built-in browser |
|---|---|---|---|---|
| User sees the page | No (headless) | Yes | No (image only) | Yes, full interactive browser |
| Element-level capture | Yes (selectors) | Yes (DOM access) | No (flat image) | Yes (structured references) |
| Privacy model | Full DOM exposed | Full DOM exposed | Screenshot only (limited) | Sanitized references, no form values |
| Authentication support | Manual (cookies, login) | Inherits browser session | N/A | Inherits user's browsing session |
| Prompt injection risk | High (raw DOM to model) | High (raw DOM to model) | Low (image only) | Low (data-only serialization) |
| Setup required | Install + configure | Install extension | Minimal | Built-in, no setup |
Integration patterns with Semibot's broader ecosystem
The built-in browser does not operate in isolation. It connects to Semibot's skill system, workspace context, and approval model to create end-to-end workflows that span web content and local work.
- Browser + deep-research skill. Select a paragraph from a news article and ask the agent to verify the claims. The deep-research skill can cross-reference the web content against multiple sources, producing a fact-check report with citations. The browser provides the starting point; the skill provides the verification methodology.
- Browser + arxiv skill. Select a citation from a blog post and ask the agent to find the original paper. The agent uses the selected text as a search query, invokes the arxiv skill to locate the paper, and produces a summary with a BibTeX citation. The browser bridges the gap between informal web content and formal academic sources.
- Browser + workspace files. Select data from a web dashboard and ask the agent to update a local spreadsheet. The agent captures the structured data from the browser, formats it, and writes it to a file in the workspace—subject to folder grants and approval. This bridges web data collection and local data processing.
- Browser + secretary delegation. You can delegate a recurring web monitoring task to the secretary: "Check this pricing page weekly and tell me if it changes." The secretary uses browser automation (not element selection) to capture the page state, compares it against the baseline, and surfaces changes. Element selection is for interactive use; automated monitoring uses browser automation under the hood.
- Browser + specialist. Create a specialist that knows how to analyze a specific type of web content—a competitor's product page, a financial dashboard, a regulatory filing. Assign it a Work Item with browser-captured references as input. The specialist applies its method to the captured data, producing structured analysis instead of ad-hoc commentary.
Failure modes and boundary conditions
The built-in browser has specific technical boundaries that users should understand to avoid frustration:
- Dynamic content timing. If a page loads content asynchronously (infinite scroll, lazy-loaded images, SPAs), the element selection may capture a loading state rather than the final content. The user must wait for the page to stabilize before selecting. There is no automatic "wait for load" in selection mode.
- Cross-origin iframe blindness. Many modern dashboards embed content from third-party domains in iframes. If the iframe is cross-origin, the DOM reader cannot access its contents. The user sees the iframe but cannot select elements inside it. Same-origin iframes work with coordinate transformation.
- Canvas and WebGL content. Dashboards that render charts using Canvas, WebGL, or SVG-as-image cannot have their internal data extracted through DOM reading. The screenshot captures the visual output, but the agent cannot access the underlying data points. For data extraction, the user should look for an underlying data table or API endpoint.
- Selection scope. Selection only works in the currently visible viewport. Content that requires scrolling to see is not captured. For long pages, the user must scroll, select, and repeat—each selection captures only what is visible at that moment.
- Stale references. Captured references are historical snapshots. If the page updates after capture, the reference still shows the old content. This is by design (privacy and consistency), but it means the agent cannot "re-read" the same element to check for updates. Re-capture is required.
- Complex page layouts. Pages with overlapping absolute-positioned elements, z-index tricks, or modal overlays can confuse the selection highlighter. If the wrong element is highlighted, the user can try hovering more precisely or use shift-click to adjust. This is a visual selection problem, not an agent problem.
- Chromium dependency. The browser is Chromium-based. Sites that specifically detect and block Chromium automation may behave differently than in a standard browser. This is rare but possible for sites with aggressive bot detection.
Theoretical depth: the content-as-data boundary
The browser's privacy model is grounded in a specific security principle: web content must be treated as data, not as instructions. This distinction matters because modern LLMs process all input text in the same attention window—a carefully crafted web page could, in theory, inject instructions that override the system prompt. By serializing web content as structured data (text content, element role, bounding box) rather than passing raw HTML, the browser prevents prompt injection from web pages.
The CDP (Chrome DevTools Protocol) architecture provides the technical foundation for this isolation. The DOM reader script runs in an isolated execution world—a separate JavaScript context that shares the DOM but not the page's JavaScript environment. This means the reader can access element text and attributes but cannot be intercepted by the page's event handlers, mutation observers, or prototype overrides. The selection mode's click interception uses the same isolation: it captures the element at the click coordinates without dispatching the click event to the page.
The URL sanitization follows OWASP guidelines for sensitive data in URLs. Query parameters often contain session tokens, API keys, or tracking identifiers. Hash fragments may contain OAuth tokens in single-page authentication flows. By stripping these before storage, the browser ensures that captured references do not inadvertently persist credentials in the conversation history.
Limitations
- Cannot capture content from cross-origin iframes or closed Shadow DOM. Same-origin iframes with complex nesting may also have capture limitations depending on the page's Content Security Policy.
- Canvas and video elements support visible-area screenshots only; internal state is not accessible. Charts rendered as Canvas (common in analytics dashboards) produce screenshots but not extractable data points.
- Selection only works in the current visible area. Scrolled-off content requires the user to scroll first. There is no "select all content on this page" operation.
- Browser element references cannot prove source code file or component ownership—they are web data, not workspace artifacts. A captured React component from a documentation site is text, not the component's source code.
- The CDP-based architecture means the browser requires Chromium support. Non-Chromium web views are not supported. Sites that detect and block Chromium automation may behave unexpectedly.
- Text extraction is limited to visible text nodes. Text rendered as images (common in infographics, memes, and some older sites) is captured in screenshots but not as selectable text.
- Multi-select captures elements in order but does not automatically detect relationships between them (e.g., which table row a cell belongs to). Structural relationships must be inferred by the agent from the captured context.
- The browser shares the desktop application's process. Extremely heavy pages with many DOM nodes can slow the browser and, by extension, the application. Performance on such pages may degrade.
How the agent sees: semantic snapshots, not pixels
Element selection in Semibot's built-in browser does not depend on a vision model reading screens. When you point at an element, the agent receives a structured semantic snapshot: the element's accessible role, name, and text; the page's sanitized address; capture time and structural path; plus a strictly bounded auxiliary screenshot (at most 640×480, trimmed to the target with a small margin, capped at 180 KB, conservatively omitted for inputs or sensitive regions). Selection runs through a control-state machine: the agent pauses page automation before entering pick mode, and each capture request binds a control revision and a page revision. Once the page changes, an old reference is a historical snapshot—re-execution requires re-observing and re-verifying the element; an old coordinate is not an execution permission.
The security posture rests on one principle: web content is data, never instructions. Reading form values, scripts, and hidden text is prohibited; text and attribute captures have length caps; URL sanitization strips credentials, query strings, and hashes; screenshots persist to an app-private directory with restrictive permissions; the renderer receives only snapshots—never the debugging endpoint or executable node handles—which doubles as a structural defense against prompt injection.
The boundaries are stated plainly: open shadow DOM can be reached; same-origin iframes need coordinate translation; cross-origin iframes, closed shadow roots, and the interiors of canvas and video cannot be claimed as DOM—only their visible region can be screenshotted. On the engineering side the module shipped with 141 desktop regressions and 50 style regressions, roughly 93% line and 84% branch coverage on the new capture modules—and an explicit note that self-assessment scores are not a release review.
FAQ
Can the agent click things on the page?
Only through explicit browser automation tools, which require approval. Element selection captures references without triggering page clicks.
Can the agent read my password fields?
No. The DOM reader explicitly excludes input values and sensitive regions.
Can I select multiple elements?
Yes. Shift-click enables multi-select. Each element gets its own reference with preserved order.
What happens if the page changes after I select?
The captured reference is a snapshot. It persists in history. Re-executing requires re-observing the page.
Can the agent fill out forms on the page?
Yes, through browser automation tools, which require explicit approval. Element selection captures references; form filling is a separate automation action. The agent can type into fields, select dropdowns, and submit forms—but each action requires your approval because it modifies external state.
Does this work with single-page applications (SPAs)?
Yes. The browser renders JavaScript, so SPAs work as they would in a normal browser. However, dynamic content that loads after the initial render may need time to appear before selection. If you see a loading spinner, wait for it to complete before selecting.
Can I use the browser to scrape data from multiple pages?
For interactive use, you can navigate to each page and select elements manually. For automated multi-page scraping, delegate a task to the agent using browser automation tools—the agent navigates, captures, and collects results across pages. This requires approval for navigation actions.
How does the browser handle pop-ups and modal dialogs?
Pop-ups and modals rendered as part of the page DOM are accessible through element selection. Browser-native dialogs (alert, confirm, prompt) are handled by the automation layer and may pause interaction until dismissed.
Can I select content inside a PDF viewer embedded in the page?
It depends on the viewer. PDF viewers that render as Canvas (most modern ones) cannot have their internal text extracted through DOM reading. The screenshot captures the visual content. For text extraction, download the PDF and use the PDF skill instead.
Does the browser store my browsing history?
The browser maintains a session history for navigation (back, forward). Captured element references are stored in the conversation history. URLs are sanitized before storage. The browser does not send browsing data to external servers.
Can I use the browser with a proxy or VPN?
The browser uses the system's network configuration. If your system is configured to use a proxy or VPN, the browser routes through it. Semibot does not configure network proxies independently.
How is this different from "browse with Bing" in ChatGPT?
ChatGPT's browsing mode fetches public URLs and extracts text. Semibot's built-in browser is a full interactive browser where you can see the page, select specific elements, authenticate with your own sessions, and interact with dynamic content. The key difference is user-directed selection: you choose exactly what the agent sees, rather than the agent crawling the full page.
Can the agent read content that requires a login?
Yes. The browser inherits your browsing session. If you are logged into a service, the agent can see and select content on authenticated pages—subject to the privacy model (no form values, sanitized URLs). The agent does not store or transmit your credentials.
