19 min read
AI assisted

ego-browser and Aside Fail Differently

Two browser agents on the same missions: structure and defects

To let an AI agent manage a shopping cart or organize a conversation list, it needs to read a logged-in page, operate its controls, and check what actually changed. Browser automation tools provide that connection. The tool determines which tabs the agent operates, how the user can intervene, and how failures and changes can be checked and reversed.

ego-browser and Aside both let an agent observe and operate a browser, but they organize the work differently. ego-browser puts the agent's tabs in a task space separate from user windows; Aside operates actual tabs in the user's browser. Alongside individual operations, Aside also offers exec, which accepts a whole task in natural language. They can handle the same tasks while differing in how the agent shares the browser with the user and how much the caller manages directly.

Comparing them helps establish what those differences mean in practice. Can the user keep browsing while the agent works? Can the caller detect a failed operation? Can changed data and tabs be put back in order afterward? I examined these questions through a Coupang shopping cart flow, popup and canvas tasks, conversation archiving and restoration, and behavior probes on controlled pages.

One difference emerged in how failures become visible. ego-browser returned success after missing an offscreen element or clicking another element. With Aside, TypeError and SyntaxError interrupted the work. Both wasted six rounds on cold runs, but finding and correcting the problems involved different work. Order effects prevented a speed verdict. This comparison uses the observed behavior and failure modes to explain the constraints of using each tool.

This is a comparison of ego-browser (ego lite) 0.4.7.4 and Aside CLI 1.26.902.1732 on real site missions. Environment: macOS 15.7.9 arm64, Chromium 150.

The structures differ

First, consider the workspace and control model. ego separates the agent's tabs from the user's; Aside operates the user's real tabs.

ego-browser Aside
Entry point ego-browser nodejs <<'EOF' aside repl / aside exec / 3 MCP tools
Isolation task space — separate from user windows, inherits login only none. a tab from openTab is a user tab
Ownership agent/agentDelegatedToUser/user + handoff protocol no unified ownership or handoff model
State script locals redeclared per call; browser state persists in the task space REPL scope persists (watch for name collisions)
Page JS js(string) = CDP Runtime.evaluate Playwright API (partial)
Raw access cdp(method, params) none
Runtime limits Node APIs available sandboxed — no import/require
Site knowledge site-skills runtime (with wiring problems) 51 built-in skills

If the user is browsing alongside the agent, that one row is effectively the whole decision. ego won't touch their windows. Aside can't avoid it.

Observation APIs

How each one "sees" a page differs too.

ego    snapshotText()  → indented tree + [ref=N, loc=..., url=...]
                          refs valid only in the latest snapshot (CDP backendNodeId)

Aside  snapshot(page)  → { tree, diff }
                          refs look like e1, resolved lazily via page.locator('e1')
                          annotatedScreenshot() overlays refs onto the image

Aside's diff cuts return volume, but conditionally.

Change type tree diff Savings
Page navigation 56,069 chars 56,069 chars 0%
In-page change 50,818 chars 5,529 chars 89.1%

That 89.1% assumes "full snapshot after every action." If the same judgment can be made with one targeted query, the return volume drops further. I didn't put targeted queries and diff head to head on identical tasks, so I'm not claiming a savings multiple or a general advantage.

ego's defects

click() doesn't check actionability

Measured with probes on a locally controlled page.

Condition ego click(sel) Aside page.click(sel)
Offscreen element returns success · no click · no scroll auto-scrolls, then clicks
0×0 element (inline) returns success · no click Timeout waiting for element to be ready
Normal onscreen element fine, isTrusted: true fine
Disabled onscreen element passes the check, fires an isTrusted: false event rejected

click(selector) sends input to a computed viewport coordinate without checking the resolved element's actionability. Offscreen, nothing happens. If another element occupies the point, it clicks that element and returns success.

elementFromPoint before click: overlay
click result log: [{"id":"overlay","isTrusted":true}]

"0×0 means it won't click" is an overgeneralization. A position: fixed 0×0 element remains the hit target for elementFromPoint and clicks normally. Conversely, cover the center of a visible element with an overlay and it clicks the overlay and returns success. That case is worse: a no-op is noticeable, but an unrelated control firing a side effect is invisible to the caller.

The root cause is that elementCenter('#btn') returns {x:124.4, y:3065.5} without scrolling, against a viewport height of 977. waitForElement() also returns true for offscreen elements. The coordinate math and the input primitives are there; the layer that binds them safely isn't.

This burned a lot of rounds on the Coupang cart delete button. Using it means wrapping every click with your own scroll, size, and hit-test check.

async function safeClick(sel, opts) {
  const S = JSON.stringify(sel)
  await js(`document.querySelector(${S})?.scrollIntoView({block:'center'})`)
  await wait(0.3)
  const st = await js(`(() => {
    const e = document.querySelector(${S}); if (!e) return 'no-element'
    const r = e.getBoundingClientRect()
    if (r.width === 0 || r.height === 0) return 'zero-size'
    if (r.top < 0 || r.bottom > innerHeight) return 'offscreen'
    const hit = document.elementFromPoint(r.left + r.width/2, r.top + r.height/2)
    return (hit === e || e.contains(hit)) ? 'ok' : 'occluded'
  })()`)
  if (st !== 'ok') throw new Error('click precheck failed: ' + st)
  await click(sel, opts)
}

CSS selectors only — it won't work on @N refs. And don't put this check in front of uploadFile. A hidden 0×0 <input type="file"> is a common pattern and currently works fine.

captureScreenshot() stops responding in some state

It failed five times in a row during the canvas mission and in probes right after, roughly 15 seconds each.

try1: ERR 15004ms — CDP request timed out: Page.captureScreenshot
try2: ERR 15004ms — CDP request timed out: Page.captureScreenshot
try3: ERR 15005ms — CDP request timed out: Page.captureScreenshot

Same on about:blank, on example.com, on a local page. Yet pageInfo() called immediately after in the same tab answered in under a second, and js, click, hover, snapshotText, and waitForElement were all fine. The CDP channel was alive; only this command didn't answer.

I re-measured while varying window state. After activating the app once, all five attempts succeeded in 43–70ms — foreground, background, minimized, and after clearing every task space. The failing state never reproduced. So what 0.4.7.4 established stops here. This helper fails 5/5 in some state and succeeds 5/5 in another, and the condition separating the two was not identified. Focus alone doesn't explain it; background and minimized both succeeded.

It still matters in practice. When an agent drives a browser, the window is usually occluded or behind something. snapshotText() gives back the accessibility tree, which covers UI expressed as DOM — but for what is painted rather than marked up (<canvas>, WebGL, images, layout judgments), a screenshot is the only general observation path ego exposes. In the canvas mission the first move was captureScreenshot(), and after two failures and 32.4 seconds the arm pulled pixels directly through canvas.toDataURL('image/png') instead. That escape hatch existed only because the target was a canvas. On an ordinary page there is none.

The site-skills runtime can't read its own bundled learnings

ego.helpers holds 52 helpers and SKILL.md documents a subset. elementCenter, iframeTarget, snapshotRaw, and the entire site-skills runtime are undocumented. Call only the lookup APIs and you see empty results; call runSiteTool() and the runtime tells you why.

Error: site skill not found: "github"
  searched: /skills/ego-browser/learnings
  EGO_BROWSER_AGENT_WORKSPACE: unset

The default CLI runtime's working root is /, so it searches a path that doesn't exist. Point the environment variable at the installed skill directory and loading and dispatch work — but wiring alone doesn't restore the feature. Run the bundled GitHub tools afterward and search_repos returns [] despite having results, while get_repo_stats fills only forks and returns empty strings for the rest. The bundled selectors no longer match the current DOM.

All three confirmed ego defects are already filed on the upstream issue tracker.

Aside's defects

repl --help documents a function that doesn't exist

$ aside repl --help
Notes:
  ... call openTab(url) or getTabs() first.
Examples:
  aside repl --account u1 "const tabs = await getTabs(); console.log(tabs.length)"

getTabs doesn't exist.

$ aside repl "const tabs = await getTabs(); console.log(tabs.length)"
ReferenceError: getTabs is not defined

The real name is listBrowserTabs. Copy the documented example verbatim and it fails. The bigger problem is what's missing. The global scope actually holds annotatedScreenshot, attachBrowserTab, installPageScript, snapshot, sleep, and more — while --help mentions exactly one real function (openTab) plus one that isn't there. The functions that make Aside worth using as Aside are documented nowhere.

The Playwright surface is partial

Working and non-working are mixed together.

  • Works: locator().evaluateAll(), filter({hasText}), waitForLoadState(), page.pdf(), page.close()
  • Doesn't: setViewportSize() (real browser tab, so no viewport control), page.waitForTimeout(), getTabs/listTabs

There's no way to know in advance which half you're on. In the real-data mission below, the first three rounds went to finding out.

Persistent scope and orphaned tabs

Because REPL scope persists across calls, redeclaring a const throws SyntaxError. The catch is that a failed call still consumes the name. A declaration from a call that died on an error stays registered in scope, so the retry throws SyntaxError instead of the original error.

On the tab side, ownership disappears across session boundaries. Reproduced three times.

Attempt Returned Actually
repl closing its own tab fine closed
repl attach then closeTab "no current open tabs" not closed
repl attach then page.close() "closed: …" not closed
asking exec to do it "no permission, user opened this tab" not closed

The user has to close it by hand. By contrast ego cleans up with one completeTaskSpace(id, {keep:false}), and explicitly hard-stops when the user has taken control through the GUI.

A REPL global exposes stored credentials in plaintext

Serializing one particular global object wholesale on an active profile prints stored OAuth credentials in the clear, along with billing settings, communication settings, and permission rules. The credentials.json file itself is protected -rw-------; the runtime global goes around that protection.

This surfaced while checking which image models were available — printing values instead of Object.keys(). I'm not publishing a reproduction command. Revoking an exposed refresh token is the safe move, and the vendor-side fix is either removing the credential fields or masking token patterns on serialization.

They fail differently

This is where the two tools' characters show most clearly.

ego Aside
Source of wasted rounds selectors, timing, false success its own API and syntax
Character fighting the DOM fighting its own tooling
Failure mode quiet (returns success) loud (throws)

Cold-run wasted rounds tied 6 to 6. Both sides hit selector problems; what differed is the rest. Aside's were its own API and syntax, ego's were timing and false success.

The quiet one is harder to debug. An error is something the agent reads and routes around; a false success only surfaces one or two steps later.

Can you verify cleanup?

After an agent uses a browser, cleanup remains — and the structural difference shows up again here. I left two tabs open at the end of one call, then queried from an entirely separate call.

### 1) open 2 tabs in one call, exit without closing
tabs opened: 2

### 2) query leftover tabs from a completely separate call
leftover tabs: []

aside repl "<code>" stands up a fresh browser session per call, and tabs opened via openTab vanish when the process ends. You get an empty array whether cleanup happened or not. There is no way to verify "did it clean up."

ego is the opposite. A task space survives the process. Two task spaces created during reconnaissance were still sitting there later and had to be closed by hand.

Tabs it opened User tabs
ego (task space) must be closed explicitly → the check carries information outside the task space, untouched
Aside (repl) auto-destroyed at process exit → the check carries no information leaving them is correct; closing them is an intrusion

This is a difference in how verifiable the report is, not in hygiene. Aside is structured so residue never accumulates; ego can leave residue but requires an explicit close.

Delegation didn't beat primitives

Aside has one layer ego doesn't. Instead of pushing Playwright-style JS through repl, you can hand the whole task to exec in natural language. ego has no equivalent — everything happens in the caller's turn.

I gave the same popup mission to both layers. They tied. exec satisfied every predicate in 67 seconds and returned exactly the JSON shape requested; the repl track scored the same.

They used different tools to get there.

repl track exec track
Observation page.evaluate, locator.innerText snapshot(page) → ref tree
Popup acquisition page.waitForEvent('popup') listBrowserTabs() + attachBrowserTab()
Manipulation page.locator(css) page.locator('e4') — a ref from the snapshot

The popup object the repl track caught has a different API surface from a normal Playwright Page: no waitForLoadState, no waitForTimeout, two TypeErrors along the way. The exec track never went down that path.

But snapshot, attachBrowserTab, and sleep are all present in the repl scope too.

$ aside repl "console.log(['snapshot','attachBrowserTab','sleep'].map(n => n+'='+typeof globalThis[n]).join(', '))"
snapshot=function, attachBrowserTab=function, sleep=function

They weren't unavailable — the arm didn't know they existed. The exec agent starts by reading built-in skill documentation; a repl user never receives it. On the same engine, the delegation layer gets the docs and the primitive layer doesn't. That's an asymmetry in discoverability, not capability, and it's cheap to fix.

The model conditions weren't held constant, though. exec ran with --effort medium while the primitive track used the default model. There's no guarantee the two tracks ran at the same model tier, so whether exec diagnosed its errors instantly because of structure or because of the model isn't something this experiment can settle.

One more thing to know going in: when Aside's model setting is provider: claude-code, that agent's inference comes out of the user's Claude subscription. The side-car layers handling OCR, summarization, and memory consolidation ride the same path. The shipping default is different, so it can be set back.

Capability missions — mostly ties

With the efficiency comparison unresolved (below), the question became "can it do this at all." Scoring moved to success/failure and block codes only. Four more missions.

Mission What it tests ego Aside
Separate window operate a real popup, reflect the result in the parent 5 5
Canvas coordinates hit a painted target with no DOM node, first try 5 5
Real user data mutation archive 3 specified conversations, then restore them 5 2
Remote execution drive the locally logged-in browser from an SSH session works works

Grades run from 5 (completed, cleaned up, zero intervention) down to 0 (never arrived), with a separate F for false success.

All three axes I expected to separate them turned out not to be axes. I thought the popup would expose the ownership difference; both opened it, searched, selected an item, and confirmed the value propagated back into the parent window's input. I thought ego couldn't attach from an SSH session, since it opens no TCP port and uses a unix socket under $TMPDIR — but on macOS TMPDIR is per user, not per session, and both CLIs worked from an ssh localhost session. And both hit the canvas target with a single click, each converting between canvas-local and viewport coordinates correctly.

The real-data mission is the one that split

Archiving three ChatGPT conversations and putting them back. It's the only mission that mutates real user data, so the targets were fixed at three with their IDs written into the spec.

ego Aside
Archived 3 3
Restored 3 0
Non-target conversations touched none (evidence submitted) none (evidence submitted)
Decision rounds 26 (over the cap of 20) 20 (hit the cap)

The Aside arm terminated with three conversations still archived, and wrote plainly that manual restoration was required.

Three things got in the way. Clicking the conversation options button through page.click closed the menu instantly and activated the conversation link instead; opening it required dispatching a synthetic pointerdown. With no page.focus, waitForTimeout, or keyboard, the first three rounds went to API exploration. And decisively, the arm looked at the archived-items dialog, concluded "the rows carry no conversation ID," and spent rounds on a workaround to recover titles. Querying the same dialog again, the IDs were right there as a[href^="/c/"].

That last one is an observation failure, not a tool limit. The first two are real Aside defects; the third is the arm's own misdiagnosis. "This task is impossible with Aside" is not a conclusion this data supports.

Some blame is the mission designer's. The round cap and the restore obligation collide head-on, and the spec never said which wins. ego broke the cap and confirmed the restore; Aside honored the cap and left the data mutated. Both followed instructions, so that grade gap should be discounted.

One trap caught both equally. Because the archived list updates with a delay, both arms judged at round 20 that the conversations were still archived. They had already been unarchived. Pressing all three in a row registers only one; the calls have to be split.

Efficiency came back inconclusive

Each tool ran the Coupang mission twice, four runs total. All four succeeded, zero blocks, zero captchas. Then the pre-registered order-effect test made the comparison unusable.

Run Position Rounds Wasted
ego cold first 25 6
ego warm second 8 0
Aside cold second 18 6
Aside warm first 12 1

Because the schedule used ABBA alternation, each arm's warm run landed opposite its cold run, tangling position effects with learning effects. The within-arm position delta reached 17 rounds for ego against a baseline of 2 — INCONCLUSIVE. With N=1 per cell they can't be separated; resolving it needs 8 warm runs and there are 2. So no claims about rounds, time, or bytes.

Three measurement traps worth recording. Don't trust the tool's own timing: one arm reported duration_ms: 0 on 13 of 18 rounds, producing a "6.7x faster" figure that reversed once wall-clock measurement was forced. Don't compare total bytes: a single failed round (querySelectorAll('*')) returned 1,082,464 bytes, 99% of the total, and excluding it flips the direction. Judge blocking as a time series: sampling once after domcontentloaded misses Akamai's delayed block, and URL plus title doesn't work either, since Coupang sometimes returns a normal title while blocking.

What to pick

Situation Pick Why
Default ego-browser + a separate subagent clear isolation and ownership model
User is browsing at the same time ego won't touch user tabs
Human intervention mid-task ego handoff is defined as a protocol
Need external npm modules ego the Aside REPL is sandboxed
Raw CDP ego cdp()
Work that must roll back real user data ego short restore path, verifiable cleanup
Clicking offscreen elements ego with a mandatory wrapper, or Aside ego fails silently
Observing painted output (canvas, images, layout) Aside ego's captureScreenshot depends on state
Slack, Gmail, Notion chores Aside many go through APIs without a tab

If I had to fix one default, it's ego-browser with a separate subagent. The isolation and ownership model is clear, and on this account and these missions I did not observe Aside's exec delegation adding anything over a subagent.

But choosing ego means writing your own click() wrapper before you start. Skip that and you carry silent misfires. By raw defect weight ego is the heavier of the two; it's still the default because those defects can be worked around, while an ownership model can't be bolted on later.

Limits of this record

  • Efficiency. The order-effect test didn't pass. Six more warm runs are needed
  • Sample size. The capability missions ran once each. No repetition
  • Reproducibility. The target site's DOM has A/B tests and responsive branches; initial load bytes differed per browser at the same URL and moment (222,927 / 252,588)
  • Attribution on the real-data mission. Aside's failure was an observation failure colliding with a round cap. Reading it as a tool limit is wrong