ego-browser and Aside Fail Differently
To let an AI agent manage a shopping cart or organize a conversation list, it needs to read a logged-in page, operate its controls, and check what actually changed. Browser automation tools provide that connection. The tool determines which tabs the agent operates, how the user can intervene, and how failures and changes can be checked and reversed.
ego-browser and Aside both let an agent observe and operate a browser, but they organize the work differently. ego-browser puts the agent's tabs in a task space separate from user windows; Aside operates actual tabs in the user's browser. Alongside individual operations, Aside also offers exec, which accepts a whole task in natural language. They can handle the same tasks while differing in how the agent shares the browser with the user and how much the caller manages directly.
Comparing them helps establish what those differences mean in practice. Can the user keep browsing while the agent works? Can the caller detect a failed operation? Can changed data and tabs be put back in order afterward? I examined these questions through a Coupang shopping cart flow, popup and canvas tasks, conversation archiving and restoration, and behavior probes on controlled pages.
One difference emerged in how failures become visible. ego-browser returned success after missing an offscreen element or clicking another element. With Aside, TypeError and SyntaxError interrupted the work. Both wasted six rounds on cold runs, but finding and correcting the problems involved different work. Order effects prevented a speed verdict. This comparison uses the observed behavior and failure modes to explain the constraints of using each tool.
This is a comparison of ego-browser (ego lite) 0.4.7.4 and Aside CLI 1.26.902.1732 on real site missions. Environment: macOS 15.7.9 arm64, Chromium 150.
The structures differ
First, consider the workspace and control model. ego separates the agent's tabs from the user's; Aside operates the user's real tabs.
| ego-browser | Aside | |
|---|---|---|
| Entry point | ego-browser nodejs <<'EOF' |
aside repl / aside exec / 3 MCP tools |
| Isolation | task space — separate from user windows, inherits login only | none. a tab from openTab is a user tab |
| Ownership | agent/agentDelegatedToUser/user + handoff protocol |
no unified ownership or handoff model |
| State | script locals redeclared per call; browser state persists in the task space | REPL scope persists (watch for name collisions) |
| Page JS | js(string) = CDP Runtime.evaluate |
Playwright API (partial) |
| Raw access | cdp(method, params) |
none |
| Runtime limits | Node APIs available | sandboxed — no import/require |
| Site knowledge | site-skills runtime (with wiring problems) | 51 built-in skills |
If the user is browsing alongside the agent, that one row is effectively the whole decision. ego won't touch their windows. Aside can't avoid it.
Observation APIs
How each one "sees" a page differs too.
ego snapshotText() → indented tree + [ref=N, loc=..., url=...]
refs valid only in the latest snapshot (CDP backendNodeId)
Aside snapshot(page) → { tree, diff }
refs look like e1, resolved lazily via page.locator('e1')
annotatedScreenshot() overlays refs onto the imageAside's diff cuts return volume, but conditionally.
| Change type | tree | diff | Savings |
|---|---|---|---|
| Page navigation | 56,069 chars | 56,069 chars | 0% |
| In-page change | 50,818 chars | 5,529 chars | 89.1% |
That 89.1% assumes "full snapshot after every action." If the same judgment can be made with one targeted query, the return volume drops further. I didn't put targeted queries and diff head to head on identical tasks, so I'm not claiming a savings multiple or a general advantage.
ego's defects
click() doesn't check actionability
Measured with probes on a locally controlled page.
| Condition | ego click(sel) |
Aside page.click(sel) |
|---|---|---|
| Offscreen element | returns success · no click · no scroll | auto-scrolls, then clicks |
| 0×0 element (inline) | returns success · no click | Timeout waiting for element to be ready |
| Normal onscreen element | fine, isTrusted: true |
fine |
| Disabled onscreen element | passes the check, fires an isTrusted: false event |
rejected |
click(selector) sends input to a computed viewport coordinate without checking the resolved element's actionability. Offscreen, nothing happens. If another element occupies the point, it clicks that element and returns success.
elementFromPoint before click: overlay
click result log: [{"id":"overlay","isTrusted":true}]"0×0 means it won't click" is an overgeneralization. A position: fixed 0×0 element remains the hit target for elementFromPoint and clicks normally. Conversely, cover the center of a visible element with an overlay and it clicks the overlay and returns success. That case is worse: a no-op is noticeable, but an unrelated control firing a side effect is invisible to the caller.
The root cause is that elementCenter('#btn') returns {x:124.4, y:3065.5} without scrolling, against a viewport height of 977. waitForElement() also returns true for offscreen elements. The coordinate math and the input primitives are there; the layer that binds them safely isn't.
This burned a lot of rounds on the Coupang cart delete button. Using it means wrapping every click with your own scroll, size, and hit-test check.
async function safeClick(sel, opts) {
const S = JSON.stringify(sel)
await js(`document.querySelector(${S})?.scrollIntoView({block:'center'})`)
await wait(0.3)
const st = await js(`(() => {
const e = document.querySelector(${S}); if (!e) return 'no-element'
const r = e.getBoundingClientRect()
if (r.width === 0 || r.height === 0) return 'zero-size'
if (r.top < 0 || r.bottom > innerHeight) return 'offscreen'
const hit = document.elementFromPoint(r.left + r.width/2, r.top + r.height/2)
return (hit === e || e.contains(hit)) ? 'ok' : 'occluded'
})()`)
if (st !== 'ok') throw new Error('click precheck failed: ' + st)
await click(sel, opts)
}CSS selectors only — it won't work on @N refs. And don't put this check in front of uploadFile. A hidden 0×0 <input type="file"> is a common pattern and currently works fine.
captureScreenshot() stops responding in some state
It failed five times in a row during the canvas mission and in probes right after, roughly 15 seconds each.
try1: ERR 15004ms — CDP request timed out: Page.captureScreenshot
try2: ERR 15004ms — CDP request timed out: Page.captureScreenshot
try3: ERR 15005ms — CDP request timed out: Page.captureScreenshotSame on about:blank, on example.com, on a local page. Yet pageInfo() called immediately after in the same tab answered in under a second, and js, click, hover, snapshotText, and waitForElement were all fine. The CDP channel was alive; only this command didn't answer.
I re-measured while varying window state. After activating the app once, all five attempts succeeded in 43–70ms — foreground, background, minimized, and after clearing every task space. The failing state never reproduced. So what 0.4.7.4 established stops here. This helper fails 5/5 in some state and succeeds 5/5 in another, and the condition separating the two was not identified. Focus alone doesn't explain it; background and minimized both succeeded.
It still matters in practice. When an agent drives a browser, the window is usually occluded or behind something. snapshotText() gives back the accessibility tree, which covers UI expressed as DOM — but for what is painted rather than marked up (<canvas>, WebGL, images, layout judgments), a screenshot is the only general observation path ego exposes. In the canvas mission the first move was captureScreenshot(), and after two failures and 32.4 seconds the arm pulled pixels directly through canvas.toDataURL('image/png') instead. That escape hatch existed only because the target was a canvas. On an ordinary page there is none.
The site-skills runtime can't read its own bundled learnings
ego.helpers holds 52 helpers and SKILL.md documents a subset. elementCenter, iframeTarget, snapshotRaw, and the entire site-skills runtime are undocumented. Call only the lookup APIs and you see empty results; call runSiteTool() and the runtime tells you why.
Error: site skill not found: "github"
searched: /skills/ego-browser/learnings
EGO_BROWSER_AGENT_WORKSPACE: unsetThe default CLI runtime's working root is /, so it searches a path that doesn't exist. Point the environment variable at the installed skill directory and loading and dispatch work — but wiring alone doesn't restore the feature. Run the bundled GitHub tools afterward and search_repos returns [] despite having results, while get_repo_stats fills only forks and returns empty strings for the rest. The bundled selectors no longer match the current DOM.
All three confirmed ego defects are already filed on the upstream issue tracker.
Aside's defects
repl --help documents a function that doesn't exist
$ aside repl --help
Notes:
... call openTab(url) or getTabs() first.
Examples:
aside repl --account u1 "const tabs = await getTabs(); console.log(tabs.length)"getTabs doesn't exist.
$ aside repl "const tabs = await getTabs(); console.log(tabs.length)"
ReferenceError: getTabs is not definedThe real name is listBrowserTabs. Copy the documented example verbatim and it fails. The bigger problem is what's missing. The global scope actually holds annotatedScreenshot, attachBrowserTab, installPageScript, snapshot, sleep, and more — while --help mentions exactly one real function (openTab) plus one that isn't there. The functions that make Aside worth using as Aside are documented nowhere.
The Playwright surface is partial
Working and non-working are mixed together.
- Works:
locator().evaluateAll(),filter({hasText}),waitForLoadState(),page.pdf(),page.close() - Doesn't:
setViewportSize()(real browser tab, so no viewport control),page.waitForTimeout(),getTabs/listTabs
There's no way to know in advance which half you're on. In the real-data mission below, the first three rounds went to finding out.
Persistent scope and orphaned tabs
Because REPL scope persists across calls, redeclaring a const throws SyntaxError. The catch is that a failed call still consumes the name. A declaration from a call that died on an error stays registered in scope, so the retry throws SyntaxError instead of the original error.
On the tab side, ownership disappears across session boundaries. Reproduced three times.
| Attempt | Returned | Actually |
|---|---|---|
repl closing its own tab |
fine | closed |
repl attach then closeTab |
"no current open tabs" | not closed |
repl attach then page.close() |
"closed: …" | not closed |
asking exec to do it |
"no permission, user opened this tab" | not closed |
The user has to close it by hand. By contrast ego cleans up with one completeTaskSpace(id, {keep:false}), and explicitly hard-stops when the user has taken control through the GUI.
A REPL global exposes stored credentials in plaintext
Serializing one particular global object wholesale on an active profile prints stored OAuth credentials in the clear, along with billing settings, communication settings, and permission rules. The credentials.json file itself is protected -rw-------; the runtime global goes around that protection.
This surfaced while checking which image models were available — printing values instead of Object.keys(). I'm not publishing a reproduction command. Revoking an exposed refresh token is the safe move, and the vendor-side fix is either removing the credential fields or masking token patterns on serialization.
They fail differently
This is where the two tools' characters show most clearly.
| ego | Aside | |
|---|---|---|
| Source of wasted rounds | selectors, timing, false success | its own API and syntax |
| Character | fighting the DOM | fighting its own tooling |
| Failure mode | quiet (returns success) | loud (throws) |
Cold-run wasted rounds tied 6 to 6. Both sides hit selector problems; what differed is the rest. Aside's were its own API and syntax, ego's were timing and false success.
The quiet one is harder to debug. An error is something the agent reads and routes around; a false success only surfaces one or two steps later.
Can you verify cleanup?
After an agent uses a browser, cleanup remains — and the structural difference shows up again here. I left two tabs open at the end of one call, then queried from an entirely separate call.
### 1) open 2 tabs in one call, exit without closing
tabs opened: 2
### 2) query leftover tabs from a completely separate call
leftover tabs: []aside repl "<code>" stands up a fresh browser session per call, and tabs opened via openTab vanish when the process ends. You get an empty array whether cleanup happened or not. There is no way to verify "did it clean up."
ego is the opposite. A task space survives the process. Two task spaces created during reconnaissance were still sitting there later and had to be closed by hand.
| Tabs it opened | User tabs | |
|---|---|---|
| ego (task space) | must be closed explicitly → the check carries information | outside the task space, untouched |
Aside (repl) |
auto-destroyed at process exit → the check carries no information | leaving them is correct; closing them is an intrusion |
This is a difference in how verifiable the report is, not in hygiene. Aside is structured so residue never accumulates; ego can leave residue but requires an explicit close.
Delegation didn't beat primitives
Aside has one layer ego doesn't. Instead of pushing Playwright-style JS through repl, you can hand the whole task to exec in natural language. ego has no equivalent — everything happens in the caller's turn.
I gave the same popup mission to both layers. They tied. exec satisfied every predicate in 67 seconds and returned exactly the JSON shape requested; the repl track scored the same.
They used different tools to get there.
repl track |
exec track |
|
|---|---|---|
| Observation | page.evaluate, locator.innerText |
snapshot(page) → ref tree |
| Popup acquisition | page.waitForEvent('popup') |
listBrowserTabs() + attachBrowserTab() |
| Manipulation | page.locator(css) |
page.locator('e4') — a ref from the snapshot |
The popup object the repl track caught has a different API surface from a normal Playwright Page: no waitForLoadState, no waitForTimeout, two TypeErrors along the way. The exec track never went down that path.
But snapshot, attachBrowserTab, and sleep are all present in the repl scope too.
$ aside repl "console.log(['snapshot','attachBrowserTab','sleep'].map(n => n+'='+typeof globalThis[n]).join(', '))"
snapshot=function, attachBrowserTab=function, sleep=functionThey weren't unavailable — the arm didn't know they existed. The exec agent starts by reading built-in skill documentation; a repl user never receives it. On the same engine, the delegation layer gets the docs and the primitive layer doesn't. That's an asymmetry in discoverability, not capability, and it's cheap to fix.
The model conditions weren't held constant, though. exec ran with --effort medium while the primitive track used the default model. There's no guarantee the two tracks ran at the same model tier, so whether exec diagnosed its errors instantly because of structure or because of the model isn't something this experiment can settle.
One more thing to know going in: when Aside's model setting is provider: claude-code, that agent's inference comes out of the user's Claude subscription. The side-car layers handling OCR, summarization, and memory consolidation ride the same path. The shipping default is different, so it can be set back.
Capability missions — mostly ties
With the efficiency comparison unresolved (below), the question became "can it do this at all." Scoring moved to success/failure and block codes only. Four more missions.
| Mission | What it tests | ego | Aside |
|---|---|---|---|
| Separate window | operate a real popup, reflect the result in the parent | 5 | 5 |
| Canvas coordinates | hit a painted target with no DOM node, first try | 5 | 5 |
| Real user data mutation | archive 3 specified conversations, then restore them | 5 | 2 |
| Remote execution | drive the locally logged-in browser from an SSH session | works | works |
Grades run from 5 (completed, cleaned up, zero intervention) down to 0 (never arrived), with a separate F for false success.
All three axes I expected to separate them turned out not to be axes. I thought the popup would expose the ownership difference; both opened it, searched, selected an item, and confirmed the value propagated back into the parent window's input. I thought ego couldn't attach from an SSH session, since it opens no TCP port and uses a unix socket under $TMPDIR — but on macOS TMPDIR is per user, not per session, and both CLIs worked from an ssh localhost session. And both hit the canvas target with a single click, each converting between canvas-local and viewport coordinates correctly.
The real-data mission is the one that split
Archiving three ChatGPT conversations and putting them back. It's the only mission that mutates real user data, so the targets were fixed at three with their IDs written into the spec.
| ego | Aside | |
|---|---|---|
| Archived | 3 | 3 |
| Restored | 3 | 0 |
| Non-target conversations touched | none (evidence submitted) | none (evidence submitted) |
| Decision rounds | 26 (over the cap of 20) | 20 (hit the cap) |
The Aside arm terminated with three conversations still archived, and wrote plainly that manual restoration was required.
Three things got in the way. Clicking the conversation options button through page.click closed the menu instantly and activated the conversation link instead; opening it required dispatching a synthetic pointerdown. With no page.focus, waitForTimeout, or keyboard, the first three rounds went to API exploration. And decisively, the arm looked at the archived-items dialog, concluded "the rows carry no conversation ID," and spent rounds on a workaround to recover titles. Querying the same dialog again, the IDs were right there as a[href^="/c/"].
That last one is an observation failure, not a tool limit. The first two are real Aside defects; the third is the arm's own misdiagnosis. "This task is impossible with Aside" is not a conclusion this data supports.
Some blame is the mission designer's. The round cap and the restore obligation collide head-on, and the spec never said which wins. ego broke the cap and confirmed the restore; Aside honored the cap and left the data mutated. Both followed instructions, so that grade gap should be discounted.
One trap caught both equally. Because the archived list updates with a delay, both arms judged at round 20 that the conversations were still archived. They had already been unarchived. Pressing all three in a row registers only one; the calls have to be split.
Efficiency came back inconclusive
Each tool ran the Coupang mission twice, four runs total. All four succeeded, zero blocks, zero captchas. Then the pre-registered order-effect test made the comparison unusable.
| Run | Position | Rounds | Wasted |
|---|---|---|---|
| ego cold | first | 25 | 6 |
| ego warm | second | 8 | 0 |
| Aside cold | second | 18 | 6 |
| Aside warm | first | 12 | 1 |
Because the schedule used ABBA alternation, each arm's warm run landed opposite its cold run, tangling position effects with learning effects. The within-arm position delta reached 17 rounds for ego against a baseline of 2 — INCONCLUSIVE. With N=1 per cell they can't be separated; resolving it needs 8 warm runs and there are 2. So no claims about rounds, time, or bytes.
Three measurement traps worth recording. Don't trust the tool's own timing: one arm reported duration_ms: 0 on 13 of 18 rounds, producing a "6.7x faster" figure that reversed once wall-clock measurement was forced. Don't compare total bytes: a single failed round (querySelectorAll('*')) returned 1,082,464 bytes, 99% of the total, and excluding it flips the direction. Judge blocking as a time series: sampling once after domcontentloaded misses Akamai's delayed block, and URL plus title doesn't work either, since Coupang sometimes returns a normal title while blocking.
What to pick
| Situation | Pick | Why |
|---|---|---|
| Default | ego-browser + a separate subagent | clear isolation and ownership model |
| User is browsing at the same time | ego | won't touch user tabs |
| Human intervention mid-task | ego | handoff is defined as a protocol |
| Need external npm modules | ego | the Aside REPL is sandboxed |
| Raw CDP | ego | cdp() |
| Work that must roll back real user data | ego | short restore path, verifiable cleanup |
| Clicking offscreen elements | ego with a mandatory wrapper, or Aside | ego fails silently |
| Observing painted output (canvas, images, layout) | Aside | ego's captureScreenshot depends on state |
| Slack, Gmail, Notion chores | Aside | many go through APIs without a tab |
If I had to fix one default, it's ego-browser with a separate subagent. The isolation and ownership model is clear, and on this account and these missions I did not observe Aside's exec delegation adding anything over a subagent.
But choosing ego means writing your own click() wrapper before you start. Skip that and you carry silent misfires. By raw defect weight ego is the heavier of the two; it's still the default because those defects can be worked around, while an ownership model can't be bolted on later.
Limits of this record
- Efficiency. The order-effect test didn't pass. Six more warm runs are needed
- Sample size. The capability missions ran once each. No repetition
- Reproducibility. The target site's DOM has A/B tests and responsive branches; initial load bytes differed per browser at the same URL and moment (222,927 / 252,588)
- Attribution on the real-data mission. Aside's failure was an observation failure colliding with a round cap. Reading it as a tool limit is wrong