Experimental record / run 007 / experiment VE-004 / rev A

Does more page context help Jev?

[ FIRST-PARTY ][ PRE-REGISTERED ][ FOLLOW-UP TO RUN 006 ]

RUN 006 found Jev picked the right element on 38.4% of Mind2Web’s web-task steps. A public reply argued the setup was too thin: a short text cue per element, run offline, where a real agent gives richer context. This record tests the part an offline run can test. Same 9,378 steps, four arms: 50 or 10 candidates, each with or without where it sits on the page. The design and how it would be read were committed first.

PAGE CONTEXT, 50 CANDIDATES
+3.9 pp
BEST ARM, RIGHT ELEMENT
42.3%
BEST ARM, TASKS FULLY RIGHT
63 / 1,341
SAME CHOICE AS RUN 006
92.7%
The pre-registered headline
Adding page context changed Jev’s element accuracy by +3.9 pp with 50 candidates (95% interval +3.1 pp to +4.8 pp) and +3.4 pp with 10 (+2.7 pp to +4.2 pp). By the rule fixed in advance, the claim that richer context helps is supported. The best arm still picked the right element on 42.3% of steps and completed 4.7% of tasks (3.7% to 6.0%).
01 / Four arms, same steps

Context helps with either candidate count. Cutting the list to 10 is a trade: where the right element survives the cut, Jev picks it more often (rings), but the ranker’s top 10 misses it on 32.7% of steps, and those count as wrong (dots).

ALL STEPS (95% INTERVAL)STEPS WHERE THE RIGHT ELEMENT IS IN THE TOP 1030%35%40%45%50%55%60%RIGHT ELEMENT (AXIS STARTS AT 30%)50 CANDIDATESArm A, 50 candidates, description only: 38.4% of all steps (37.2%–39.6%); 50.0% where the right element is in the top 10A · description only38.4%50.0%Arm C, 50 candidates, + page context: 42.3% of all steps (41.0%–43.6%); 54.7% where the right element is in the top 10C · + page context42.3%54.7%10 CANDIDATESArm B, 10 candidates, description only: 35.3% of all steps (34.2%–36.5%); 52.5% where the right element is in the top 10B · description only35.3%52.5%Arm D, 10 candidates, + page context: 38.8% of all steps (37.5%–40.0%); 57.6% where the right element is in the top 10D · + page context38.8%57.6%
ARMCANDIDATESCONTEXTRIGHT ELEMENT95% (TASK BOOTSTRAP)IN TOP 10 ONLYRIGHT STEPTASKS FULLY RIGHTINPUT TOKENSCOST / 1K STEPS
A50no38.4%37.2% to 39.6%50.0%36.9%52 (3.9%)1,550$0.065
B10no35.3%34.2% to 36.5%52.5%34.0%41 (3.1%)652$0.027
C50yes42.3%41.0% to 43.6%54.7%40.9%63 (4.7%)2,731$0.115
D10yes38.8%37.5% to 40.0%57.6%37.5%56 (4.2%)893$0.038

“In top 10 only” counts the 6,312 steps where the right element is among the ranker’s top 10, so every arm can get it right. No arm can exceed 86.3% on all steps, and the 10-candidate arms cannot exceed 67.3%. Context costs input tokens: arm C used 1.76× arm A’s input tokens, $0.11 per thousand steps at list price.

02 / The four comparisons
COMPARISONRIGHT ELEMENT, DIFFERENCE95% (PAIRED TASK BOOTSTRAP)ONLY NEW ARM RIGHTONLY BASELINE RIGHTIN TOP 10 ONLYREADING
Add page context, 50 candidates (C vs A)+3.9 pp+3.1 pp to +4.8 pp921555+4.6 pphelps
Add page context, 10 candidates (D vs B)+3.4 pp+2.7 pp to +4.2 pp764442+5.1 pphelps
10 instead of 50, no context (B vs A)−3.0 pp−3.8 pp to −2.2 pp532817+2.5 pphurts
10 instead of 50, with context (D vs C)−3.5 pp−4.4 pp to −2.7 pp575904+2.9 pphurts

A comparison reads “helps” or “hurts” only when its interval excludes zero, as the protocol fixed. Context does not only add right answers: with 50 candidates it fixed 921 steps and broke 555.

03 / What changed from RUN 006
  • The answers repeat. Arm A sent RUN 006’s element question again, hours later. Jev chose the same element on 92.7% of steps, and accuracy was 38.4% then and 38.4% now. Only 261 steps flipped between right and wrong; adding context flipped 1,476. The context effect is not run-to-run noise.
  • Inferring the action fixes RUN 006’s design flaw. RUN 006 asked for the action type separately, without the element. Reading it off the chosen element instead gave the right action on 96.2% of right elements in arm A, lifting step accuracy from 27.3% to 36.9% and whole tasks from 29 to 52, while element accuracy stayed at 38.4%.
  • Confidence still flags errors only moderately. Error AUROC was 0.735 in arm A and 0.730 with context.
04 / What this does and does not answer

The critique was right that the cue was thin: a short description of where each element sits moved accuracy by three to four points, well outside noise. But even with context, Jev still chose the wrong element on 57.7% of steps, and a task needs every step right.

Two parts of “real use” cannot be tested here. Jev accepts text only, so screenshots are out. And Mind2Web is recorded: a wrong choice leads nowhere, so retrying and self-correction need a live environment. The task-success figures are for an agent that never retries, not for one that can.

05 / Method

Pre-registered before any request. RUN 006's steps and element descriptions, sent under four interleaved arms: the ranker's top 50 or top 10 candidates, each described alone or followed by its page region (nearby text and enclosing sections). One Choice question per request over the candidates in page order plus none_of_the_above; state is the task and the last five demonstrated actions. The action type is inferred from the chosen element's tag. Comparisons are paired, with intervals from 10,000 resamples of whole tasks. 37,512 requests, $2.30 at list price, all resolved to jev-1.13.0. Example option with context: [input] | search | Search by city... || in: Change Location | within: combobox.

06 / Limitations
  • Each step starts from the demonstrated history on a recorded page. A live agent that makes a mistake sees different pages afterwards and may retry; this record cannot say how Jev recovers from its own errors.
  • Page context is a short region per element in JevBench's own format — nearby text and enclosing section names — not the page HTML or a screenshot. Jev accepts text only, so screenshots could not be tested with it.
  • Candidates come from the Mind2Web authors' ranker, which lists the right element among its top 50 on 86.3% of steps and its top 10 on 67.3%; the rest are counted wrong, as in Mind2Web's own scoring.
  • The action type is inferred from the chosen element by a fixed rule, which gives the right action for the right element on 95.8% of steps. Typed text and selected option values are not evaluated.
  • The context arms' instructions add one sentence explaining the context format, so each context comparison measures the context and that sentence together.
  • Latency is client-observed round-trip time from a cloud runner whose region was not recorded; all four arms were interleaved, so they are comparable with each other.
  • Cost is derived from recorded input tokens at list price, not reconciled to an invoice.
07 / What this run does not prove
  • It does not measure any particular agent product, including Browser Use's Jev integration, whose prompts, page representation and recovery logic differ.
  • It does not show what richer context than this — full HTML, accessibility trees, several rounds, or retries — would achieve.
  • It does not show how Jev compares with any other model; no other model was run under this setup.
  • It is not an independent result: JevBench chose the setup, wrote the protocol, ran Jev and wrote the analysis.
08 / Data and licence

Mind2Web: Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H. and Su, Y. (2023), Mind2Web: Towards a Generalist Agent for the Web, NeurIPS 2023. Dataset licensed CC BY 4.0. The example option above is from the test split and is shown only to illustrate the format.

Why the test data is not here. Mind2Web’s authors ask that the unzipped test split not be redistributed, so JevBench publishes only step ids, arms, Jev’s choices, confidences and scores. The frozen inputs can be rebuilt byte for byte from the official downloads with the published script; their hashes are in the manifest.