The model evaluates itself
During the interview, the model checks its own work against your source of truth: “I get this number for revenue last month — can you enter the value from your dashboard so I can validate myself?” When the numbers match, that answer is saved as a verified eval pair. This is build-measure-learn applied to context: slower on the initial build, but you see the value of each incremental piece of context with real questions, in real time.Live verification against your dashboard
Each domain closes with a verify step. The agent answers a handful of the domain’s questions against your live warehouse twice — context off and context on — and then confirms the context-on answer against a dashboard your team already trusts.- A match becomes the strongest kind of eval seed: a real number, pinned to a date, with the SQL
that produced it. The blessed SQL lands in a gitignored
evals/verified/sidecar next to the seed. - A mismatch is harvested back into the context as a caveat and a
correctionseed — the workflow reports the discrepancy and its cause; it never silently changes a business definition. - The off → on → truth delta is printed so the analyst sees the payoff before committing to the next domain.
Transcript
Transcript
- Live dashboard verification for Shorelane Commerce. The agent is asked to calculate Finance AOV and refund rate, then reconcile them against the live Sigma dashboard.
- It needs the authenticated Sigma workbook, Shorelane – Business, so it opens it in a dedicated local Chrome profile. The page lands on the Sigma login screen.
- The agent asks the user to sign in manually. It never handles or inspects credentials. Nothing leaves the local machine and nothing is shared with Nodal. Sign-in — and two-factor authentication — happen in the visible browser, on the user’s own computer.
- The authenticated workbook loads on Page 1. The user sends “ready”.
- The agent reads the active dashboard controls: the reporting period is Last 12 Months. That label resolves to exact calendar bounds — September 1, 2025 through August 31, 2026 — across all channels, in USD.
- Finance AOV: GMV divided by every order placed in the window. Finance refund rate: refunds issued in the window divided by GMV from orders placed in the same window.
- Both values come from fresh, read-only BigQuery queries. Saved snapshot values are never reused as answers.
- The dashboard is then captured independently through Sigma, preserving its actual display precision. AOV: 3.8k in Sigma. Refund rate: 0.3028531% from BigQuery, displayed as 0.30% in Sigma. Both match within the dashboard’s display-rounding precision.
- The active filters, dashboard capture, and reconciliation report are saved. If a value disagreed, the workflow would report the mismatch and its cause without silently changing the business definition.
- In your new context repo, you now have a validated evaluation seed for future regression testing.
How the dashboard gets read
Automated reading is optional and consent-gated. When you say yes, setup adds a local browser binding (Chrome DevTools MCP — a bridge that lets the agent navigate a visible Chrome window, not the developer panel) with a dedicated profile so the session stays separate from your normal browsing. You sign in yourself; Nodal never types, stores, requests, or even looks at a credential, and never asks for a tokenized share link. Without the binding, the flow is identical — the analyst reads the value and the active filters off the dashboard and the interview uses that. Your choice is saved in.nodal.local.json as one
of ask_when_needed, manual, or automated, so setup doesn’t keep asking. Consent text, the
restart requirement, and troubleshooting live in the open-source
dashboard verification guide.
Filters first, then values
The single most common “mismatch” between a dashboard and a query is two correct numbers for two different windows. So the capture records the active filters, date grain, and window anchor alongside every value, and the reconciliation compares windows before values. A window mismatch is reported in plain language — “the dashboard shows fully-elapsed months through Jul 31; the query summed through Aug 12” — not as a data discrepancy. Values are read from data, not pixels, in a fixed order — network payload → embedded chart data → DOM → per-widget export → vision on a screenshot as a last resort — and every value carries the extraction tier it came from. A tile showing$5.9M is recorded at display precision; the
reconciliation never manufactures decimals. Playbooks ship for Plotly and Sigma, and the
skill writes a replay playbook per dashboard so later runs are deterministic.
The local eval delta
Beyond the in-session check, the same questions can be re-run any time — context on and off — across a whole domain:eval_harness/ reports the accuracy difference — a concrete number you can show.
The delta is honest in both directions. In the demo, context made
one revenue question better and another worse — a reference file routed every revenue question
to a table with no channel dimension. Catching that during the interview, while the analyst is in
the room to fix it, is the point.
