Skip to main content
The distinctive thing about building context through an interview is that the measurement comes for free. Every disambiguation the analyst makes (“active client means X, not Y”) is at once a context entry and a labeled eval pair. Building context is harvesting ground truth.

The model evaluates itself

During the interview, the model checks its own work against your source of truth: “I get this number for revenue last month — can you enter the value from your dashboard so I can validate myself?” When the numbers match, that answer is saved as a verified eval pair. This is build-measure-learn applied to context: slower on the initial build, but you see the value of each incremental piece of context with real questions, in real time.

Live verification against your dashboard

Each domain closes with a verify step. The agent answers a handful of the domain’s questions against your live warehouse twice — context off and context on — and then confirms the context-on answer against a dashboard your team already trusts.
  • A match becomes the strongest kind of eval seed: a real number, pinned to a date, with the SQL that produced it. The blessed SQL lands in a gitignored evals/verified/ sidecar next to the seed.
  • A mismatch is harvested back into the context as a caveat and a correction seed — the workflow reports the discrepancy and its cause; it never silently changes a business definition.
  • The off → on → truth delta is printed so the analyst sees the payoff before committing to the next domain.
Here it is end to end against a real Sigma workbook: Screen recording at 2560 px — use the player’s fullscreen control, or open it in its own tab ↗ to read the terminal and dashboard text. The agent is asked for Finance AOV and refund rate for Shorelane Commerce and to reconcile them against the live Shorelane – Business workbook. It opens the workbook in a dedicated local Chrome profile, lands on the Sigma login screen, and hands control to the user — who signs in and completes two-factor auth themselves. Once the page is ready, the agent reads the active controls (Last 12 Months) and resolves that label to exact calendar bounds, runs fresh read-only BigQuery queries for both metrics, captures the dashboard independently at its display precision, and reports the match: 3,824.21∗∗againstatileshowing∗∗3,824.21** against a tile showing **3.8k, and 0.3028531% against 0.30% — both within display rounding. The filters, capture, and reconciliation report are saved, and the context repo now holds a validated seed for future regression testing.
  1. Live dashboard verification for Shorelane Commerce. The agent is asked to calculate Finance AOV and refund rate, then reconcile them against the live Sigma dashboard.
  2. It needs the authenticated Sigma workbook, Shorelane – Business, so it opens it in a dedicated local Chrome profile. The page lands on the Sigma login screen.
  3. The agent asks the user to sign in manually. It never handles or inspects credentials. Nothing leaves the local machine and nothing is shared with Nodal. Sign-in — and two-factor authentication — happen in the visible browser, on the user’s own computer.
  4. The authenticated workbook loads on Page 1. The user sends “ready”.
  5. The agent reads the active dashboard controls: the reporting period is Last 12 Months. That label resolves to exact calendar bounds — September 1, 2025 through August 31, 2026 — across all channels, in USD.
  6. Finance AOV: GMV divided by every order placed in the window. Finance refund rate: refunds issued in the window divided by GMV from orders placed in the same window.
  7. Both values come from fresh, read-only BigQuery queries. Saved snapshot values are never reused as answers.
  8. The dashboard is then captured independently through Sigma, preserving its actual display precision. AOV: 3,824.21fromBigQuery,displayedas3,824.21 from BigQuery, displayed as 3.8k in Sigma. Refund rate: 0.3028531% from BigQuery, displayed as 0.30% in Sigma. Both match within the dashboard’s display-rounding precision.
  9. The active filters, dashboard capture, and reconciliation report are saved. If a value disagreed, the workflow would report the mismatch and its cause without silently changing the business definition.
  10. In your new context repo, you now have a validated evaluation seed for future regression testing.

How the dashboard gets read

Automated reading is optional and consent-gated. When you say yes, setup adds a local browser binding (Chrome DevTools MCP — a bridge that lets the agent navigate a visible Chrome window, not the developer panel) with a dedicated profile so the session stays separate from your normal browsing. You sign in yourself; Nodal never types, stores, requests, or even looks at a credential, and never asks for a tokenized share link. Without the binding, the flow is identical — the analyst reads the value and the active filters off the dashboard and the interview uses that. Your choice is saved in .nodal.local.json as one of ask_when_needed, manual, or automated, so setup doesn’t keep asking. Consent text, the restart requirement, and troubleshooting live in the open-source dashboard verification guide.

Filters first, then values

The single most common “mismatch” between a dashboard and a query is two correct numbers for two different windows. So the capture records the active filters, date grain, and window anchor alongside every value, and the reconciliation compares windows before values. A window mismatch is reported in plain language — “the dashboard shows fully-elapsed months through Jul 31; the query summed through Aug 12” — not as a data discrepancy. Values are read from data, not pixels, in a fixed order — network payload → embedded chart data → DOM → per-widget export → vision on a screenshot as a last resort — and every value carries the extraction tier it came from. A tile showing $5.9M is recorded at display precision; the reconciliation never manufactures decimals. Playbooks ship for Plotly and Sigma, and the skill writes a replay playbook per dashboard so later runs are deterministic.
The public Shorelane reconciliation report is what one run produces: five exact matches, one window mismatch explained to the dollar, and one value not verifiable from that dashboard. On the ambiguous question “what was revenue last month?”, context-off chose net_revenue (21.6% above the dashboard); context-on chose recognized_revenue and matched it exactly.

The local eval delta

Beyond the in-session check, the same questions can be re-run any time — context on and off — across a whole domain:
The open-source eval_harness/ reports the accuracy difference — a concrete number you can show.
The delta is honest in both directions. In the demo, context made one revenue question better and another worse — a reference file routed every revenue question to a table with no channel dimension. Catching that during the interview, while the analyst is in the room to fix it, is the point.

Format-agnostic

The harness reads ACF, Kaelio KTX, dbt models and docs, raw markdown, or an agent data-analysis skill, normalizes them all into one representation, and measures the delta the same way — so you can evaluate context you already have, not just ACF. Non-ACF context arrives with no labeled seeds, which is exactly why the interview is worth running even then: it’s what mints the ground truth any format can then be graded against.

From one-shot to continuous

The one-shot, run-locally eval delta and dashboard reconciliation are free. Continuous re-evaluation, drift detection pinned to the commit that caused it, and observability across a team are the hosted product — see enterprise evaluation.