DATAHUB AGENT HACKATHON · OPEN / WILDCARD
Attest

Verify what your AI claims about your data.

An auditor for AI agents that make claims about data. Verdicts are decided in deterministic code — zero decided by a model — published by a human, and written back into DataHub as durable assertions.

Explore the recorded audit Watch the demo View the code

The replay is the real interface, driven by one real audit captured off the wire — nothing to install, nothing live. It refuses any question it has no recorded answer for rather than inventing one.

01 · The failure

Nobody checks the sentence.

An agent writes “customer_profile holds no PII.” The sentence sounds like the end of a task. It is actually a claim — and the catalog can check it. The table is tagged PII. Its email column is tagged PII. The catalog said so the whole time; nobody asked it.

CONTRADICTED

“customer_profile holds no PII.”

The catalog holds what the claim denies: urn:li:tag:PII, plus glossary terms EmailAddress and PersonName. The claim is wrong, and the evidence is citable.

INSUFFICIENT-COVERAGE

“orders_export holds no PII.”

The next table has no tags at all. “We reviewed it and it’s clean” and “nobody ever looked” are different facts — but without a third verdict, they are indistinguishable. That is why Attest has three verdicts, not two.

02 · The demo

Three minutes, end to end.

If the embed doesn’t load: watch on YouTube.

03 · How it works

Deterministic in the middle, human at the gate.

01

A claim comes in

Paste agent output. Attest parses it into discrete, checkable claims about named datasets.

02

Deterministic check

Each claim is evaluated against the live DataHub catalog in plain code. No model touches the verdict.

03

Three verdicts

Every verdict cites the catalog field that decided it.

SUPPORTED CONTRADICTED INSUFFICIENT-COVERAGE
04

A human publishes

Publish and accept-correction are separate decisions. Published verdicts are written back to DataHub as durable assertions — the next agent inherits them from the catalog, not from Attest.

04 · The numbers

Receipts, not headlines.

100%
accuracy on 40 labeled claims — zero off the diagonal, across all 12 claim-type × verdict cells
67.5%
when a checker is deliberately sabotaged — asserted by a test on every CI run
The pairing is the argument: an invariant nobody can falsify is decoration.
0
verdicts decided by a model
17/17
seeded datasets where the MCP read diverges from the GraphQL one — measured, then written up
2
issues filed upstream, one with a fix proposed

06 · Engaging with the MCP server

MCP discovers. A human resolves. GraphQL verifies. Deterministic code decides.

MCP
discovery · advisory

Searches the live catalog by name and offers candidate datasets.

A human
selects the URN

Picks one. It must then appear verbatim in the agent’s own text.

GraphQL
authoritative evidence

Reads the snapshot every claim about that dataset is decided against.

Checkers
the verdict, in code

Date math, set membership, string comparison. No model decides.

The DataHub MCP Server is in the product. It backs Attest’s URN picker, so a human finds a dataset by typing part of its name against the live catalog instead of pasting a 90-character URN — replacing a static list generated from our own seed, which describes one catalog and is useless against any other. just discover runs it against yours and exits zero.

Nothing it returns is evidence, and that boundary is asserted rather than described. The only value discovery passes on is the entity URN, which must still appear verbatim in the agent’s own text before any claim can be made about it. So a wrong pick produces claims about an explicitly wrong URN — visible in the report and in the published artifact — never a resolution error laundered into catalog disagreement. No checker, snapshot, cache or pipeline node can import that module; the import graph is walked and asserted.

It does not back the verdict read, and that is a conclusion the measurement forced rather than a preference we started with. Attest’s catalog read already had the one-method seam an adapter would implement, so we built to it and ran the real server against all 17 seeded datasets. The server runs — compatibility is the failure everyone expects here and is not what happened; every call answered, for every dataset, without error. Field parity against the reference read diverges on 17 of 17 datasets (136 mismatches, mean 8.00 per dataset over 491 comparisons), and four of five true claims change verdict — including customer_profile.email is PII reading back Contradicted.

Its read tools are built to feed a language model, and each optimisation for that job removes something a deterministic checker needs. That makes this a finding about structured consumers, not a defect for the server’s intended use, where the compaction is a feature: a transport that is lossy for a language model is inverting for a checker, because the checker’s precision is exactly what makes it unable to shrug. Three of the four mechanisms behind it are fixable upstream, so we wrote them up with reproductions.

just spike-mcp exits non-zero by design: the day it goes green, the finding has expired and the decision is worth reopening. Full write-up and per-dataset diffs: docs/mcp-evaluation.md.

07 · Outside our own catalog

We ran it somewhere we did not control.

67 datasets 7 platforms 15 claims metadata we didn’t author

Every number above was measured against a catalog we wrote, with labels applying a policy we wrote. The benchmark says that about itself, and it is the honest weak point. So Attest was also run against DataHub’s own showcase-ecommerce datapack — 67 datasets over 7 platforms, metadata nobody here wrote a line of. Fifteen claims, $0.0077.

Nothing in it is scored: no accuracy, no macro-F1, no confusion matrix. The golden benchmark is a conformance gate where 100% is the expected result, and importing that apparatus onto fifteen unlabelled claims on a foreign catalog would measure nothing. The question was narrower — does Attest produce defensible verdicts against metadata it did not design, and where does it hit its own limits? It found one before a single audit ran.

15 OF 67 UNAUDITABLE — THEN CLOSED, AND RE-MEASURED

A gap our own seed could never expose.

Attest’s GraphQL query had no CorpGroup arm, so a group-owned dataset arrived with an empty owner object and every claim about it was refused. The guard failed closed, which is correct. Its diagnosis — malformed catalog response — was wrong: the response is fine, the query was incomplete, and Attest could not tell the difference. No test here could have caught it. The seed emitted corpuser owners exclusively, so the offline tier, the live tier and the 12-cell matrix were all green and the gap was invisible from inside. It is now closed, and closed seed-first: a group-owned dataset was seeded before the query arm was added, so the fix does not rest on the same seed that hid the problem. The close is measured rather than asserted — a two-arm census over one loaded catalog reads 67/67 readable with the arm and 52/67 without it, 15 recovered, 0 still refused, 0 regressed.

ONE CLAIM WAS RIGHT BY LUCK — NOW REFUSED

Detecting it cost a point, which is the correct direction.

A sentence about a column’s glossary label came back from the decomposer as a schema claim, dropping the term URN entirely. The schema checker then answered a question nobody asked — does this column exist? — and returned the verdict the trial expected. It was banked as a match and named as luck. A deterministic family guard now refuses that transcription outright, as No-Claim, before any checker sees it. So the two figures the receipt never nets have converged rather than improved: 14 of 15 outcomes matched what was predeclared, and 14 of 15 answered the question actually asked — the same fourteen. The match count fell from 15 to 14 because a lucky match became a visible refusal. That is the number getting more honest, not the system getting worse.

The catalog is a near-miss on all three PII signals Attest recognises: a glossary term named literally PII, filed under a node called Classification; contains_pii where the policy names hasPII; and no completeness marker anywhere. So a column named cust_email, carrying a term whose own description reads “Subject to PII handling requirements”, comes back Insufficient-Coverage where a person would say PII. That is a declared position — structure is a declaration, a name is a guess — meeting a real catalog. The trial does not resolve it. It shows the price, with a receipt.

What this does not prove

One small curated datapack is not “works in production”, fifteen claims are not a sample of anything, and the claims were still written by the same person who wrote the checkers — the trial moves the catalog outside our control, not the claims.

The write-up, and the receipts every figure opens in — the trial and the two-arm census. The pre-fix run is kept as the superseded baseline, unedited, so the before/after pair has a before.

08 · Try it

Runs on your machine, by design.

just setup && just check      # offline, no DataHub, no API key
just up && just seed          # DataHub Core v1.5.0.6, vendored pinned compose
just demo                     # UI and API on :8003

Needs ~8 GB free. The first bring-up pulls ~12.6 GB of images, which is network-dependent; a cold boot to a healthy GMS after that is ~4.5 minutes. Local by design — a per-machine catalog is what gives every reader a clean story.