Published on

Making "Talk to Your Data" Actually Work

Authors
Line art in Daana blue on a black background: a single glowing core sends continuous lines out to stacked storage blocks, a table grid, a bar-chart dashboard and a small conversational agent shape, every line tracing back to the same origin.

The Ask and the Quiet Failure

A client asked us recently to make "talk to your data" work. They didn't want it for their own analysts, who at least know which dashboard to distrust. They wanted it for their clients: people outside the company, typing a question into a box and expecting a number they can act on.

The demo is the easy part: point a modern model at a warehouse and it writes plausible SQL in minutes. The hard question is built on what, exactly. An agent answering for your business has to know what your business means by "customer," "active," "churned," and that meaning has to be written down somewhere for the agent to read.

Anthropic has published the most candid account of what happens next. Their internal analytics agent runs on skills, markdown documents that tell Claude how the data model works. In their own write-up, they report offline accuracy drifting from about 95% at launch to about 65% over a month, as those documents fell behind a data model that changed daily. Anthropic fixed it, and we'll get to how. Still, a team with that much engineering depth, documenting a data model it owns, lost thirty points of accuracy in a month.

The failure is quiet, and that's what makes it expensive. An agent with stale context doesn't throw an error. It answers, fluently, with a number that was right last month. Inside a company that costs an awkward Slack thread. When the person asking is your client, the wrong number is the product.

"Talk to your data" works when the platform underneath is built for it, and most of the ones we're asked to add AI to aren't. Teams stack skills, semantic layer configs and context documents over a warehouse nobody modeled, and each one is a second copy of the truth. Second copies drift. The durable fix declares meaning once, in a model, and generates every layer from it, including the layer the agent reads, so the agent's context comes from the same definitions the platform was built from and changes in the same commit.

Context Isn't Enough makes the broader case that context alone can't rescue a messy platform. The open question is what to build instead.

The Move Everyone Is Already Making

You've probably seen some version of this thread. Someone from sales pastes a screenshot into the data channel. The agent says net revenue retention last quarter was 112%. The board dashboard says 104%. Underneath, the question every data team learns to dread: "why doesn't this number match the dashboard?"

A data engineer replies that the agent is probably reading the old revenue table, deprecated back in June. Someone asks whether anyone updated the skill file. Nobody answers that one. Meanwhile the agent's original reply sits at the top of the thread with a thumbs-up from exactly one person (the one who built the agent).

What happens next is predictable: somebody opens the skill file and adds a line, use revenue_v2 instead. Next month it's a line about which status codes count as active. Some teams put this in a skill file, some in a semantic layer config, some in a long context document the agent loads before every question. The format varies, and the move is the same: write down what the agent needs to know, next to a warehouse that doesn't know it.

Each of those documents is a map, drawn over territory the mapmaker didn't build and doesn't control. The map drifts because nothing ties it to the territory: when an engineer deprecates a table, the pipeline changes in the same commit, but the skill file changes only if somebody remembers it exists. Context Isn't Enough walks through why that gap widens as the map grows, so we won't re-derive it here.

And yes, the move is reasonable. A skill file takes an afternoon; modeling the warehouse properly takes a quarter, a budget conversation, and a long meeting about who owns the word "customer." When leadership asks you to "just add AI," writing down what you already know is the fastest way to get an agent answering, and the demo goes great for the first few weeks.

The trouble shows up later, once the warehouse has moved and the map hasn't. That's the moment Anthropic hit, and their response is the strongest version of this move anyone has published. Before arguing with it, look at how far it gets.

Anthropic's Actual Fix, and Its Ceiling

Start with what Anthropic achieved, because it's a lot. According to their write-up, about 95% of their business analytics queries are now automated through Claude, at roughly 95% accuracy in aggregate. Without skills, accuracy on their evals didn't exceed 21%.

Skills are also where the drift came from, and Anthropic says so plainly:

"Skill docs describe a data model that changes daily, so without active maintenance they're wrong within weeks."

That's the slide from about 95% to about 65% over a month. Keep the two numbers apart: 95% coverage is how many queries the agent handles; the 95% to 65% drift was offline accuracy on their evals, before they fixed anything. Their own phrasing: the drop continued until they "treated this as an engineering problem." Then they did, starting with where the files live:

"colocating skill markdown files in the same repo as our transformation models, so the PR that changes a model is the same PR that updates the doc describing it."

The rest is tooling. A code-review hook flags any reporting-model change that doesn't touch a skill file (roughly 90% of their data-model PRs now include a skill change in the same diff), and a scheduled agent scans stakeholder channels for corrections and opens a PR with a one-line fix to the relevant doc.

That worked, and it's good engineering: Anthropic found the cause, put the documents where the changes happen, and automated what could be automated. If you run a skills-based agent today, copying this setup is the right next step.

Now look at what all of that machinery is for: keeping two artifacts in step, narrowing the gap between a data model and the documents describing it without closing it. The hook can tell that a skill file was touched. It can't tell whether the skill file now says the right thing.

A problem you fix by moving files between repositories is an architectural problem. Colocation works because the distance between the model and its description was causing the drift, and shrinking that distance shrinks the drift. Follow the same reasoning one step further and the distance goes to zero: the agent reads the model itself, and there's nothing left to sync.

The ceiling matters more for everyone who isn't Anthropic. They own their data model end to end, keep it in one repository, and have a data team that can build review hooks as a side project. Now picture the team asked last week to "just add AI" to a warehouse assembled over ten years from four vendors' exports and a lot of hand-written SQL: who owns the model there, and which repository would the skill files even go in? For that team the colocation fix isn't on the menu. They need the model first.

The Durable Fix: Declare Meaning Once

Take a question every subscription business gets asked weekly: how many active subscriptions do we have?

Your warehouse doesn't store "active". It stores the facts "active" is made from: when a subscription started, when it ended or is due to end, what it's worth each month, maybe a status code from the billing system that half the team trusts. "Active" is a rule applied to those facts: started, not yet ended, worth more than zero, so free trials don't count. That rule probably lives in several places at once, a CTE in the revenue model, a calculated field in the BI tool, a paragraph in a skill file, each maintained by someone different, at a different speed.

The alternative is to declare the meaning once, in a form both people and machines can read. Here's what that looks like in DMDL, the model description language we use at Daana:

entities:
  - id: "SUBSCRIPTION"
    name: "SUBSCRIPTION"
    definition: "A customer subscription"
    description: "Represents an active or historical subscription to a service plan"
    attributes:
      - id: "SUBSCRIPTION_START_DATE"
        name: "SUBSCRIPTION_START_DATE"
        definition: "When subscription activated"
        type: "START_TIMESTAMP"

      - id: "SUBSCRIPTION_END_DATE"
        name: "SUBSCRIPTION_END_DATE"
        definition: "When subscription expires or was cancelled"
        type: "END_TIMESTAMP"

      - id: "MONTHLY_VALUE"
        name: "MONTHLY_VALUE"
        definition: "Monthly subscription value with currency"
        effective_timestamp: true
        group:
          - id: "MONTHLY_VALUE"
            name: "MONTHLY_VALUE"
            definition: "The monetary amount"
            type: "NUMBER"
          - id: "MONTHLY_VALUE_CURRENCY"
            name: "MONTHLY_VALUE_CURRENCY"
            definition: "Currency code (USD, EUR, SEK)"
            type: "UNIT"

Look for the rule. The word "active" appears once, in a description; the rule itself isn't in the file. What the file holds is what a subscription is, and the facts that describe it: when it starts, when it ends, and a monthly value that's allowed to change over time (effective_timestamp: true tells the platform to keep its history, so last March's price is still there when someone asks about last March). That's deliberate: "active subscription" is a question asked of these facts, declared one layer up, built from attributes that already have a single definition, and neither the facts nor the rule is restated anywhere else.

The model says what a subscription is. It doesn't say where one comes from. That's the job of a second, equally short file, the mapping, which points each declared attribute at a column (or an expression over columns) in a source system. Here's SUBSCRIPTION mapped from a billing system's subscriptions table:

entity_id: SUBSCRIPTION

mapping_groups:
  - name: billing_subscriptions
    tables:
      - connection: billing
        table: billing.subscriptions
        primary_keys:
          - subscription_id
        ingestion_strategy: FULL
        attributes:
          - id: SUBSCRIPTION_START_DATE
            transformation_expression: activated_at
          - id: SUBSCRIPTION_END_DATE
            transformation_expression: COALESCE(cancelled_at, expires_at)
          - id: MONTHLY_VALUE
            transformation_expression: price_cents / 100.0
            attribute_effective_timestamp_expression: price_changed_at
          - id: MONTHLY_VALUE_CURRENCY
            transformation_expression: currency
            attribute_effective_timestamp_expression: price_changed_at

The billing system's quirks live here and nowhere else: prices stored in cents, an end date that's either a cancellation or an expiry. If billing renames a column, this file changes and the model doesn't. From the model and the mapping together, the generator produces the SQL that loads subscriptions and keeps their history, the documentation, and the context the agent reads.

Now count who reads that declaration: a person reviewing the model, the LLM answering a question about subscriptions, and the generator producing the SQL, the documentation and the agent's context. One definition, three readers, no step where anyone copies meaning between artifacts. When the definition changes, everything generated from it changes in the same commit, the property Anthropic's hook approximates from the outside.

DMDL is our way of declaring a model; others do it with different syntax and tools, but the essence is the same: meaning written once, in a form a machine can generate from, so documentation and implementation are one artifact.

It isn't free: deciding what a subscription is, with finance, product and whoever owns billing, is harder than writing a skill file this afternoon. For a team with one source system and no data person, it's overkill: write the skill file and get on with your week.

Generating Every Layer, Including the One the Agent Talks To

A declaration on its own is a nicely formatted YAML file. It earns its keep when the platform is generated from it, and that platform has three distinct jobs.

The first is to keep what the source systems recorded, exactly as they recorded it, quirks included: if billing sends status code 4 for "paused," this layer holds a 4. The second is to hold one agreed definition of each business concept (customer, subscription, order) and the relationships between them, built from declarations like the one above. The third is to shape data for a specific question or tool: a revenue dashboard, a finance extract, a churn feature table.

At Daana we call these layers DAS (Data As System sees it), DAB (Data As Business sees it) and DAR (Data As Requirements needs it): source, business and consumption. The separation means a change in one place stays in one place: when a source system renames a column, DAS absorbs it and the business definition of a subscription doesn't move.

The agent's context is no exception. It's data and definitions shaped for one consumer, the LLM, which makes it a DAR artifact like any dashboard. A team writing skill files by hand is hand-building that one shape while generating none of the others, and the files drift for the same reason hand-built marts do. On a declared platform the agent's context is generated from the DAB like everything else in DAR, no fourth thing bolted onto the side.

For the agent's queries, the shape that helps most is Francesco Puppini's Unified Star Schema, from his book with Bill Inmon, The Unified Star Schema: one bridge table connecting every entity to every other it can legitimately reach, with a generated list of which measures can be sliced by which dimensions. Generated from the DAB, it gives the agent one place to query, with the model's answerable questions already listed. The Consumption Layer Should Generate Itself covers the mechanism.

Doesn't the model change daily too?

That's the obvious objection, straight from Anthropic's sentence about a data model that changes daily. If the model is the single source, won't it have to change daily as well?

Mostly not: a business information model, built on the business's processes, changes slowly. The data team adds entities and attributes over time, but the definition of an active customer rarely changes. Big changes to the core model tend to follow big changes to the business itself: picture an online bookstore buying a chain of physical bookshops, suddenly facing cash registers, stores, staff on the floor and a new logistics operation. That's a real event with a board meeting attached, not a weekly occurrence. Most of the time the core business objects stay stable.

The layer that moves quickly is DAR: new questions bring new metrics, new fact tables and new specialized requirements, and they keep coming. Built on a modeled DAB, that work is straightforward and deterministic, because the entities, relationships and history it needs are already declared. Most DAR use cases can be declared too, putting the fast-changing layer under the same discipline as the slow one.

So on a declared platform, most of the daily change lands one layer up, in DAR, or arrives in the DAB as a new attribute or entity that leaves the existing definitions alone.

When something does change, generation decides how it travels. Drift between the layers can't happen quietly: a broken declaration fails loudly at its source, and a wrong one is wrong in exactly one place, which is the place you fix it. Compare that with the skill-file world, where a wrong definition surfaces as a disagreement between two numbers in a Slack thread, weeks later, and someone has to work out which copy was wrong.

What This Closes for the Agent (and What It Doesn't)

The agent reads the same definitions the platform was built from, so there's no month in which accuracy can slide while the documents catch up. The question from that Slack thread, why the agent's number doesn't match the dashboard, mostly stops coming up, because both numbers are generated from the same declaration.

It also gives the agent structure to walk. During a standup, a churn number on a dashboard looked wrong. An agent traced it down through the declared layers, from the DAR metric to its DAB definition, through the DAS source data and into the ingestion logs, and had found the root cause before the standup ended. That story deserves its own write-up, and it's getting one. It worked because every step of the walk was a declared relationship: the agent didn't have to guess which table fed the metric or which source fed the table, because the platform could tell it.

Generation has a limit, though, and it's the one that matters most. It enforces a definition consistently. It can't make the definition correct.

Say the model declares an active subscription as started, not yet ended, and worth more than zero. Finance, it turns out, also excludes paused subscriptions, and nobody told the modeler. On a declared platform, every layer generated from that model inherits the same mistake. The dashboard, the finance extract and the agent all agree, and they're all wrong by the same amount. Consistency looks a lot like correctness from the outside.

That makes a badly modeled business layer more dangerous than a messy warehouse, in one specific way. In the skill-file world, the disagreement between the agent and the dashboard is irritating, but it's also a signal: somebody posts a screenshot, and the thread exposes the problem. On a generated platform that signal is gone. A confidently wrong agent at platform scale produces no thread, because there's nothing to disagree with.

The hard work moves upstream, into modeling. What does the business mean by "active"? Who decides? What happens to paused accounts? Those are questions for people who understand the business, and no generator answers them. Generation takes the copying, syncing and reconciling off those people's plates. When they get a definition wrong, the error lives in one declaration, a much better place to find it than in four documents and a dashboard.

What to Check Before You Add Another Skill

Next time the agent gets a number wrong, you'll feel the pull to open the skill file and add a line. Before you do, ask one question about the fact the agent needed: where does it live?

There are only a few possible answers. It might be declared once in a model the platform is generated from, in which case the fix belongs there. It might live in code, a filter in a CTE somebody wrote two years ago, in which case the skill file you're about to edit is a second copy of it. Or it lives only in a document or the skill file itself, in which case nothing in the platform enforces it.

The last two answers are where drift comes from. A new line in the skill file fixes today's answer and adds one more copy to keep in step, the trade Anthropic made deliberately and then built machinery to manage. That trade can be the right one: with one source system and a handful of metrics, a well-kept skill file is a perfectly good platform for an agent. The cost grows with every copy, and few teams we meet have counted theirs.

So count them. This week, take the five questions your agent, or your dashboards, gets asked most, and for each one find every place the definition behind the answer is written down: the transformation SQL, the BI tool's calculated fields, the semantic layer config, the skill files, the wiki. Note which of those places are generated from a single declaration and which were written by hand to match it.

If every answer traces back to one declaration, adding a skill is cheap and safe, because the skill reads from the same source as everything else. If the count comes back at four or five hand-written copies per definition, the next skill file won't fix the agent. It will add one more copy, and the next Slack thread is already on its way.

If you'd like to see what the declared version looks like in practice, why we built Daana is where that story starts.