Data journalism increasingly relies on reporters to uncover trends, disparities and accountability stories within large structured datasets. In practice, exploring such data is slow and brittle: hundreds of variables, evolving schemas, and opaque coding conventions force journalists to write non‑trivial analysis code while their hypotheses shift. Typical “ask in English, get SQL” promises break down because LLMs often mis‑match schemas, misinterpret domain semantics or units, and make silent assumptions.
DataWeave tackles these problems with four tightly coupled components:
- Conversational layer that accepts natural‑language exploratory queries;
- Schema‑grounding layer that maps the dialogue to dataset metadata (tables, columns, types, units);
- Analytical‑planning layer that translates intent into a sequence of statistical operations such as aggregation, grouping or correlation;
- Executable‑query layer that renders the plan into SQL or Pandas code and returns inspectable results.
Instead of treating the LLM as an autonomous answer engine, DataWeave treats it as an interactive partner whose outputs can be inspected, corrected and steered. Users can view the generated query, adjust column mappings or modify the analysis plan, then re‑run to validate a hypothesis.
A real‑world case study involved professional journalists exploring the U.S. Department of Education’s Integrated Postsecondary Education Data System (IPEDS). IPEDS contains thousands of variables, frequent schema updates, and complex educational hierarchies with unit conventions. Reporters used conversational prompts to isolate subsets of interest; the system automatically produced SQL with appropriate unit conversions, displayed results in an interactive table, and helped reveal systematic differences in scholarship allocation across institution types.
During deployment we iterated the architecture: the original single‑turn Q&A was replaced by multi‑turn dialogue; schema caching moved from static files to a live metadata service; error handling was extended to detect schema drift automatically. These changes tripled the system’s usage time in a newsroom and cut the error rate to about 5 %.
Three design principles emerge from our experience:
- Keep LLM output auditable; every generated query must be visualised and manually editable;
- Continuously align to the latest schema to mitigate drift;
- Make the analytical plan explicit so domain knowledge can be injected at each step.
Review