More on the topic…
AI has made it trivial for non-technical people to generate data code and analytics on the fly. Instead of submitting tickets and waiting weeks for a chart, someone can now ask an AI agent in plain language and get results instantly. The scale is staggering—millions of people are writing billions of lines of code this way, and at this point more non-technical folks are probably generating data code than using traditional BI tools. The speed and accessibility are genuinely valuable, which is why nobody's going back to the old process.
The catch is that this speed creates chaos. All these generated queries, dashboards, and reports live scattered across Slack messages, HTML files on laptops, and one-off chat sessions. There's no audit trail showing what context the AI had access to, which tables it queried, or what filters it applied before spitting out an answer. Nobody reviews the code. Nobody can reproduce it. The numbers look polished and confident on a board deck, but there's zero visibility into whether they're actually correct. You can't evaluate quality you can't see, and you definitely can't enforce consistent metric definitions across thousands of independent AI-generated analyses.
This leaves data teams stuck between two bad options: either let people generate whatever they want and surrender any hope of governance, or force them back into rigid traditional BI tools that can't do much of anything. The article calls this "AI Data Sprawl"—a growing pile of plausible-looking numbers and ungoverned code that nobody can trace, yet somehow the data team still owns the mess. The author works at Hex and hints they're building generative data apps with better context, controls, and collaboration features to address this, but frames the problem as largely unsolved across the industry.
Questions about this article
No questions yet.