talk · community record
Diving Through Data at OpenAI: How a Data Agent Navigates 70,000 Datasets
Bonnie Xu of OpenAI on the internal data agent her team built. Nearly the whole company uses the data platform: over 600 petabytes processed a day across roughly 70,000 datasets, with about 200,000 queries run in production daily, and growing fast. When ChatGPT launched the question was how many weekly active users there were; now it is how many instant checkout users are on Chat Pro in Japan -- the same shape of question, far more nuanced, and today it costs five Slack threads and two meetings. Why table discovery is the hard part at that scale, where similarly named tables hold different cohorts, team-specific views and columns added weeks later.