Skip to content
Rush Commerce
AI & Automation4 min read

OpenAI's Data agent: your semantic layer is the product

OpenAI shipped a Data agent in ChatGPT Work that queries your warehouse and builds dashboards. What it returns depends on definitions you probably never wrote.

OpenAI shipped a Data agent in ChatGPT Work on September 10 that connects to your warehouse, investigates what moved, and builds dashboards from a conversation. No SQL, no BI tool to learn. It will work for some teams immediately. Whether it works for yours depends on something most small companies never wrote down: what your numbers actually mean.

What actually happened

Per OpenAI's announcement, the Data agent connects to approved sources including Amazon Redshift, Datadog, Google BigQuery, ClickHouse, Databricks, MongoDB, Redis and Snowflake, and pulls files from Google Drive and SharePoint into the analysis. It builds and edits dashboards in Omni, Oracle BI, Power BI, Sigma, Tableau and ThoughtSpot. You install it as Data from the Plugins directory; admins enable it under Workspace settings → Plugins.

Two details carry more weight than the demo.

First, the context. OpenAI says the agent takes your business terms, metric definitions, custom calculations and data relationships from semantic layers and trusted sources — naming Databricks Genie Ontology, dbt, GitHub, Snowflake Horizon and existing BI dashboards. That is where "revenue" gets defined.

Second, the access model. Admins choose which connections are available and which roles can use them, and queries enforce the connected account's existing permissions, including table, row and column restrictions. The agent inherits the grants of whatever account you wired up. It adds no permission layer of its own.

OpenAI says nearly all of its product team and over two-thirds of its go-to-market org already use data agents in ChatGPT Work, and names NTT DATA, Thermo Fisher Scientific and ServicePiston among alpha participants. The announcement publishes no pricing and no accuracy benchmarks. It does say you can review the evidence behind each finding — the right affordance, and an admission that you should.

Why your semantic layer decides what the Data agent says

Here is the failure mode nobody demos. You ask "why did revenue slow in August." Somewhere in your warehouse sit gross bookings, net of refunds, recognized revenue, and a Stripe payouts table that lags two days. If none of those is labeled, the agent picks one. It picks confidently, produces a clean chart, and is wrong in a way that survives three meetings because the chart looks finished.

dbt, Genie Ontology and Snowflake Horizon are not ceremony. They are where you say refunds are subtracted, trials are excluded, the fiscal month ends on the 28th. Teams that wrote that down get an analyst. Teams that did not get a very fast guesser.

The permissions model cuts the same way, and harder for small companies. If the connection uses a service account with broad SELECT — and in a ten-person shop it usually does, because that is the account someone created to make Metabase work in 2024 — then everyone who can use that connection reads everything that account can read. Payroll in an HR table. Customer PII in a support export. Your warehouse grants just became your AI access policy, and you have not read them since you wrote them.

The work is the work it has always been, now with a deadline:

  • Write the definitions down, starting with the five numbers your team argues about.
  • Scope a read-only account to the tables the agent should see — not the one your ETL uses.
  • Check the evidence on the first ten answers, not the tenth. The cheapest time to learn it misread your schema is before someone forwards the dashboard.

The bottleneck is finally your definitions rather than the tooling. Better problem. Still yours.

Key takeaways

  • OpenAI launched the Data agent in ChatGPT Work on September 10, connecting to Redshift, BigQuery, ClickHouse, Databricks, MongoDB, Redis, Snowflake and Datadog
  • It builds and edits dashboards in Omni, Oracle BI, Power BI, Sigma, Tableau and ThoughtSpot
  • Business context comes from semantic layers — Databricks Genie Ontology, dbt, GitHub, Snowflake Horizon — so undefined metrics get defined by the model instead
  • Queries enforce the connected account's existing table, row and column permissions; the agent adds no access layer of its own
  • A broad service-account grant becomes an org-wide read policy the moment you connect it — scope a dedicated read-only account first
  • The announcement publishes no pricing and no accuracy benchmarks; verify the evidence trail on early answers

Your numbers should mean one thing. We build the data layer under AI analytics — defined metrics, scoped read access, a warehouse that answers the same question the same way twice. See how we set up data infrastructure, or tell us which number your team argues about.

Sources: OpenAI: Now everyone can put data to work, ChatGPT Work for data teams.

  • #openai
  • #chatgpt-work
  • #data-agent
  • #analytics
  • #semantic-layer
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.