Databricks Acquires Row Zero: Live, Governed Spreadsheets Come to Genie

Business teams keep the spreadsheet they think in. Data teams keep permissions, freshness and lineage intact.

The quarter is closing. The finance team needs a forecast by Friday.

Someone opens Databricks, runs a query, exports a CSV, and drops it into a spreadsheet. Within an hour there are pivots, three scenarios, and a tab called "final_v2".

That spreadsheet is now the most important file in the company for the next five days. It is also completely disconnected from the data it came from.

That gap is the reason Databricks acquired Row Zero.

Spreadsheets are not the problem

It is easy for data teams to treat spreadsheets as a failure of discipline. Build a dashboard, the thinking goes, and people will stop exporting.

They never stop exporting.

A spreadsheet lets you think with your hands. You change one cell and watch the model move. You add a column without asking anyone. You try a bad idea in ten seconds and throw it away.

Dashboards cannot do that. They answer questions someone already thought of. Spreadsheets let you ask the question you just thought of.

So the honest goal is not to remove spreadsheets. It is to stop the export.

The spreadmart: the last ungoverned layer

The moment governed data leaves the platform as a file, four things quietly break.

Freshness. The numbers were true at the moment of export. By Wednesday they are a snapshot of Monday, and nothing in the file says so.

Permissions. Unity Catalog can mask a salary column or restrict a region. A CSV cannot. Once the file is on a laptop, it can be emailed to anyone.

Lineage. Ask where the 12.4 million came from and the answer is a filename. There is no path back to the table, the transformation, or the run that produced it.

Definitions. Each analyst applies their own rule for refunds, currency, and cutoff dates. Every copy becomes a slightly different version of the truth.

The industry has a name for what this creates: spreadmarts. Small, private data marts living in files, outside every control the data team built.

For years this was tolerated because the consumer was a human who could apply judgement. That assumption is ending.

Faded disconnected spreadsheets on the left converging into one live connected spreadsheet on the right

What Row Zero actually does differently

Row Zero is a spreadsheet built to work on live data rather than on a file you downloaded.

A normal spreadsheet loads rows into memory on your machine, which is why it slows down somewhere past a million rows and why your copy immediately starts drifting from the source.

Row Zero pushes the work down to the data platform. The grid is the interface. The query engine stays where the data lives. You scroll, filter, and pivot against millions of rows because the compute is not on your laptop.

The important part is not the row count. It is that there is no copy. The sheet is a view onto governed tables, so permissions and freshness still apply while you work.

The Genie to spreadsheet flow

Inside Databricks, this connects to something that already exists.

Genie answers questions in plain language. Ask for net revenue by region for last quarter, and Genie resolves what your company means by net revenue through Genie Ontology, then queries the governed tables behind it.

Today that conversation usually ends in a chart. The next step is usually an export, because a chart cannot answer "what if we shift two deals into October".

With a spreadsheet in that path, the flow keeps going:

  1. Ask the question in plain language.
  2. Genie resolves the business definition and returns the result.
  3. Open it as a live sheet instead of a download.
  4. Pivot, model, and forecast in the grid.
  5. The numbers stay connected to the source, with permissions and refresh intact.

The modeling still happens in a spreadsheet, because that is where modeling belongs. What changes is that nobody had to break the link to do it.

Humans and agents need the same version of the truth

This matters more now because humans are no longer the only readers.

Picture an AI agent that watches pipeline revenue and flags accounts at risk. It reads live governed tables. Meanwhile the sales lead is planning from a spreadsheet exported last Thursday.

Both are confident. Both are working from different numbers. The meeting becomes an argument about whose file is right, and nobody can prove it.

Multiply that by twenty agents and forty spreadsheets and you get an organisation that cannot agree on its own performance.

The fix is not smarter models. It is one shared definition, one set of permissions, and one source of freshness, used by people and agents alike.

The three layers underneath

A live spreadsheet is only as trustworthy as what sits below it. In the Databricks stack, that is three layers doing three different jobs.

Three stacked layers labelled Unity Catalog, Genie Ontology and Unity Gateway with a spreadsheet above them

Unity Catalog governs the data. Who can see which table, which column is masked, which rows a region manager is allowed to read, and where every number came from.

Genie Ontology holds the shared context. It records what the business means by revenue, churn, or active customer, so the same word resolves the same way every time.

Unity Gateway governs AI access. It is the control plane where AI tools and MCP services are treated as things you grant and audit, rather than integrations someone wired up privately.

Stack a spreadsheet on top of those three and it stops being a private copy. It becomes another governed interface onto the same trusted data.

Note on scope: Genie, Unity Catalog governance, and Row Zero integration are enterprise workspace capabilities. Databricks Free Edition will not give you the Genie to spreadsheet flow. The modeling discipline underneath it is still worth practising there, and the exercise below does exactly that.

Where this fits The Context Advantage

We look at moves like this through four questions: Choice, Context, Control, and Cost.

Choice. This is the clearest one. Finance keeps spreadsheets. Analysts keep SQL. Engineers keep PySpark. Agents keep their own interface. Nobody is told to abandon the tool they think in.

Context. The definitions live in one place instead of being re-invented in each file, so the same question gets the same answer no matter who asks.

Control. Access rules follow the data into the grid. Revoking a permission actually revokes it, instead of leaving copies behind on laptops.

Cost. The expensive part of ungoverned spreadsheets was never storage. It was the hours spent reconciling three numbers before a board meeting, and the decisions made on the wrong one.

You can read more on that framework at The Context Advantage.

Practice this in Free Edition

You cannot run Row Zero in Free Edition, but you can build the layer it depends on: one clean table with one official definition, so no spreadsheet needs to guess.

Start with a small orders table.

from pyspark.sql import functions as F

orders = spark.createDataFrame(
    [
        (1, "West", 1000.0, "paid",     "2026-09-05"),
        (2, "West",  500.0, "refunded", "2026-09-06"),
        (3, "East",  750.0, "paid",     "2026-09-07"),
        (4, "East",  300.0, "pending",  "2026-09-08"),
    ],
    ["order_id", "region", "amount", "status", "order_date"],
)

orders.write.mode("overwrite").saveAsTable("workspace.default.orders_raw")

Now write the definition down once, as a view, instead of leaving it in four spreadsheets.

CREATE OR REPLACE VIEW workspace.default.net_revenue_by_region AS
SELECT
  region,
  SUM(amount) AS net_revenue
FROM workspace.default.orders_raw
WHERE status = 'paid'
GROUP BY region
SELECT * FROM workspace.default.net_revenue_by_region ORDER BY region

East returns 750 and West returns 1000. Refunded and pending orders are excluded, on purpose, in one place.

Now try the comparison that makes the point.

SELECT region, SUM(amount) AS everything_total
FROM workspace.default.orders_raw
GROUP BY region
ORDER BY region

West now reads 1500 instead of 1000. Nobody made an arithmetic mistake. The second query simply used a different definition of revenue, and it is exactly the mistake a downloaded file invites.

That is the whole argument for governed spreadsheets, in two queries.

What to take from this

Row Zero is not a story about a better grid. It is a story about where the definition of a number lives.

If the definition lives in the platform, a spreadsheet is a safe interface and an agent is a safe consumer. If the definition lives in whoever exported the file last, every new interface multiplies the confusion.

Spreadsheets will run the business for another forty years. The work in front of data teams is making sure the numbers inside them are still connected to something true.

Brick by brick.

https://youtu.be/jvs5ljOW2DQ

Continue learning