Thirteen years, one architectural argument, and why the 2026 funding round is really an engineering story
There is a moment most data engineers remember.
You are sitting at a desk at 8 in the morning, coffee going cold, watching a job that was supposed to finish at 2. Somewhere in the middle of the night a file arrived with one extra column, and the whole chain stopped. The warehouse team is waiting on you. The analysts are waiting on the warehouse team. The dashboard that the business will look at in an hour is still showing yesterday.
Nobody designed that morning on purpose. It came out of an assumption the whole industry accepted for years: storage for raw data is one system, storage for trusted data is another system, and you move data between them on a schedule while humans wait.
On August 13, 2026, Databricks closed a $5 billion round at a $190 billion valuation. The round was led by Coatue, joined by Blackstone, MGX, T. Rowe Price, and new investor Sixth Street Growth. The company had announced the term sheet a month earlier at $188 billion, and by the close the valuation had ticked up. It is a very large number, and large numbers usually get covered as finance news.
We want to cover it as an engineering story, because that is what it actually is. The valuation is the market pricing thirteen years of a single argument: that the split which created those bad mornings was never necessary.
This is the long version of that story, from a research project in a Berkeley lab to the company that a lot of the data world now builds on. It is also, honestly, a thank you note. We spend our days writing lessons about this platform. It has been a pleasure to teach.

In 2009, a group of researchers at the UC Berkeley AMPLab started working on a system called Spark.
The problem they were looking at was simple to state. MapReduce had made large scale data processing possible, but it wrote intermediate results to disk between every stage. For a single pass over a lot of data, that was fine. For anything iterative, like the machine learning work the lab cared about, it was painful. You paid the disk cost again and again for the same data.
Spark's idea was to keep working data in memory across stages, and to give people a small set of operations they could compose. That is a modest sounding change. In practice it moved certain jobs from hours to minutes, and it made a class of work feel possible that had not felt possible before.
Two decisions in those early years mattered more than the performance numbers.
The first was that Spark was open sourced in 2010, early, while it was still rough. The second was that in 2013 it was donated to the Apache Software Foundation, where it became a top level project. The researchers gave away the thing they could have kept.
That posture, build in the open and let other people depend on it, shows up again and again in the rest of this story. It is worth noticing now, because it explains a lot of later decisions that would otherwise look strange for a company of this size.
Databricks was founded in 2013 by seven people from that lab: Ali Ghodsi, Ion Stoica, Matei Zaharia, Patrick Wendell, Reynold Xin, Andy Konwinski, and Scott Shenker.
Seven founders is unusual. Most companies would treat that as a governance problem. In this case it was a fair reflection of how the work was actually done, and it set the culture. Research first, publish, then productize.
The first commercial product arrived in June 2014, alongside a $33 million Series B led by New Enterprise Associates. It was a hosted way to run Spark, with notebooks attached.
That combination sounds ordinary in 2026 because everyone copied it. At the time it was quietly important. Running Spark yourself meant standing up a cluster, tuning it, keeping it alive, and explaining to your manager why it fell over. Databricks turned that into something you could open in a browser and start typing in. Notebooks meant your exploration, your code, and your results lived in the same place, which is how people actually think when they are working out a problem.
For a few years that was the company: the best place to run Spark. A useful business, and a limited one. The interesting part is what they did next, because it required arguing with the market instead of serving it.
Here is the constraint that shaped everything.
Object storage, meaning S3 and its equivalents, is wonderful. It is cheap, it is effectively unlimited, and you can point many engines at the same bytes. It is also, on its own, a bag of files. It gives you no transactions, no schema enforcement, and no reliable way to know what a table looked like an hour ago. Two writers can produce a mess. A reader can see a half finished write.
So teams did the sensible thing. They kept the cheap lake for raw data, and they loaded the important, trusted, queryable data into a warehouse. Two systems, two copies, two sets of permissions, and a pipeline in the middle that had to run overnight. Those bad mornings were the natural consequence.
Delta Lake, which Databricks open sourced in April 2019, attacked the assumption directly. The idea is not complicated once you see it. Keep the data in open Parquet files on object storage, and add a transaction log that records, in order, which files belong to the table. Readers consult the log, not the directory. Writers append to the log atomically.
From that one mechanism you get a surprising amount:
If you want to feel this rather than read about it, our Delta Lake lesson has you break a table on purpose and then recover it with time travel. That exercise is the whole point of the format in about ten minutes.
The reason Delta mattered so much is that reliability was the actual reason people kept a separate warehouse. Once the lake became trustworthy, the second system became a choice rather than a requirement.
With Delta in place, Databricks made a claim: you do not need a lake and a warehouse. You need one place, with warehouse behaviour on lake economics. They called it the lakehouse.
The industry did not immediately agree. It sounded like a marketing word, and for a while the performance gap was real. Warehouses had spent decades on query execution.
So the company closed the gap in public. Photon, a vectorized query engine written in C++ and introduced around 2020, made SQL on Delta tables genuinely fast. Databricks SQL followed, reaching general availability in late 2021 and giving analysts warehouse style endpoints and BI connections over the same tables the engineers were writing.
That is the pattern to notice. The response to "your architecture is slower" was not a louder claim. It was an engine.
The lakehouse also needed a way of working, not just a place to store things. That is where the medallion pattern came in: raw data lands in bronze, gets cleaned and conformed into silver, and gets shaped for consumption in gold. It is a plain idea that solved a real organizational problem, because it gave a team a shared vocabulary for what a table is allowed to promise. We wrote a full medallion architecture explainer and a hands-on medallion lesson because it is still the single most useful mental model for a new data engineer.
Every platform hits this wall. You succeed, adoption spreads, and suddenly there are four hundred tables, nobody is sure which ones are safe to use, and a security review is coming.
Unity Catalog, which rolled out through 2023, was the answer: one metadata and permission layer across workspaces, with a three level namespace of catalog, schema, and table, plus lineage, auditing, and search. Permissions moved from "who can read this storage path" to "who can select from this table, and which columns can they see."
Two things about Unity Catalog deserve credit.
The first is that it treated governance as a platform feature instead of an add-on product. Lineage is computed because the platform ran the query, not because someone maintained a diagram.
The second is that in June 2024, Databricks open sourced it. A company with a strong commercial position gave away the layer that holds the keys. That was a real decision with real revenue implications, and it is consistent with 2010 and 2013.
If governance is where you are right now, start with our Unity Catalog guide for the concepts and the Unity Catalog lesson for the practice.
Pipelines were the next honest weak point. Most teams were writing notebooks that called notebooks, wiring them into schedules, and holding the dependency graph in their heads.
The shift here was declarative. Instead of writing the steps, you declare the tables you want, the queries that define them, and the quality expectations they must meet. The platform works out the order, the incremental processing, and the recovery. Lakeflow, announced in June 2024, brought ingestion, transformation, and orchestration under that one idea, and Lakeflow Declarative Pipelines became the way most new pipelines get written.
This is the change people underestimate. A script tells you what happened. A declaration tells you what is supposed to be true. Only the second one can be checked automatically, and only the second one can be safely re-run by something that is not you. Our Lakeflow guide and the data quality lesson cover the expectations side, which is where the value actually shows up.
Two supporting moves fit here. In June 2024, Databricks acquired Tabular, the company founded by the creators of Apache Iceberg, in a deal widely reported at around a billion dollars. Rather than win a format war, they bought both sides of it and started working toward interoperability, so that a table can be read by more engines rather than fewer. And in February 2025 they acquired BladeBridge to help convert legacy warehouse code, because a platform argument only wins if migration is possible for people with twenty years of SQL.
In June 2023, Databricks acquired MosaicML, a model training company, in a deal reported at about $1.3 billion. At the time some read it as a company chasing the moment.
Looking back, the logic was straightforward. If the useful AI applications inside a company depend on that company's own data, then whoever governs the data is standing in the right place to serve the models. The data platform was already there. It needed serving, training, evaluation, and a way to treat models like managed assets.
What followed came fast:
The 2026 Summit continued in the same direction, with Genie Ontology, Genie One, Genie Agents, and Unity Catalog Metrics. The theme was consistent: give the AI layer the same semantics and permissions the data layer already has, so that an agent answering a question is reading the governed definition rather than guessing.
We covered the two pieces we think matter most for engineers in the Unity AI Gateway article and in the piece on the data stack being redesigned for agents.
Here is the whole arc in one view.
| Year | What happened | Why it mattered |
|---|---|---|
| 2009 | Spark research begins at the UC Berkeley AMPLab | Iterative work stops paying the disk cost on every stage |
| 2010 | Spark open sourced | Other people can build on it, early |
| 2013 | Spark donated to Apache. Databricks founded by seven researchers | The engine belongs to the community, the company sits beside it |
| 2014 | First hosted product, with notebooks | You stop babysitting clusters and start working |
| 2019 | Delta Lake open sourced | The lake becomes reliable enough to trust |
| 2020 | Photon engine and Databricks SQL | Warehouse speed on lake tables |
| 2021 | Databricks SQL reaches general availability | Analysts and engineers share one set of tables |
| 2023 | Unity Catalog rollout. MosaicML acquired | One permission layer, and a path into AI |
| 2024 | Lakeflow and AI/BI. Unity Catalog open sourced. Tabular acquired | Declared pipelines, open governance, no format war |
| 2025 | Lakebase, Agent Bricks, MLflow 3.0. Tecton acquired | Operational data and agents join the platform |
| 2026 | Genie Ontology and Agents, Unity AI Gateway | Agents get governed meaning, not just access |
Read as a list of releases, it looks like a company shipping features. Read as a shape, it is one idea being extended into every layer it can reach.
graph TD
A["Cheap open storage"] --> B["Transactions and schema<br/>Delta Lake"]
B --> C["One place for BI and engineering<br/>Lakehouse"]
C --> D["One permission and lineage layer<br/>Unity Catalog"]
D --> E["Declared pipelines instead of scripts<br/>Lakeflow"]
E --> F["Models and agents on governed data<br/>AI Gateway, Genie, Agent Bricks"]Each step only became possible because the one before it held. You cannot put a warehouse on a lake until the lake is reliable. You cannot govern one place until there is one place. You cannot let an agent write to production until intent is declared and quality is checked. You cannot let an agent answer a business question until the business definitions are governed.
That is why the platform feels coherent to teach. It is not a bundle of acquisitions. It is a sequence.
Another way to read the same sequence is to ask, for each wave, what it stopped people from suffering.
| Wave | What it shipped | The problem it removed |
|---|---|---|
| Spark | In-memory distributed processing | Re-reading the same data from disk on every stage |
| Hosted Spark and notebooks | A browser you could work in | Standing up and babysitting your own clusters |
| Delta Lake | Transactions and schema on object storage | Corrupt tables, half-written files, no way back |
| Photon and Databricks SQL | Fast SQL over lake tables | Copying trusted data into a second system for BI |
| Medallion pattern | A shared shape for tables | Nobody knowing which table is safe to use |
| Unity Catalog | One permission, lineage and audit layer | Path-level security and hand-drawn lineage diagrams |
| Lakeflow | Declared tables and expectations | Notebooks calling notebooks on a schedule |
| Serving, MLflow, Genie, Agent Bricks | AI next to the governed data | Exporting data to a separate AI stack |
| Lakebase and Tecton | Transactional and real time reads | Bolting an operational database onto an analytics platform |
Now the funding, which reads very differently once you know the engineering timeline it sits on.
| Date | Amount | Valuation | Notable investors |
|---|---|---|---|
| June 2014 | $33M | not disclosed | New Enterprise Associates |
| December 2016 | $60M | not disclosed | New Enterprise Associates |
| 2019 | $250M | about $2.75B | Andreessen Horowitz, Coatue |
| February 2021 | $1B | $28B | Franklin Templeton |
| August 2021 | $1.6B | $38B | Counterpoint Global |
| September 2023 | over $500M | $43B | T. Rowe Price, with NVIDIA as a strategic investor |
| December 2024 | $10B | $62B | Thrive Capital as a lead, broadly syndicated |
| September 2025 | $1B | over $100B | Series K, co-led by Andreessen Horowitz, Insight Partners, MGX, Thrive Capital and WCM |
| December 2025 | over $4B | $134B | Series L, existing and new investors |
| July 2026 | undisclosed at signing | $188B | term sheet led by Coatue |
| August 2026 | $5B | $190B | Coatue, with Blackstone, MGX, T. Rowe Price and Sixth Street Growth |
The revenue line underneath it moved in the same direction, and it kept accelerating. Around $1.5 billion annualized in 2023. Crossing a $3 billion run rate with positive free cash flow at the end of 2024. More than $4.8 billion by the end of 2025, growing over 55 percent year over year. A $5.4 billion run rate reported in February 2026, at more than 65 percent growth. About $6.9 billion by June 2026, with AI products alone at roughly a $1.7 billion run rate. And, alongside the August close, more than a $7 billion run rate growing over 80 percent year over year.
Read that last part again. A company at this size growing faster than it was a year earlier is rare, and it is the actual reason the valuation moved.
Two details about this round are worth pausing on.
The first is that the $188 billion and $190 billion numbers are not two rounds. On July 16, 2026, Databricks announced it had signed a term sheet for a strategic round at a $188 billion valuation, led by Coatue, without disclosing the amount. Ali Ghodsi was direct about the stage it was at, noting the money was not in the company's hands yet. On August 13 the same financing closed, at $190 billion, with the size confirmed at $5 billion. One process, two moments.
The second, reported by TechCrunch, is that the company set out to raise about $1 billion and ended up taking $5 billion because of investor demand. Databricks said the capital goes toward products that help companies build and manage AI applications, and it named Unity AI Gateway, Genie, and Lakebase specifically.
There is also a decision worth respecting here. In June 2026, CEO Ali Ghodsi said plainly that 2026 was a bad year to go public, pointing at how much IPO capital other large private companies were absorbing, and that Databricks would wait for a quieter window. A company with this much momentum choosing patience over a headline is a signal about how it is being run.
If you have never been to a Data + AI Summit, the thing that comes across is not the size. It is the pace.
Every year the keynote does two things at once. It ships something significant, and it tells you where the platform is heading. 2024 was pipelines and governance in the open. 2025 was applications, agents, and operational data. 2026 was semantics for agents, with ontology and governed metrics.
For those of us who write about this platform, Summit is also a stress test. You get roughly a week to work out which announcements change how someone should build on Monday and which are further out. That is where most of our writing gets its schedule, including the piece on Lakemeter and estimating pipeline cost before you build, which came out of exactly this kind of reading.
The rhythm has a side effect worth naming for anyone learning. Databricks moves fast enough that memorizing menus is a losing strategy. Understanding why a layer exists is not, because the reasons change far more slowly than the interfaces. That is the entire premise of how we write.
Praise is worth more when it is not blind, and a story like this is easy to tell as if every step were obvious. It was not.
The lakehouse claim was ahead of the product for a while. Early on, teams who moved BI workloads onto the lake did hit queries that were slower than the warehouse they left, and telling them the architecture was correct did not help them at 9 in the morning. Photon and Databricks SQL closed that gap, but the gap was real first.
Unity Catalog migration was genuine work. Teams with years of workspace-level tables and path-based permissions had to plan a move, and plenty of them put it off long enough that it turned into an audit problem.
The platform is also large now. A newcomer opening it in 2026 sees jobs, pipelines, dashboards, catalogs, serving endpoints, agents, and a database, and it is fair to feel lost. That is a cost of range, and it is exactly why we teach in a fixed order instead of touring the menus.
And cost has been a recurring complaint. Serverless made things easy to start and easy to leave running. The recent attention to estimating and attributing spend is the company responding to a real, repeated piece of feedback rather than a new idea it invented.
None of this undercuts the achievement. It just means the achievement was earned by fixing things, which is the only way engineering ever earns anything.
We would rather be useful than dramatic, so here is our read on the direction, based on what has already shipped.
Agents become normal users of the platform. Not a demo, an account. Which means they need permissions, an audit trail, a spend limit, and definitions they cannot misread. Every piece of that already exists for humans. The work of the next few years is making it true for software.
Context becomes the scarce resource. A model is easy to call. Knowing which table is the trusted one, what "active customer" means in this company, and which numbers a director already signed off on is the hard part. Governed metrics and ontologies are Databricks building exactly that. This is the same argument our sister book The Context Advantage makes for the human side of the job: the value is not in the prompt, it is in the context you can hand over with confidence.
Cost moves into design. For years cost was something you discovered in the monthly bill. Tools that estimate before you build, and metrics that attribute spend to a pipeline, turn cost into a design input like latency. Expect more of this, and expect interviewers to ask about it. Our Databricks pricing guide is the plain English version if that is new to you.
Open formats keep winning, quietly. Delta and Iceberg interoperability, an open catalog, open MLflow. The commercial value is moving up to governance, execution, and the agent layer, and the storage layer keeps getting more open. That is good for anyone who has to make a decision they will live with for a decade.
BricksNotes exists because we learned this platform the slow way, and thought the path could be gentler.
We are not affiliated with Databricks. We are engineers who teach. What we admire, honestly, is a set of choices that were not obviously good for a company: giving Spark to Apache, open sourcing Delta Lake, open sourcing Unity Catalog, buying the Iceberg team instead of fighting them, and answering performance criticism with an engine instead of a press release. Those are the decisions of people who intend to still be right in ten years.
So, genuinely: congratulations, and thank you. Not for the valuation, which is a consequence. For the thirteen years of work that made the valuation reasonable, and for building it in a way that let the rest of us read the source.
If you are new, do not start with the news. Start with the ground.
If you are studying for a certification, our certification guide lays out the paths and costs, and the free Data Engineer Associate practice exam will tell you honestly where you stand in 90 minutes.