Databricks built an agentic Customer Data Platform where your data already lives, not in another copy
Most companies do not have a customer data problem. They have a customer data copy problem.
The data already lives in the lakehouse. It is clean. It is governed. And then a separate Customer Data Platform comes along and copies all of it somewhere else. A new home, new rules, new costs.
At the Databricks Data and AI Summit, Databricks introduced a different idea. It is called CustomerLake. It is an agentic Customer Data Platform built directly into the lakehouse. Not bolted on. Not another copy.
Let us walk through what that actually means, and why it matters for the people who build and maintain data systems every day.
A traditional CDP usually works by pulling your customer data out of the warehouse and into its own world.
That sounds harmless. It is not.
When data is copied, three quiet costs appear. You now store the same data twice, so you pay twice. You now govern the same data in two places, so rules drift apart. And you now break lineage, so you can no longer trace where a number came from.
Databricks calls this the CDP tax. It is the slow, ongoing price of duplication. You feel it in storage bills, in audit meetings, and in the moment someone asks "is this customer count correct?" and nobody is sure.
The cheapest copy of data is the one you never made.
CustomerLake keeps the customer data where it already is and brings the customer tools to it.
It offers the things a marketing or growth team expects from a CDP. A Customer 360 view that brings every signal about a person into one profile. Identity resolution that decides which records belong to the same human. Audience segmentation to group people by behavior. Activation to push those audiences out to the channels that reach them. And personalization to tailor the experience.
The difference is the foundation. All of this is built on Unity Catalog. So the same governance, permissions, and lineage you already trust for your tables now cover your customer profiles and audiences too.
If you have studied how Unity Catalog organizes access and lineage in the governance chapters of the book, this will feel familiar. CustomerLake is that same idea, pointed at customer data.
What makes CustomerLake agentic is two kinds of AI agents doing the heavy work.
Profile Agents take raw, messy customer signals and turn them into business-ready Customer 360 profiles. Think of the work you normally do to clean, join, and resolve identity across sources. The agent helps shape that into a usable profile.
Campaign Agents work on the activation side. They build audiences, suggest the next best action for each person, launch campaigns, and then keep watching the results to improve them.
You can picture this as the medallion idea applied to people. Raw signals are the bronze layer. Profile Agents refine them toward silver and gold. Campaign Agents act on the gold profiles. The pattern you already practice with tables now shapes how customer engagement is built.
Databricks uses a phrase here that is worth slowing down on. Infinity Campaigns.
A normal campaign has a start and a stop. You design it, launch it, wait, and review.
An Infinity Campaign is always on. The Campaign Agent keeps adjusting audiences and actions in real time as new behavior arrives. It does not wait for the next planning cycle. It adapts as the data changes.
That only works if the data underneath is fresh and trusted. Which is exactly why building this on the lakehouse matters.
CustomerLake is not a walled garden.
It connects out to the marketing tools teams already use, like Adobe, Meta, Braze, LiveRamp, and The Trade Desk. Lakehouse Federation lets it reach data across systems without forcing another copy.
So you get two things that usually fight each other. Openness to the tools the business loves, and a single governed foundation underneath. Unity Catalog stays the source of truth. The channels are just where the work goes out.
It is easy to read this as a marketing announcement. It is not, at least not only.
CustomerLake is really a statement about architecture. It says the cleanest path is to bring intelligence to the data instead of moving the data to the intelligence.
If your organization already uses the Medallion Architecture and Unity Catalog, CustomerLake is a natural next step, not a new platform to learn from scratch. The foundations you have been building are the foundations it runs on.
CustomerLake is in Private Preview today. So you do not need to act on it tomorrow. But it is worth understanding the direction, because the direction is clear. Fewer copies. One governed source. Agents doing the repetitive work close to the data.
That is a future worth preparing for. And the preparation is the same calm, careful work it has always been. Understand your data, govern it well, and keep one trusted copy.
If you want to build that foundation properly, the Unity Catalog and Medallion chapters in the book are the place to start.