Two verbs, two very different engineering problems
The request came in at 9:12 on a Tuesday morning. "Can you add a refund_reason column to the payments table, backfill it from the raw files, and tell the finance dashboard about it?"
Six months ago that was a two day task. Read the ticket. Open the notebook. Search Slack for the person who owns the raw files. Find the table. Guess the column names. Run something. Break something. Fix it.
That Tuesday, an agent did the first half in about forty seconds. It found the payments table, read its schema, found the raw source location, noticed the column already existed in the JSON payload, and proposed a MERGE statement. Then it stopped and waited for a human to say yes.
That is the shape of the new work. The agent discovers. The agent applies. And the quality of both steps depends almost entirely on what we, the data engineers, left behind for it to read.
Most conversations about AI in data platforms collapse into one word: automation. That word hides the interesting part.
An agent doing real work on a data platform does two very different things.
It discovers. It looks around. It asks what tables exist, what they mean, who owns them, how fresh they are, what depends on them, and what happened the last time someone touched them.
Then it applies. It writes a query, creates a table, runs a MERGE, triggers a pipeline, or updates a dashboard.
Discovery is a read problem. Applying is a trust problem. They fail in different ways, and they need different work from us.
Discovery fails quietly. Applying fails loudly. The quiet failure is the one that costs you more.
When an agent gives a wrong answer about your data, the instinct is to blame the model. Usually the model is fine. It simply had nothing good to read.
Think about what an agent actually sees when it looks at your lakehouse. It sees table names, column names, comments, tags, owners, lineage, and recent history. That is the whole world from its point of view. If your table is called tbl_pmt_f_v2 and no column has a comment, the agent is guessing. A new teammate would be guessing too.
This is why Unity Catalog stopped being a governance checkbox and became the thing your agents read. Names, comments, and tags are no longer documentation. They are the interface.
You can improve discovery today with plain SQL.
COMMENT ON TABLE main.finance.payments IS
'One row per completed payment. Source: Stripe events. Grain: payment_id.';
ALTER TABLE main.finance.payments
ALTER COLUMN refund_reason
COMMENT 'Free text reason supplied by support. Null when not refunded.';
ALTER TABLE main.finance.payments
SET TAGS ('domain' = 'finance', 'certified' = 'true');Three statements. Nothing clever. But an agent reading that table now knows the grain, the source, the meaning of a confusing column, and whether a human has blessed it. That is context, and context is the whole game. We wrote about why in Context engineering is becoming a real job skill for data engineers.
A useful habit: before you let an agent near a domain, read your own catalog the way it would.
from pyspark.sql import functions as F
tables_df = spark.sql("SHOW TABLES IN main.finance")
# Any table without a comment is a table your agent will guess about.
described_df = spark.sql("DESCRIBE EXTENDED main.finance.payments")
missing_comments_df = described_df.filter(
(F.col("col_name") != "") & (F.col("comment").isNull())
)
missing_comments_df.show(truncate=False)If that last query returns a long list, you do not have an AI problem. You have a description problem, and it is fixable this week.
Discovery is safe. Applying is not.
The moment an agent can write, the question changes from "is it correct?" to "what happens when it is wrong?" And it will be wrong sometimes, the same way people are wrong sometimes.
Good engineering here is old engineering. Four habits carry most of the weight.
Make every action repeatable. If the agent runs the same step twice, the table should look the same. That is idempotency, and it is the difference between a retry and an incident. We walked through the patterns in The job ran twice. Why safe reruns decide if a pipeline is production ready.
Let the agent write to its own space first. A staging schema the agent owns is much calmer than a certified table it can touch. Promotion stays a separate, reviewed step.
Prefer MERGE over overwrite. A merge with a clear key is easy to reason about and easy to undo. An overwrite on the wrong path is a bad afternoon.
MERGE INTO main.finance.payments AS target
USING staging_agent.payments_refunds AS source
ON target.payment_id = source.payment_id
WHEN MATCHED THEN
UPDATE SET target.refund_reason = source.refund_reason;Keep time travel close. Delta Lake gives you a version history, and that history is your undo button.
DESCRIBE HISTORY main.finance.payments;
-- Read the table as it was before the agent touched it.
SELECT COUNT(*) FROM main.finance.payments VERSION AS OF 412;
-- Restore if the change was wrong.
RESTORE TABLE main.finance.payments TO VERSION AS OF 412;If you are new to these mechanics, the Delta Lake lesson and the incremental processing lesson cover them from the beginning, and both run on Databricks Free Edition.
There is one more piece, and it is easy to skip.
An agent that discovers and applies without checking its own work is just a faster way to create problems. Expectations, row counts, null checks, and freshness rules are how the agent learns whether the thing it just did was good.
The nice part is that these checks are readable. When a check fails with a clear name, the agent can explain what went wrong instead of shrugging. Our data quality lesson shows the patterns, and The schema changed overnight covers what happens when the source moves under you.
Governed access matters too. When agents call models and tools through one governed door, you get logs, limits, and a way to answer "who did that?". We explained that shift in Unity AI Gateway is GA, and the wider architectural change in The data stack is being redesigned for AI agents.
None of this makes the data engineer smaller. It moves the work.
Less time typing the query. More time deciding what a table means, what an agent is allowed to do with it, and how you would know if it went wrong. Naming, describing, and bounding become the high value skills. They always were, quietly. Now they have an audience that reads everything you write.
If you want the longer argument about how context becomes an advantage rather than an afterthought, that is the subject of our companion book, The Context Advantage. And if you want the hands-on foundations, Thinking in Data Engineering with Databricks walks the same ground with runnable examples, from your first DataFrame through Unity Catalog and governed AI serving.
Pick one table you own. Not the whole catalog. One table.
Give it a comment that says what one row means. Comment its three most confusing columns. Tag its domain and owner. Then ask an agent a question about it and see how much better the answer is.
That is the loop. The agent discovers what you described. The agent applies what you allowed. The judgment stays with you.
Start with Unity Catalog if you want the governance groundwork, or Start Here if you are beginning from zero. Either way, the next table you describe well is the next thing your agents get right.