Most pipelines are no longer typed by hand. The next step is just as big. The same prompt that drafts your Lakeflow pipeline can also wrap it in a Lakeflow Job, with schedules, execution modes, and chained dependencies.
Somewhere in the last two years, a quiet line was crossed.
More than 60% of new data pipelines in Databricks workspaces now start in Genie Code. Not in an empty notebook. Not in a copied template from an old project. A person describes what they need, and an AI agent drafts the pipeline.
If you have used it, you know it writes pipelines well. Sources, transformations, medallion layers, expectations. Clean, readable code that a reviewer can actually follow.
But here is the part many teams have not noticed yet.
Genie Code can also create the Lakeflow Job that wraps the pipeline. The schedule. The execution mode. The retries. The notifications. The chain of dependencies that connects your pipeline to the rest of the workflow.
One prompt. That is it.
This article explains why that matters, how it works, and how to try it yourself in Databricks Free Edition.
A pipeline on its own is a definition. It says what the data becomes.
Raw events land in Bronze. Cleaned and deduplicated records move to Silver. Business-ready aggregates arrive in Gold. Expectations drop the bad rows. That is the what.
But a definition does not run itself. Somewhere, a human still has to answer the harder questions.
When does this run? Every hour? Every morning at six? Only when new files arrive?
What must finish first? Does the customer table need to refresh before the orders pipeline starts?
What happens after? Does a dashboard refresh? Does a quality audit run? Does someone get paged when it fails?
These are not coding questions. They are operating questions. And until they are answered, you do not have a production pipeline. You have a very good script.
This is the gap a Lakeflow Job closes.
Think of it this way.
A Lakeflow Declarative Pipeline is the recipe. Ingredients, steps, quality checks. It describes the data and nothing else.
A Lakeflow Job is the kitchen schedule. It says when cooking starts, which dish must be plated before another begins, who gets called if the oven fails, and whether the kitchen runs all day or only for the dinner rush.
When you wrap a pipeline in a Lakeflow Job, you gain four things.
Scheduling and triggers. Run on a cron schedule, on file arrival, or continuously. The job owns the timing, so nobody has to remember to press run.
Execution modes. Triggered mode starts the work, finishes it, and stops. You pay for what you use. Continuous mode keeps the pipeline alive for near real-time freshness. You can mix both inside one job.
Dependency chaining. A job is a graph of tasks. Validate the input files. Run the pipeline. Run the quality audit. Refresh the downstream tables. Each task waits for the ones it depends on.
Operational behavior. Retries on failure. Notifications to the right channel. A run history you can inspect when something looks wrong at 2 AM.
This is the difference between code that works and a system you can trust.
The reason this matters now is that the boundary between writing the pipeline and wiring the orchestration has collapsed into a single conversation.
A weak prompt asks for code:
Build me a pipeline that loads orders into Delta tables.A strong prompt describes the whole operation. Here is the anatomy:
Create a Lakeflow Declarative Pipeline that reads order event files
from my Volume at /Volumes/workspace/default/book_data/orders/
into a Bronze streaming table.
Clean the records into Silver: enforce valid order_ids, drop rows
with negative amounts, and deduplicate on order_id keeping the
latest event_time.
Aggregate daily revenue by region into a Gold table.
Then wrap this pipeline in a Lakeflow Job named daily_orders_job.
Run it in triggered mode every day at 06:00.
Retry twice on failure and notify the data team channel.
After the pipeline succeeds, run an OPTIMIZE on the Gold table.Read what that prompt actually contains. It is not code. It is a description of intent, written in plain language, covering the four things that make work production ready.
The data contract. Sources, targets, and the medallion layers between them.
The quality rules. What valid means, and what happens to rows that break the rules.
The schedule. When the work runs and in which mode.
The operating behavior. Retries, notifications, and what runs downstream.
When a prompt carries all four, Genie Code can generate both artifacts: the pipeline definition and the job definition that orchestrates it.
Two things appear, and it is worth understanding each one.
The first is the pipeline itself. Declarative tables or views, expectations expressed as constraints, and the flow from Bronze to Gold. If you have built these by hand, the shape will feel familiar from Chapter 7.6: Ingestion and ETL Pipelines, where we build the same pattern step by step in Free Edition.
The second is the job definition. This is the piece people used to click together by hand in the Jobs and Pipelines UI. Tasks, dependencies, schedule, compute choices, and alert rules, expressed as configuration you can review, version, and change.
A generated job for the prompt above looks roughly like this as a graph:
daily_orders_job (triggered, 06:00 daily)
|
v
run orders_pipeline (Bronze -> Silver -> Gold)
|
v
optimize gold_daily_revenue
|
v
notify data team on failure at any stepNothing here is magic. Every piece maps to something you can inspect and edit. That is the point. The agent drafts the wiring. You keep the judgment.
This is the one decision in the job definition that changes your cost and freshness the most, so it deserves its own moment.
Triggered mode is the default for batch thinking. The pipeline starts, processes what is new, and shuts down. A daily revenue table does not need to exist between runs. Triggered mode respects that, and your bill reflects it.
Continuous mode keeps the pipeline running. New records flow through in seconds to minutes. Fraud signals, live inventory, operational dashboards. When the business question cannot wait for the next schedule, continuous is the honest answer.
The trap is choosing continuous because it sounds better. It is not better. It is more expensive and more demanding to operate, and most tables in most companies are perfectly fine refreshed a few times a day. Pick the mode from the business need, not from the label.
If cost shape interests you, Databricks Serverless Compute: The End of Cluster Babysitting goes deeper on how serverless changes this math.
A single pipeline is rarely the whole story. Real workflows look like a small city of tasks.
The customers table refreshes first, because three pipelines read it. The orders pipeline runs next. Then the revenue aggregates. Then a data quality audit compares today against the trailing week. Only then does the Genie space that answers business questions get its refresh, so nobody asks a question against half loaded data.
In a Lakeflow Job, each of these is a task, and the dependencies between them are explicit. Task B simply declares that it depends on task A. The scheduler handles the rest, in parallel where it can, in sequence where it must.
This matters for trust. When dependencies live in a job definition instead of in someone''s memory, a new engineer can read the workflow like a map. And when something fails, the run history shows exactly which task broke and which ones never started.
Debugging those failures is its own skill. Chapter 17: Observability and Debugging teaches you to read run timelines and find the real bottleneck, and The dashboard was wrong before anyone knew shows what good alerting looks like in practice.
It is worth being precise about the new division of labor.
The agent is excellent at structure. Boilerplate, syntax, sensible defaults, consistent naming. The 60% number exists because this part of the work was always the slowest to type and the easiest to get subtly wrong.
You are still responsible for intent. Four reviews before anything generated reaches production.
First, the schedule. Does 06:00 match when the source data actually arrives? A pipeline that runs before its inputs land processes nothing and reports success.
Second, the dependencies. Does the generated order match the real relationships? An agent can chain tasks in the order you wrote them, not the order the business needs.
Third, the quality rules. Expectations are where trust is enforced. Read every one. Chapter 13: Data Quality Engineering gives you the patterns to judge them against.
Fourth, the failure behavior. Who gets notified, and is a retry the right response? Some failures should stop the line and page a human. The job ran twice explains why safe reruns decide whether a retry heals the pipeline or corrupts it.
This is the same theme that runs through What Happens When Two AI Agents Disagree?. Agents draft. Humans hold authority. The teams that get this balance right move fast without giving up control.
Everything below runs in Databricks Free Edition. No paid workspace needed.
Step 1. Create sample data. Open a notebook and write a small orders file to your Volume.
import csv
rows = [
["order_id", "region", "amount", "event_time"],
["o1", "west", 120.0, "2026-10-08 09:00:00"],
["o2", "east", 75.5, "2026-10-08 09:05:00"],
["o2", "east", 80.0, "2026-10-08 09:20:00"],
["o3", "west", -10.0, "2026-10-08 09:30:00"],
]
path = "/Volumes/workspace/default/book_data/orders/orders_day1.csv"
with open(path, "w", newline="") as f:
csv.writer(f).writerows(rows)Notice the two traps we planted. Order o2 arrives twice, and o3 has a negative amount. A good pipeline handles both without being told twice.
Step 2. Prompt for the pipeline. Open Genie Code in your workspace and paste the one prompt pattern from earlier, adjusted to these paths. Ask for Bronze, Silver with deduplication and a positive amount expectation, and Gold with daily revenue by region.
Step 3. Prompt for the job. In the same conversation, ask Genie Code to wrap the pipeline in a Lakeflow Job. Name it, set a daily 06:00 triggered schedule, two retries, and a downstream OPTIMIZE on the Gold table.
Step 4. Review before you run. Read the expectations. Check the schedule. Confirm the dependency order. This review habit is the whole game.
Step 5. Run and observe. Trigger the job manually from Jobs and Pipelines. Watch the task graph light up in order. Then open the Gold table and check that o2 appears once with the latest amount, and o3 is gone.
If you want to build this pattern by hand first so the generated version makes sense, Chapter 7.6: Ingestion and ETL Pipelines walks through it, and Chapter 19: Production Orchestration covers scheduling and dependencies in depth. Our guide to building a medallion pipeline in SQL is a good companion read.
Step back and the pattern is larger than one feature.
First, agents learned to answer questions about data. That is the Genie analyst story we covered in Databricks Genie Explained: The AI Analyst You Have to Onboard.
Then, agents learned to write the pipelines themselves. That is the 60%.
Now, agents are learning to operate those pipelines: schedule them, chain them, and watch them. The prompt in this article is an early, practical version of that.
Each step moves the human further up the stack, from typing code to reviewing intent to setting the rules agents operate within. If that direction interests you, The Future of Data to Decisions Is Autonomous explores where the road leads, and Genie Code can now convert your legacy SQL shows the same agentic pattern applied to migration work.
The engineers who thrive in this world are not the ones who type the fastest. They are the ones who can look at a generated pipeline and a generated job, and know in five minutes whether to trust them. That judgment comes from understanding the fundamentals the code is built on.
If you are newer to Databricks and want the foundations before the agents, start with the first free chapters of the book. Chapter 1: Start Here sets up your Free Edition workspace, and Chapter 2: Spark DataFrames teaches the data model every pipeline is built on.
When you are ready to go deeper, the full path through ingestion, quality, orchestration, and governance is in Thinking in Data Engineering with Databricks, written so every concept lands on something you can run yourself.
And if you want to practice designing pipelines before you generate them, the PySpark Pipeline Designer lets you sketch the architecture and see the code, which is exactly the skill that makes you a sharp reviewer of whatever an agent hands you next.
The tools will keep getting better at drafting. Your edge is knowing what good looks like.