One small CSV, one question, and a pipeline that survives a rerun
It was a Tuesday evening. I had a laptop, a free Databricks workspace, and a small CSV file of 100 orders that a friend had exported from a hobby store. Nothing about it was impressive. That is exactly why it worked.
Most people who want to break into data engineering start by collecting things. A course. A certification syllabus. A list of twelve tools. Then the weeks pass and nothing is built, because a list is not a project.
So I gave myself a boundary instead of a plan. Two evenings. One small file. One question that the data should be able to answer at the end. Whatever I did not finish in two evenings would not get done.
Here is what happened, and how you can repeat it tonight.
The first mistake beginners make is cleaning too early. You open the CSV, you see a bad date, and you fix it in the file. Now your pipeline has a human step in the middle, and it can never run on its own.
I loaded the file untouched into a bronze table. Wrong types, duplicates, blank customer names, all of it. Bronze is a landing zone, not a promise of quality.
raw_orders = (
spark.read
.option("header", "true")
.csv("/Volumes/workspace/default/book_data/first_project_orders.csv")
)
raw_orders.write.mode("overwrite").saveAsTable("workspace.default.bronze_orders")Then I did the boring thing that turned out to be the most useful thing all evening. I counted.
print(raw_orders.count())
raw_orders.groupBy("order_status").count().show()100 rows in the file. 100 rows in the table. Four status values, one of them misspelled. That single check is the habit that separates someone who runs code from someone who runs a pipeline. If you want the reasoning behind this layered approach, the medallion architecture guide explains why bronze, silver and gold exist at all.
By the end of evening one I had a table, a row count, and a short list of problems written in a notebook cell. That is a real deliverable.
Evening two started with cleaning, and I only fixed the problems I had actually written down. Types cast properly. Duplicate order IDs removed. Rows with no customer or no amount pushed into a separate rejected table instead of being thrown away.
That rejected table mattered more than I expected. Six rows failed. Because I kept them, I could tell the story of why the silver count did not match bronze. Deleting bad rows silently is how trust in a pipeline dies.
from pyspark.sql import functions as F
typed_orders = (
spark.table("workspace.default.bronze_orders")
.withColumn("order_amount", F.col("order_amount").cast("double"))
.withColumn("order_date", F.to_date("order_date"))
)
clean_orders = typed_orders.filter(
F.col("customer_name").isNotNull() & F.col("order_amount").isNotNull()
).dropDuplicates(["order_id"])
rejected_orders = typed_orders.exceptAll(clean_orders)Then one question, not five. Revenue per day. One gold table, small enough to read with your eyes.
SELECT
order_date,
COUNT(*) AS order_count,
ROUND(SUM(order_amount), 2) AS daily_revenue
FROM workspace.default.silver_orders
GROUP BY order_date
ORDER BY order_dateThe last hour was the part almost no beginner project includes, and the part every interviewer asks about. I ran the whole thing twice.
Nothing broke, because I had used overwrite on the bronze load and MERGE for the update file instead of blind appends. Reruns are the real test of a pipeline. A pipeline that only works the first time is a script. The incremental processing lesson goes deeper on replaceWhere and idempotent writes if you want the full pattern.
Total time: about four hours across two evenings.
I have watched a lot of people study data engineering for months and still freeze when asked what they have built. The reason is that reading gives you vocabulary and building gives you judgment. Only one of those survives contact with a real table.
A small finished project gives you three things a course cannot:
You get a story. Not "I learned Spark" but "I loaded 100 messy orders, six failed validation, and here is how I kept them visible."
You get a rerun scar. You learn what happens when a job runs twice, because it happened to you at 10pm.
You get momentum. The next project is easier, because the fear is gone.
This is the whole idea behind how BricksNotes is written. Practice first, then the concept, then the reason it matters at work. It is also the argument at the centre of Thinking in Data Engineering with Databricks: the engineers who move fastest are not the ones who memorised the most features, they are the ones who understand how systems behave when something goes wrong. The first three lessons are free if you want to feel the style before deciding.
You do not have to find data or invent requirements. We packaged the exact project, including the messy 100 row CSV and a small update file for practising MERGE, as a guided walkthrough.
Six steps, all runnable in Databricks Free Edition, with a progress checklist so you can stop after step three and come back tomorrow.
Start here: Your first data engineering project in Databricks Free Edition.
If you have never opened a workspace before, do the Free Edition setup first. It takes about ten minutes and costs nothing.