How to build reliable file-based ingestion pipelines that handle schema changes, recover unexpected data, and scale to millions of files.