<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Posts on Ghost in the data</title><link>https://ghostinthedata.info/posts/</link><description>Ghost in the data</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>Ghost in the data</copyright><lastBuildDate>Mon, 31 Aug 2026 09:00:00 +1000</lastBuildDate><atom:link href="https://ghostinthedata.info/posts/index.xml" rel="self" type="application/rss+xml"/><item><title>The Style Isn't the Skill</title><link>https://ghostinthedata.info/posts/2026/2026-08-31-style-isnt-the-skill/</link><pubDate>Mon, 31 Aug 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-08-31-style-isnt-the-skill/</guid><author>Chris Hillman</author><description>What Bruce Lee's messy fights, broke years and busted back teach data engineers about mastering a craft, and about your next self-evaluation.</description><content:encoded>&lt;p>I was a kid when I first watched Bruce Lee, soon after and my friend signed up for kung fu lessons.&lt;/p>
&lt;p>What hooked me wasn&amp;rsquo;t the fighting. It was the speed at which Bruce moved. The more you watched, though, the more you noticed something sitting behind the speed. He was obviously smart, fast, cool, and clearly thinking several moves ahead of everyone else in the room. Even as a kid you could tell the fighting was the visible bit of something much bigger.&lt;/p>
&lt;p>Looking back, the thing I admired was the confidence. Not swagger. This was different. It was the kind of confidence that comes off someone who has poured a frightening amount of themselves into one craft. You can see it on people in all areas of life. I&amp;rsquo;m sure you&amp;rsquo;ve seen it in your industry. It could be a senior data engineer walking a junior through a query plan without looking at the screen. A modeller who hears the grain of a new fact table and already knows which join is going to hurt. It radiates, and it can&amp;rsquo;t be faked for long.&lt;/p>
&lt;p>Then people will say&amp;hellip;&lt;/p>
&lt;p>&lt;em>Yeah, but he&amp;rsquo;s Bruce Lee.&lt;/em>&lt;/p>
&lt;p>You&amp;rsquo;ve probably said the same about someone like him. It sounds like humility, but it isn&amp;rsquo;t. It&amp;rsquo;s a permission slip — if he&amp;rsquo;s a different species, then the work he did to become that good is not available to me, and I&amp;rsquo;m excused from trying.&lt;/p>
&lt;p>The problem with this is that it never was easy for Bruce Lee, and I&amp;rsquo;m sure it&amp;rsquo;s the same for others as well. The three moments that made Bruce Lee happened before anyone outside a few gyms and a cancelled TV show knew who he was. He wasn&amp;rsquo;t Bruce Lee yet. He was a Wing Chun kid from Hong Kong teaching out of a friend&amp;rsquo;s house in Oakland, then a sidekick on a show that lasted one season, then a man in a back brace. Enter the Dragon came out weeks after he died. Everything that matters here is from before.&lt;/p>
&lt;/br>
&lt;hr>
&lt;h3 id="1964-the-messy-win">1964: the messy win&lt;/h3>
&lt;/br>
&lt;p>In late 1964 Lee fought Wong Jack Man in Oakland, behind closed doors, in front of a handful of people. There are disagreements as to how it went. Lee&amp;rsquo;s side says he finished it in about three minutes. Wong&amp;rsquo;s side says it dragged past twenty and ended with both men exhausted. What everyone agrees on is that it was ugly, and that Lee came out of it winded and rattled. For a man who had spent the previous year telling anyone who&amp;rsquo;d listen that his technique was superior to the &amp;ldquo;classical&amp;rdquo; styles, it was a humiliating result. He&amp;rsquo;d won, by his own account. It didn&amp;rsquo;t feel like winning.&lt;/p>
&lt;p>What he did next is the interesting part. He didn&amp;rsquo;t defend the technique. He tore it up, and hit the reset switch. That fight is the generally accepted origin of Jeet Kune Do, and the first thing he changed wasn&amp;rsquo;t his kung fu. It was his conditioning. He went looking for far more extreme fitness training than any traditional school would have asked of him.&lt;/p>
&lt;p>The line he used to explain the whole shift:&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;The proper method is the one that works.&amp;rdquo;&lt;/p>&lt;/blockquote>
&lt;p>As with most things that Bruce Lee said, there is so much in those few words. It&amp;rsquo;s easy to read it as &amp;ldquo;method doesn&amp;rsquo;t matter,&amp;rdquo; and that&amp;rsquo;s the opposite of what it says.&lt;/p>
&lt;p>It says there is a proper method. It says the test of whether a method is proper is a single question: does it work? And it says, quietly, that the answer can change. A method that worked last year against last year&amp;rsquo;s opponent gets re-tested, and if it fails, it isn&amp;rsquo;t proper anymore. Loyalty to it is not a virtue.&lt;/p>
&lt;p>He made the point a second time, in a way only Bruce Lee would. He put a mock tombstone in his studio. The inscription:&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;In memory of a once fluid man, crammed and distorted by the classical mess.&amp;rdquo;&lt;/p>&lt;/blockquote>
&lt;p>He wasn&amp;rsquo;t mourning kung fu. He was mourning himself. The &amp;ldquo;classical mess&amp;rdquo; was the ceremony of the art mistaken for the art: the numbered forms, the fixed stances, the correct sequence of steps that everyone agreed was correct because everyone had always agreed. He&amp;rsquo;d drilled it for years. And in a real fight, against an opponent who didn&amp;rsquo;t know the choreography, it had crammed and distorted him.&lt;/p>
&lt;p>On one side of the spectrum there might be a template or process that was set up. Does it work? Or is the only answer &amp;ldquo;it&amp;rsquo;s how we&amp;rsquo;ve always done it&amp;rdquo;? Sometimes we don&amp;rsquo;t spend the time reflecting — it might have worked perfectly initially, but is it as effective now? There&amp;rsquo;s also the other side, where we make exceptions to the process far too often, to the point it becomes a mess and we no longer know what the technique even is.&lt;/p>
&lt;p>Bruce was ruthless about fixed &lt;em>form&lt;/em>. He was obsessive about &lt;em>fundamentals&lt;/em>. His response to a bad fight was more discipline, not less. The freedom he&amp;rsquo;s famous for was built on top of conditioning most people wouldn&amp;rsquo;t tolerate.&lt;/p>
&lt;p>The style isn&amp;rsquo;t the skill. That&amp;rsquo;s the whole thing.&lt;/p>
&lt;/br>
&lt;hr>
&lt;h3 id="one-kick-ten-thousand-times">One kick, ten thousand times&lt;/h3>
&lt;/br>
&lt;p>One of Bruce&amp;rsquo;s quotes was:&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;I fear not the man who has practiced ten thousand kicks once, but I fear the man who has practiced one kick ten thousand times.&amp;rdquo;&lt;/p>&lt;/blockquote>
&lt;p>Ten thousand kicks once is a career spent collecting — a new tool every quarter, every certification going, every architecture pattern from every conference talk. You can talk about all of them. You&amp;rsquo;ve done each of them once, in a tutorial, on a dataset that was designed to make the tutorial work.&lt;/p>
&lt;p>One kick ten thousand times is mastery of the skill. It isn&amp;rsquo;t knowing the technique. It&amp;rsquo;s having met every way the technique fails. I see this a lot meeting people with various titles: have they actually dealt with quality incidents, and how did they protect the data from it? Have they worked with a source system that offers deltas, and tested them thoroughly? The engineer who&amp;rsquo;s built them fifty times knows what happens when the source system back-dates a change, when two updates land in the same batch with the same timestamp, when the current flag and the high date disagree, and when a backfill quietly rewrites six months of history because someone ran it day by day. (I wrote &lt;a href="https://ghostinthedata.info/posts/2026/2026-02-07-self-healing/" target="_blank" rel="noopener">a whole post about that last one&lt;/a>, because it took me an embarrassing number of reps to see it coming.)&lt;/p>
&lt;p>So pick the kick. Dimensional modelling. Streaming ingestion. Query performance on the warehouse you actually run. Pipeline testing. Not all of them. One, then maybe a second. The industry will keep offering you ten thousand shiny new ones, and most of them will be gone in three years. However, honing these skills is what shines out in interviews I conduct — I can see the difference between someone who knows how to kick, versus someone who has done that kick thousands of times.&lt;/p>
&lt;/br>
&lt;hr>
&lt;h3 id="absorb-discard-add">Absorb, discard, add&lt;/h3>
&lt;/br>
&lt;p>The quote that describes how Lee built Jeet Kune Do out of Wing Chun, Western boxing, fencing, and whatever else he could get his hands on:&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;Absorb what is useful, discard what is not, add what is uniquely your own.&amp;rdquo;&lt;/p>&lt;/blockquote>
&lt;p>&lt;em>Absorb what is useful.&lt;/em> Not &amp;ldquo;learn everything.&amp;rdquo; Absorb implies digestion. You take the thing in and it becomes part of how you move, without needing to consult it. Kimball&amp;rsquo;s grain-first modelling is worth absorbing. So are idempotent loads, and the habit of writing the test before the transformation. Once absorbed, you stop noticing you&amp;rsquo;re doing them.&lt;/p>
&lt;p>&lt;em>Discard what is not.&lt;/em> Harder than it sounds, because discarding means admitting you spent time on something that didn&amp;rsquo;t pay off. The pattern that made sense at your last company and doesn&amp;rsquo;t fit this one. The tool you got certified in that your current platform doesn&amp;rsquo;t run. Lee walked away from the style his own master taught him. Most of us can&amp;rsquo;t walk away from a Confluence page.&lt;/p>
&lt;p>&lt;em>Add what is uniquely your own.&lt;/em> What have you built that isn&amp;rsquo;t in any documentation? The reconciliation check you wrote because the standard tests kept missing one specific failure. The way you explain a lineage problem to finance so they stop asking the same question every month. That&amp;rsquo;s yours. It&amp;rsquo;s the part nobody can hire off the shelf. If your self-evaluation lists things you absorbed and nothing you added, you&amp;rsquo;ve described a good student, not a good engineer.&lt;/p>
&lt;/br>
&lt;hr>
&lt;h3 id="water-and-the-shape-of-the-cup">Water and the shape of the cup&lt;/h3>
&lt;/br>
&lt;blockquote>
&lt;p>&amp;ldquo;You put water in a cup, it becomes the cup.&amp;rdquo;&lt;/p>&lt;/blockquote>
&lt;p>Water has no form of its own; it takes the form of whatever holds it. Bruce also points out that water can flow or it can crash. Formless isn&amp;rsquo;t the same as weak.&lt;/p>
&lt;p>For a data modeller this is close to a job description. The model takes the shape of the question, not the shape of the framework diagram. A star schema is right until the business hands you a many-to-many relationship that every consumer needs, and then it&amp;rsquo;s wrong, and the modeller who is loyal to the star instead of the question will build something the analysts route around within a month. I&amp;rsquo;ve argued before that there&amp;rsquo;s &lt;a href="https://ghostinthedata.info/posts/2026/2026-05-02-five-worlds-data-engineering/" target="_blank" rel="noopener">no single right way to do data engineering&lt;/a>, only the right way for the world you&amp;rsquo;re in. Lee got there fifty years earlier with the motto he put on the Jeet Kune Do emblem:&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;Using no way as way, having no limitation as limitation.&amp;rdquo;&lt;/p>&lt;/blockquote>
&lt;p>No way as way doesn&amp;rsquo;t mean no method. It means the method is chosen fresh for each fight. Kimball for the reporting layer, something flatter and wider for the feature store, streaming where latency is the product and batch everywhere else, with no shame in any of it. The limitation you refuse to accept is the one that says &amp;ldquo;we&amp;rsquo;re a Kimball shop&amp;rdquo; or &amp;ldquo;we don&amp;rsquo;t do streaming here&amp;rdquo; before anyone has looked at the problem.&lt;/p>
&lt;/br>
&lt;hr>
&lt;h3 id="1969-the-letter">1969: the letter&lt;/h3>
&lt;/br>
&lt;p>January 1969. Lee is twenty-eight. The Green Hornet had come and gone after one season. He&amp;rsquo;s picking up bit parts and teaching private lessons to actors to pay the bills. Linda is pregnant with their second child. Bruce&amp;rsquo;s Hollywood career has stalled.&lt;/p>
&lt;p>He sits down and writes himself a letter. It&amp;rsquo;s on a sheet of notepaper headed &amp;ldquo;Secret,&amp;rdquo; in red and blue ink, and he titles it &amp;ldquo;My Definite Chief Aim.&amp;rdquo; The whole thing is four sentences. In the first he declares that he will be the first highest paid Oriental superstar in the United States (his words, and the era&amp;rsquo;s). The last two set a dated target and a personal one: world fame from 1970, ten million dollars by the end of 1980, and a life lived the way he pleases.&lt;/p>
&lt;p>The sentence in the middle is the one that matters most:&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;In return I will give the most exciting performances&amp;hellip;&amp;rdquo;&lt;/p>&lt;/blockquote>
&lt;p>This isn&amp;rsquo;t a wish list. It&amp;rsquo;s a contract — here&amp;rsquo;s what I will be, here&amp;rsquo;s what I will give in return.&lt;/p>
&lt;p>I&amp;rsquo;ve read enough self-evaluations to notice a pattern. Most of them are the first sentence without the second. &amp;ldquo;I want to grow into a senior role.&amp;rdquo; &amp;ldquo;I&amp;rsquo;d like more exposure to the streaming work.&amp;rdquo; &amp;ldquo;I&amp;rsquo;m keen to lead a project next year.&amp;rdquo; All fine. All wishes. The &amp;ldquo;in return&amp;rdquo; clause is missing, and without it there&amp;rsquo;s nothing for a manager to hold onto and nothing for you to hold yourself to.&lt;/p>
&lt;p>Try writing it Lee&amp;rsquo;s way. Something like: &lt;em>By the end of June I will be the person this team hands streaming ingestion to without a second thought. In return I will have built late-arrival handling on both event pipelines, documented the watermark strategy so someone else can run it, and carried the on-call for it the first three times it breaks.&lt;/em>&lt;/p>
&lt;p>That&amp;rsquo;s a sentence a manager can act on. It&amp;rsquo;s also a sentence that scares you slightly to write, which is how you know it&amp;rsquo;s the real one. (He achieved everything in the letter, incidentally, and then died at thirty-two.)&lt;/p>
&lt;/br>
&lt;hr>
&lt;h3 id="in-the-review-room">In the review room&lt;/h3>
&lt;/br>
&lt;p>When I sit with someone to go through their self-evaluation, I&amp;rsquo;ve started asking myself a slightly ridiculous question: what would Bruce Lee do with this? Ridiculous or not, it changes what I look for.&lt;/p>
&lt;p>&lt;strong>If you&amp;rsquo;re writing the evaluation.&lt;/strong>&lt;/p>
&lt;p>Write the &amp;ldquo;in return&amp;rdquo; clause. Every aim gets a consideration.&lt;/p>
&lt;p>Name your kick. What&amp;rsquo;s the one thing you&amp;rsquo;ve done enough times that you know how it fails? If you can&amp;rsquo;t name one, that&amp;rsquo;s the finding, and it&amp;rsquo;s a useful one.&lt;/p>
&lt;p>Include something you added, not just something you absorbed. The thing you built that wasn&amp;rsquo;t in the docs.&lt;/p>
&lt;p>&lt;strong>If you&amp;rsquo;re reading it.&lt;/strong>&lt;/p>
&lt;p>Judge what works, not adherence to form. The engineer who follows every ceremony and ships fragile pipelines has the classical mess. Say so, kindly and specifically.&lt;/p>
&lt;p>But check the reps. &amp;ldquo;Judge outcomes&amp;rdquo; drifts very easily into &amp;ldquo;reward whoever got lucky.&amp;rdquo; Someone who has practised one kick ten thousand times has a track record of predictable outcomes, and that&amp;rsquo;s the difference between skill and a good quarter.&lt;/p>
&lt;p>Look for the gap between knowing and doing. Lee was fond of a line he borrowed from Goethe: knowing is not enough, we must apply; willing is not enough, we must do. Most people in a review know exactly what they should be doing. The evaluation is a measure of the distance between that and what happened.&lt;/p>
&lt;p>If you want a structured way to run that kind of exercise with your own team, I&amp;rsquo;ve written before about the &lt;a href="https://ghostinthedata.info/posts/2025/2025-11-01-discover-our-best-selves/" target="_blank" rel="noopener">Reflective Best Self process&lt;/a>.&lt;/p>
&lt;/br>
&lt;hr>
&lt;h3 id="the-tradeoff">The tradeoff&lt;/h3>
&lt;/br>
&lt;p>Judging by what works is harder than judging by ceremony. Ceremony is easy to audit: the template was filled in, the points were estimated, the meeting was attended. Skill is hard to audit, because you have to understand the work well enough to know whether the outcome was earned or lucky, and that takes time managers don&amp;rsquo;t have and technical depth some of them stopped maintaining years ago. Take the Lee approach and you sign up for that cost.&lt;/p>
&lt;p>There&amp;rsquo;s a cost on the other side too. A self-evaluation written without the mask, with a real &amp;ldquo;in return&amp;rdquo; clause and a named weakness, is a sharper tool for your own growth and a sharper tool in the wrong hands. In a healthy team it gets you coached. In an unhealthy one it gets quoted back at you.&lt;/p>
&lt;p>&lt;/br>&lt;/br>&lt;/p></content:encoded><category>Leadership</category><category>Career Development</category><category>Leadership</category><category>Career Development</category><category>Performance Reviews</category><category>Self Evaluation</category><category>Data Modelling</category><category>Streaming</category><category>Skills</category><category>Philosophy</category></item><item><title>AWS MWAA: The Practitioner's Unvarnished Guide</title><link>https://ghostinthedata.info/posts/2026/2026-07-18-mwaa-unvarnished-guide/</link><pubDate>Sat, 18 Jul 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-07-18-mwaa-unvarnished-guide/</guid><author>Chris Hillman</author><description>Everything data engineers need to know about AWS MWAA — real costs, operational gotchas, and head-to-head comparisons with Astronomer, Dagster, Prefect, Step Functions, and self-hosted Airflow.</description><content:encoded>&lt;p>I remember the first time I sat down properly with the Step Functions architecture our team inherited. Someone had designed it thoughtfully — real engineering thought had gone into it when it was built.&lt;/p>
&lt;p>But what I was looking at in the console was a map of state machines and Lambda functions that had grown well past the point where anyone could hold it in their head.&lt;/p>
&lt;p>The first production issue we had to debug on that setup took most of a working day. Not because the problem was complicated — it wasn&amp;rsquo;t — but because finding it meant jumping between CloudWatch log groups, correlating timestamps, building a mental picture of which Lambda had received what input and where the chain had broken. Every time I thought I had it, another log group. Another timestamp comparison. Another dead end.&lt;/p>
&lt;p>That&amp;rsquo;s when I suggested to the team that we test out Airflow. I remember the first time something failed on the new setup. I clicked into the task, read the log attached to that specific run, saw the error, fixed it, and restarted the DAG. Four minutes. The CDK stacks for the old architecture ran to thousands of lines. The equivalent DAG was fifty to two hundred. Same logic. Made legible.&lt;/p>
&lt;p>The difference wasn&amp;rsquo;t performance — Lambda and Airflow are comparable where it counts. It was visibility. One place to look. That&amp;rsquo;s what orchestration is actually about: not execution speed, not infrastructure elegance, but whether you can understand what&amp;rsquo;s happening in your data platform — at 9am in a post-mortem, or at 2am when something breaks and you need to know why.&lt;/p>
&lt;p>AWS MWAA — Managed Workflows for Apache Airflow — is Amazon&amp;rsquo;s answer to this problem. And it&amp;rsquo;s genuinely good at some things. But the gap between what the marketing page promises and what you&amp;rsquo;ll actually experience in production is wider than most managed services I&amp;rsquo;ve worked with. The $365/month sticker price? That&amp;rsquo;s just the beginning. The &amp;ldquo;managed&amp;rdquo; part? It manages less than you&amp;rsquo;d think.&lt;/p>
&lt;p>This guide covers everything I wish someone had told me before committing to MWAA — the real architecture, the actual costs (including the ones AWS conveniently leaves off the pricing page), the operational gotchas that&amp;rsquo;ll bite you in production, and honest head-to-head comparisons with every serious alternative: Astronomer Astro, Dagster Cloud, Prefect Cloud, Step Functions, and the self-hosted approach that might be smarter than you&amp;rsquo;d expect.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="how-mwaa-actually-works-under-the-hood">How MWAA Actually Works Under the Hood&lt;/h3>
&lt;br>
&lt;p>MWAA runs on a &lt;strong>dual-VPC architecture&lt;/strong> that&amp;rsquo;s important to understand because it explains most of the service&amp;rsquo;s constraints.&lt;/p>
&lt;p>AWS provisions Fargate containers for your schedulers, workers, and web server inside an AWS-managed ECS cluster. These containers connect to your VPC through Elastic Network Interfaces injected into your private subnets. A single-tenant &lt;strong>Aurora PostgreSQL&lt;/strong> instance stores the Airflow metadata — fully managed by AWS, but completely inaccessible to you. No connection string. No direct SQL queries. No custom dashboards built against the metadata database. If you&amp;rsquo;ve ever relied on querying &lt;code>task_instance&lt;/code> or &lt;code>dag_run&lt;/code> tables directly for custom monitoring, MWAA takes that away.&lt;/p>
&lt;p>The networking requirements are specific and unforgiving. You need two private subnets in different Availability Zones, a NAT gateway (or VPC endpoints) for internet access, and a self-referencing security group. The VPC you choose at creation is permanent — change your mind later and you&amp;rsquo;re building a new environment from scratch. Choose private webserver mode (which most production teams do for security), and accessing the Airflow UI requires a VPN or bastion host setup.&lt;/p>
&lt;p>DAGs live in an S3 bucket with versioning enabled. MWAA syncs files to the workers every 30 seconds, though new DAG files take roughly five minutes to be recognised (controlled by &lt;code>dag_dir_list_interval&lt;/code>). Updates to &lt;code>requirements.txt&lt;/code> or &lt;code>plugins.zip&lt;/code> trigger a full environment update — and that means 20 to 40 minutes of waiting while MWAA reprovisions everything. More on that particular joy later.&lt;/p>
&lt;p>A few setup details that&amp;rsquo;ll save you time: the S3 bucket must be in the same region as your MWAA environment and must have block public access enabled. Your &lt;code>dags/&lt;/code> folder sits at the root of the bucket, with &lt;code>requirements.txt&lt;/code> alongside it (not inside the &lt;code>dags/&lt;/code> folder). The &lt;code>plugins.zip&lt;/code> file follows a strict structure — if your custom operator lives at &lt;code>plugins/operators/my_operator.py&lt;/code>, the zip must mirror that path exactly. Get the structure wrong and the import will silently fail with no useful error message until you dig into the DAGProcessing logs.&lt;/p>
&lt;p>For IAM, MWAA needs an execution role with access to your S3 bucket, CloudWatch Logs, SQS (for the Celery broker), and whatever AWS services your DAGs interact with — Glue, EMR, Redshift, Lambda, whatever. The principle of least privilege matters here, but AWS&amp;rsquo;s own documentation starts with fairly broad permissions and leaves tightening to you. In practice, most teams start broad and restrict after stabilising.&lt;/p>
&lt;br>
&lt;p>&lt;strong>Environment classes&lt;/strong> span six tiers since the 2024 additions, from micro through 2xlarge:&lt;/p>
&lt;table>
 &lt;thead>
 &lt;tr>
 &lt;th>Class&lt;/th>
 &lt;th>vCPU / Memory per Worker&lt;/th>
 &lt;th>Concurrent Tasks/Worker&lt;/th>
 &lt;th>Approx. DAG Capacity&lt;/th>
 &lt;th>Base Cost/hr&lt;/th>
 &lt;/tr>
 &lt;/thead>
 &lt;tbody>
 &lt;tr>
 &lt;td>mw1.micro&lt;/td>
 &lt;td>Combined scheduler/worker: 1 vCPU, 3 GB&lt;/td>
 &lt;td>3&lt;/td>
 &lt;td>~25&lt;/td>
 &lt;td>~$0.05&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>mw1.small&lt;/td>
 &lt;td>1 vCPU, 2 GB&lt;/td>
 &lt;td>5&lt;/td>
 &lt;td>~50&lt;/td>
 &lt;td>$0.49&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>mw1.medium&lt;/td>
 &lt;td>2 vCPU, 4 GB&lt;/td>
 &lt;td>10&lt;/td>
 &lt;td>~250&lt;/td>
 &lt;td>~$0.74&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>mw1.large&lt;/td>
 &lt;td>4 vCPU, 8 GB&lt;/td>
 &lt;td>20&lt;/td>
 &lt;td>~1,000&lt;/td>
 &lt;td>$0.99&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>mw1.xlarge&lt;/td>
 &lt;td>8 vCPU, 24 GB&lt;/td>
 &lt;td>40&lt;/td>
 &lt;td>~2,000&lt;/td>
 &lt;td>~$1.49&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>mw1.2xlarge&lt;/td>
 &lt;td>16 vCPU, 48 GB&lt;/td>
 &lt;td>80&lt;/td>
 &lt;td>~4,000&lt;/td>
 &lt;td>~$1.99&lt;/td>
 &lt;/tr>
 &lt;/tbody>
&lt;/table>
&lt;p>The micro class (November 2024) is a genuine game-changer for dev/test — at roughly $37/month, you can finally have a development MWAA environment without the $365+ price floor. The catch: it combines the scheduler and worker into a single container with no autoscaling. Fine for testing DAGs, not for production.&lt;/p>
&lt;p>Worker autoscaling follows a straightforward formula: &lt;code>(running + queued tasks) / (tasks per worker) = required workers&lt;/code>. Scale-down triggers after running and queued tasks hit zero for more than two minutes, and takes two to five minutes to complete. A significant autoscaling bug — where tasks could be assigned to workers being decommissioned — persisted until May 2025. If you&amp;rsquo;re running a version from before that fix, you&amp;rsquo;ll want to upgrade.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-true-cost-of-mwaa-its-not-365-a-month">The True Cost of MWAA (It&amp;rsquo;s Not $365 a Month)&lt;/h3>
&lt;br>
&lt;p>The AWS pricing page says mw1.small costs $0.49/hr, which works out to $364.56/month. That number is misleading. Every production MWAA deployment carries infrastructure costs that the headline figure conveniently excludes.&lt;/p>
&lt;p>&lt;strong>NAT gateways are the number one hidden cost.&lt;/strong> The standard architecture requires two NAT gateways (one per AZ), costing $0.045/hr each — that&amp;rsquo;s $65.70/month before a single byte of data moves through them. Add $0.045/GB for data processing. One team documented a $12K monthly bill where 87% was unused NAT gateway costs from misconfigured routing. This is not an MWAA-specific problem, but MWAA forces you into it because of the VPC requirement.&lt;/p>
&lt;p>&lt;strong>CloudWatch metrics are the second surprise.&lt;/strong> MWAA auto-publishes per-DAG and per-task metrics, and one practitioner documented $272/month for custom metrics alone — 907 metrics at $0.30 per metric. Unless you actively restrict what gets published using &lt;code>metrics.statsd_allow_list&lt;/code>, you&amp;rsquo;ll see this line item growing steadily as you add more DAGs.&lt;/p>
&lt;p>&lt;strong>CloudWatch Logs&lt;/strong> at INFO level with active workflows adds $50–300/month in ingestion and storage charges, depending on how chatty your DAGs are.&lt;/p>
&lt;p>Here&amp;rsquo;s what a small team running a production mw1.small environment actually pays:&lt;/p>
&lt;table>
 &lt;thead>
 &lt;tr>
 &lt;th>Component&lt;/th>
 &lt;th>Monthly Cost&lt;/th>
 &lt;/tr>
 &lt;/thead>
 &lt;tbody>
 &lt;tr>
 &lt;td>mw1.small environment (24/7)&lt;/td>
 &lt;td>$364.56&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>2 NAT gateways (hourly)&lt;/td>
 &lt;td>$65.70&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>NAT data processing (~20 GB)&lt;/td>
 &lt;td>$0.90&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>CloudWatch metrics&lt;/td>
 &lt;td>$30–100&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>CloudWatch Logs (~2 GB)&lt;/td>
 &lt;td>$1.00&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Elastic IPs (2)&lt;/td>
 &lt;td>$7.30&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>S3 + metadata storage&lt;/td>
 &lt;td>$1.50&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>Total&lt;/strong>&lt;/td>
 &lt;td>&lt;strong>$470–540&lt;/strong>&lt;/td>
 &lt;/tr>
 &lt;/tbody>
&lt;/table>
&lt;p>Scale up to mw1.large with additional workers running six hours a day, extra schedulers, and full logging, and you&amp;rsquo;re looking at $2,150–2,300/month all-in. AWS&amp;rsquo;s own pricing example for a comparable setup quotes $1,047/month — but conveniently excludes NAT, CloudWatch, and networking overhead.&lt;/p>
&lt;p>Here&amp;rsquo;s that breakdown for a heavier workload — say a mid-size data team running 200+ DAGs with burst periods:&lt;/p>
&lt;table>
 &lt;thead>
 &lt;tr>
 &lt;th>Component&lt;/th>
 &lt;th>Monthly Cost&lt;/th>
 &lt;/tr>
 &lt;/thead>
 &lt;tbody>
 &lt;tr>
 &lt;td>mw1.large environment (24/7)&lt;/td>
 &lt;td>$733.68&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>19 additional workers (6 hrs/day avg)&lt;/td>
 &lt;td>$672.00&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>5 schedulers&lt;/td>
 &lt;td>$164.00&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>2 NAT gateways (hourly)&lt;/td>
 &lt;td>$65.70&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>NAT data processing (~100 GB)&lt;/td>
 &lt;td>$4.50&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>CloudWatch metrics (400+)&lt;/td>
 &lt;td>$120.00&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>CloudWatch Logs (~10 GB)&lt;/td>
 &lt;td>$5.03&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Elastic IPs, S3, misc.&lt;/td>
 &lt;td>$10.00&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>Total&lt;/strong>&lt;/td>
 &lt;td>&lt;strong>~$1,775–2,300&lt;/strong>&lt;/td>
 &lt;/tr>
 &lt;/tbody>
&lt;/table>
&lt;p>The range depends heavily on how well you control CloudWatch metric sprawl and how long your burst periods actually last. The point is: you need to model the &lt;em>real&lt;/em> infrastructure cost, not just the MWAA line item.&lt;/p>
&lt;br>
&lt;p>&lt;strong>Cost optimisation levers do exist.&lt;/strong> Replace NAT gateways with VPC endpoints — the S3 Gateway Endpoint is free, and interface endpoints process data at $0.01/GB versus NAT&amp;rsquo;s $0.045/GB. Restrict CloudWatch metrics with &lt;code>metrics.statsd_allow_list = scheduler,executor&lt;/code> or disable custom metrics entirely. Set log retention to 30–90 days instead of indefinite. For dev/test environments, implement pause/resume schedules — one team reported 70% savings by running MWAA only during business hours.&lt;/p>
&lt;p>The November 2025 launch of &lt;strong>MWAA Serverless&lt;/strong> fundamentally changes the cost equation for intermittent workloads. At $0.08/hr per task (billed per-second, minimum one minute), 2,000 tasks averaging one to two minutes costs roughly $4/month — versus $365+ for the always-on equivalent. The trade-off is substantial though: no Airflow UI, YAML-based workflow definitions, and only 80-odd AWS operators supported (no PythonOperator, no BashOperator). It&amp;rsquo;s better understood as a different product targeting a different use case than as a cheaper MWAA.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-gotchas">The Gotchas&lt;/h3>
&lt;br>
&lt;p>Practitioners consistently flag the same pain points, and understanding them before committing is worth more than any architecture diagram.&lt;/p>
&lt;p>&lt;strong>Dependency management is the highest-risk operational area.&lt;/strong> A bad &lt;code>requirements.txt&lt;/code> can crash-loop your entire environment for hours. Not minutes — hours. The constraint statement is mandatory since Airflow 2.7.2, and omitting it risks catastrophic dependency conflicts. If pip install exceeds ten minutes, the Fargate task times out and rolls back — leaving your environment stuck in an update cycle. Certain packages can break MWAA&amp;rsquo;s CloudWatch logging by overriding the &lt;code>watchtower&lt;/code> library. The only safe approach: always test with the &lt;code>aws-mwaa-local-runner&lt;/code> Docker tool before deploying dependencies to production.&lt;/p>
&lt;p>&lt;strong>Environment update times frustrate everyone.&lt;/strong> Creating an environment takes 25–30 minutes. Updating requirements or plugins: 20–40 minutes. Version upgrades: up to two hours with the environment unavailable during the process. The May 2025 &amp;ldquo;graceful updates&amp;rdquo; feature helps — it replaces components without interrupting running tasks — but the wall-clock time for the update itself hasn&amp;rsquo;t changed meaningfully.&lt;/p>
&lt;p>&lt;strong>Observability is CloudWatch-only, and navigating it is painful.&lt;/strong> MWAA creates five separate CloudWatch log groups per environment: DAGProcessing, Scheduler, Task, WebServer, and Worker. Correlating a single pipeline failure requires jumping between log groups, and CloudWatch Logs Insights queries aren&amp;rsquo;t intuitive for Airflow-style troubleshooting. You can&amp;rsquo;t redirect StatsD metrics to external tools like Datadog or Prometheus — MWAA overrides the &lt;code>statsd_host&lt;/code> configuration. External monitoring integration requires CloudWatch Logs subscriptions via Kinesis Data Firehose, adding both complexity and cost.&lt;/p>
&lt;p>If you&amp;rsquo;ve spent time in the trenches of AWS logs — and maybe this is just me — it can be a genuine nightmare to piece together what happened across services. MWAA should improve this, and in some ways it does (the Airflow UI&amp;rsquo;s task logs are excellent). But the moment you need infrastructure-level debugging, you&amp;rsquo;re back in CloudWatch hell.&lt;/p>
&lt;p>&lt;strong>The 12-hour task execution limit is a hard wall.&lt;/strong> Tasks exceeding 12 hours are killed and get stuck in the SQS queue for another 12 hours due to SQS timeout settings. The CeleryExecutor-only constraint means no per-task Docker images and no pod-level resource isolation. If you need tasks that run longer than 12 hours, you&amp;rsquo;ll need to offload compute to ECS Fargate or Lambda, using MWAA purely for orchestration.&lt;/p>
&lt;p>&lt;strong>Secrets Manager integration works but generates surprising API volume.&lt;/strong> Without configuring &lt;code>connections_lookup_pattern&lt;/code>, Airflow attempts to look up every connection in Secrets Manager — including non-existent ones. One practitioner reported 100,000+ error API calls per day per MWAA instance. Secrets stored in Secrets Manager don&amp;rsquo;t appear in the Airflow UI either, which only shows connections from its own backend.&lt;/p>
&lt;p>Other gotchas worth noting: you can&amp;rsquo;t remove &lt;code>plugins.zip&lt;/code> or &lt;code>requirements.txt&lt;/code> once added (you can only point to empty files), the VPC can&amp;rsquo;t be changed after creation, frequent DAG updates to the S3 folder can break the MWAA installation, and if VPC endpoints are accidentally deleted the environment is broken and must be entirely recreated.&lt;/p>
&lt;p>One more that catches people off guard: &lt;strong>Airflow configuration overrides&lt;/strong>. MWAA lets you set most Airflow configuration options through the console or API, but several are explicitly blocked — &lt;code>statsd_host&lt;/code>, &lt;code>statsd_port&lt;/code>, &lt;code>broker_url&lt;/code>, &lt;code>result_backend&lt;/code>, and anything related to the metadata database connection. These are locked down because MWAA manages those components. If your Airflow experience includes tuning Celery broker settings or metadata database connection pools, you&amp;rsquo;ll need to adjust your expectations about what &amp;ldquo;managed&amp;rdquo; means here. It means AWS decides those values, and some of them (particularly the default Celery broker configuration) aren&amp;rsquo;t optimal for all workload patterns.&lt;/p>
&lt;p>The practical impact: if you&amp;rsquo;ve tuned self-hosted Airflow for specific performance characteristics — aggressive task polling, custom broker acknowledgement settings, metadata database query timeouts — MWAA won&amp;rsquo;t let you replicate that tuning. For most teams this doesn&amp;rsquo;t matter. For teams at scale with specific performance requirements, it can be a dealbreaker.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="head-to-head-astronomer-astro">Head-to-Head: Astronomer Astro&lt;/h3>
&lt;br>
&lt;p>Astronomer&amp;rsquo;s Astro platform runs Apache Airflow underneath, so your DAGs, operators, and muscle memory all transfer. The differences are in the operational layer — and they&amp;rsquo;re significant.&lt;/p>
&lt;p>&lt;strong>Workers scale to zero.&lt;/strong> When no tasks are running, you pay nothing for compute. MWAA charges for the environment 24/7 regardless. Astro&amp;rsquo;s base pricing starts at $0.35/hr for the control plane, with worker costs at $0.13/hr that only accrue when tasks actually execute.&lt;/p>
&lt;p>&lt;strong>The developer experience gap is substantial.&lt;/strong> Astro CLI provides a full local Airflow environment with &lt;code>astro dev start&lt;/code> that mirrors production exactly — same Airflow version, same Python version, same provider packages. MWAA has no equivalent. The &lt;code>aws-mwaa-local-runner&lt;/code> Docker tool is helpful but doesn&amp;rsquo;t perfectly replicate the MWAA environment. DAG deployment via &lt;code>astro deploy&lt;/code> takes 5–10 minutes through a proper CI/CD pipeline; MWAA&amp;rsquo;s S3 upload plus sync cycle is slower and less integrated into developer workflows.&lt;/p>
&lt;p>&lt;strong>Astro supports the KubernetesExecutor natively&lt;/strong>, meaning you get per-task Docker images, custom resource limits, and the thin orchestrator pattern described below — without managing Kubernetes yourself. MWAA is locked to CeleryExecutor with no workaround.&lt;/p>
&lt;p>Campspot, which migrated from MWAA to Astro, reported that a critical nightly job went from over two hours to two to three minutes. Another company, Black Crow AI, recouped roughly 20% more engineering time after leaving MWAA, citing zombie task issues and the absence of Airflow-specific support.&lt;/p>
&lt;p>Astro&amp;rsquo;s &lt;strong>Observe&lt;/strong> feature provides built-in pipeline lineage, SLA tracking, and data quality monitoring — capabilities absent from MWAA. Support comes from actual Airflow core committers, versus MWAA&amp;rsquo;s generic AWS support channels.&lt;/p>
&lt;p>The migration story from MWAA to Astro is worth paying attention to because it reveals what teams actually struggle with. Campspot completed their migration in a two-week sprint — their DAGs transferred with minimal modification because both platforms run Airflow. The performance improvement came not from different code, but from better infrastructure: KubernetesExecutor allowing parallel task execution in isolated pods versus CeleryExecutor bottlenecking through shared workers. Black Crow AI&amp;rsquo;s experience was similar — the DAGs were the same, but the operational layer (scaling, monitoring, deployment speed) was categorically better.&lt;/p>
&lt;p>That said, Astronomer is a startup, not AWS. If your organisation&amp;rsquo;s procurement process requires a vendor that&amp;rsquo;s been around for 20+ years, has SOC 2 Type II (Astronomer does have this, to be fair), and won&amp;rsquo;t disappear if funding dries up — that&amp;rsquo;s a valid concern. Astronomer raised significant funding and is growing, but the risk calculus is different from buying from AWS.&lt;/p>
&lt;p>The trade-offs: Astro&amp;rsquo;s dedicated clusters cost $2.40/hr ($1,782/month at the cluster level), and networking costs (NAT, PrivateLink) are passed through from your cloud provider. For teams deeply embedded in AWS wanting consolidated billing and native IAM integration without multi-cloud requirements, MWAA&amp;rsquo;s tight ecosystem coupling has genuine value. But if developer experience and operational flexibility are priorities, Astro is meaningfully ahead.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="head-to-head-dagster-cloud">Head-to-Head: Dagster Cloud&lt;/h3>
&lt;br>
&lt;p>Dagster represents the most fundamental philosophical departure from the Airflow model — and by extension, from MWAA.&lt;/p>
&lt;p>Where Airflow (and MWAA) asks &amp;ldquo;which tasks ran successfully?&amp;rdquo;, Dagster asks &amp;ldquo;which data assets are fresh, and why?&amp;rdquo; The difference sounds academic until you&amp;rsquo;re debugging a pipeline at 7am and you want to know not just which task failed, but which downstream datasets are now stale and which business reports can&amp;rsquo;t be trusted.&lt;/p>
&lt;p>&lt;strong>Software-Defined Assets&lt;/strong> are Dagster&amp;rsquo;s core abstraction. Each asset declares its dependencies, its materialisation logic, and its freshness expectations. The framework automatically builds a dependency graph, tracks lineage at the column level, and can tell you in real time which assets are stale, which are being materialised, and which are healthy. This is observability that MWAA simply can&amp;rsquo;t match — you&amp;rsquo;d need to bolt on separate lineage tools, metadata platforms, and custom monitoring to approximate what Dagster provides natively.&lt;/p>
&lt;p>The developer experience is excellent. &lt;code>dagster dev&lt;/code> starts a local instance immediately without Docker. Branch deployments create isolated staging environments from pull requests — push a feature branch, get a dedicated Dagster instance for testing. dbt integration maps models directly into the asset graph, so your dbt transformations are first-class citizens alongside your Python pipelines.&lt;/p>
&lt;br>
&lt;p>Dagster+ pricing uses a credit-based model: one credit equals one asset materialisation. The Solo tier starts at $10/month, Starter at $100/month. For infrequent batch workloads under 30,000 materialisations per month, Dagster is dramatically cheaper than MWAA. But credits can scale unpredictably — one practitioner calculated that an 8-operation job running every five minutes consumed 69,120 credits per month, costing $2,464 at Solo rates. For high-frequency, multi-step jobs, MWAA&amp;rsquo;s predictable hourly pricing may actually be more economical.&lt;/p>
&lt;p>The ecosystem trade-off is real. Airflow has 1,600+ operators and a community of 80,000+ organisations built over a decade. Dagster&amp;rsquo;s integration library is growing rapidly but isn&amp;rsquo;t as broad. If you need a provider for an obscure SaaS API or legacy system, Airflow almost certainly has one; Dagster might not. For teams starting fresh without Airflow baggage, Dagster&amp;rsquo;s asset-centric model is compelling. For teams with hundreds of existing DAGs, the migration cost is substantial.&lt;/p>
&lt;p>There&amp;rsquo;s also a philosophical difference worth naming. Airflow thinks in terms of &lt;em>schedules and tasks&lt;/em> — &amp;ldquo;run this DAG at 6am, execute these tasks in order.&amp;rdquo; Dagster thinks in terms of &lt;em>data freshness and assets&lt;/em> — &amp;ldquo;these datasets should be no more than 2 hours stale, materialise them as needed.&amp;rdquo; Both models work, but they lead to fundamentally different approaches to monitoring, alerting, and debugging. If your team&amp;rsquo;s primary question is &amp;ldquo;did the 6am job run?&amp;rdquo; then Airflow&amp;rsquo;s model fits naturally. If your question is &amp;ldquo;is the executive dashboard showing current data, and if not, what&amp;rsquo;s stale?&amp;rdquo; then Dagster&amp;rsquo;s model is more direct.&lt;/p>
&lt;p>The deployment model also differs meaningfully. Dagster+ offers both &lt;strong>Serverless&lt;/strong> (fully managed, AWS-hosted) and &lt;strong>Hybrid&lt;/strong> (Dagster manages the control plane, your infrastructure runs the compute). The Hybrid model gives you the thin orchestrator pattern with professional management of the scheduling layer — a compelling middle ground between fully managed and fully self-hosted.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="head-to-head-prefect-cloud">Head-to-Head: Prefect Cloud&lt;/h3>
&lt;br>
&lt;p>Prefect takes the most Python-native approach to orchestration. Add &lt;code>@flow&lt;/code> and &lt;code>@task&lt;/code> decorators to regular Python functions. No DSL, no custom operators, no XCom for passing data between tasks — just Python.&lt;/p>
&lt;p>This matters most for teams where the data engineers are primarily Python developers who find Airflow&amp;rsquo;s operator model cumbersome. Dynamic workflows, event-driven triggers, and runtime conditional logic are all first-class features in Prefect. Need to loop over a list of files and process each one as a separate task? That&amp;rsquo;s a Python for-loop, not an Airflow dynamic task mapping exercise.&lt;/p>
&lt;p>Prefect Cloud&amp;rsquo;s free Hobby tier (2 users, 5 deployments) lets small teams start at $0 — versus MWAA&amp;rsquo;s $365+ minimum. The Starter tier at $100/month includes bring-your-own-compute, meaning you still pay separately for the infrastructure where your flows actually run. Prefect Cloud is a control plane, not a compute platform — similar to the thin orchestrator pattern, but with Prefect managing the scheduling and monitoring layer instead of Airflow.&lt;/p>
&lt;p>The honest limitation: Prefect&amp;rsquo;s community is a fraction of Airflow&amp;rsquo;s. Airflow pulls 30 million monthly downloads; Prefect sits around 1.8 million weekly. That gap means fewer Stack Overflow answers, fewer blog posts solving your exact problem, and fewer engineers who already know the tool when you&amp;rsquo;re hiring. For teams where hiring velocity matters, Airflow&amp;rsquo;s ubiquity is a genuine competitive advantage — and MWAA inherits that.&lt;/p>
&lt;p>Where Prefect genuinely shines is the rapid-prototyping-to-production pipeline. You can convert an existing Python script into an orchestrated workflow by adding decorators — no rewriting into operator patterns, no separating logic from configuration, no learning a new abstraction layer. One engineering team compared this directly to their Airflow experience and found that Prefect flows took roughly half the time to develop and deploy. The trade-off is less structure — Airflow&amp;rsquo;s opinionated DAG model forces a discipline that Prefect&amp;rsquo;s flexibility doesn&amp;rsquo;t enforce.&lt;/p>
&lt;p>Prefect Cloud&amp;rsquo;s Pro tier ($500/month) adds RBAC, audit logs, and custom retention policies. The Enterprise tier (custom pricing) adds SSO, dedicated infrastructure, and priority support. For context, an MWAA mw1.small environment costs $470–540/month all-in — roughly equivalent to Prefect Pro before you add your own compute infrastructure. The total cost comparison depends entirely on how much compute your flows need and where it runs.&lt;/p>
&lt;p>One more thing worth noting: Prefect 2 (the current generation) was a ground-up rewrite that broke compatibility with Prefect 1. If you&amp;rsquo;re evaluating Prefect, you&amp;rsquo;re evaluating a framework that made the decision to start over once already. That took conviction, and the result is a significantly better product — but it&amp;rsquo;s worth understanding that the ecosystem is younger than the company&amp;rsquo;s founding date suggests.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="head-to-head-aws-step-functions">Head-to-Head: AWS Step Functions&lt;/h3>
&lt;br>
&lt;p>Step Functions is the right answer more often than most Airflow advocates want to admit.&lt;/p>
&lt;p>For a 10-step workflow running once daily (~300 state transitions per month), Step Functions costs nothing within the free tier. MWAA&amp;rsquo;s minimum is $365+. Even at medium scale — 100 executions per day — costs stay under $1/month for standard workflows. The Distributed Map mode can process up to 10,000 parallel S3 objects; one AWS demo processed 560,000 CSV files in 100 seconds.&lt;/p>
&lt;p>Step Functions also runs natively serverless with zero infrastructure management. No VPC, no NAT gateways, no &lt;code>requirements.txt&lt;/code> crashes, no 20-minute environment updates. For teams building AWS-native event-driven architectures — S3 events triggering Glue jobs, Lambda processing, DynamoDB writes — Step Functions integrates with over 220 AWS services directly in the workflow definition.&lt;/p>
&lt;p>The limitations are real though. The 256KB payload limit between states constrains data-heavy handoffs. JSON-based Amazon States Language for workflow definitions is verbose and hard to debug. There&amp;rsquo;s no built-in scheduling (you need EventBridge), no backfill capability, no equivalent of Airflow&amp;rsquo;s &lt;code>catchup=True&lt;/code> for replaying historical runs. The 25,000-event history limit per standard execution bites data engineering workloads — though Express Workflows (5-minute maximum duration, $1 per million requests) eliminate that constraint for short-lived jobs.&lt;/p>
&lt;p>For complex data pipelines with interdependencies, branching logic, retries, and backfill requirements, MWAA is genuinely the better tool. For linear workflows, event-driven processing, and microservice orchestration, Step Functions wins on cost, simplicity, and operational overhead.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="beyond-named-platforms-a-pattern-worth-considering">Beyond Named Platforms: A Pattern Worth Considering&lt;/h3>
&lt;br>
&lt;p>Beyond named platforms, there&amp;rsquo;s an architectural pattern that deserves consideration alongside them — one that doesn&amp;rsquo;t come with a vendor attached.&lt;/p>
&lt;p>The idea is simple: &lt;strong>run Airflow as a lightweight scheduler and UI on minimal infrastructure, and delegate all actual compute to Kubernetes pods or ECS tasks that spin up on demand.&lt;/strong>&lt;/p>
&lt;p>In this model, your Airflow installation is deliberately &amp;ldquo;dumb.&amp;rdquo; A small EC2 instance or a modest EKS deployment runs just the webserver, scheduler, and metadata database. Workers don&amp;rsquo;t exist in the traditional sense. Instead, every task in your DAG uses the &lt;strong>KubernetesExecutor&lt;/strong> (or &lt;code>EcsRunTaskOperator&lt;/code>) to launch a purpose-built container that does the actual work, then terminates.&lt;/p>
&lt;p>Here&amp;rsquo;s why this matters:&lt;/p>
&lt;p>&lt;strong>Per-task isolation.&lt;/strong> Each task runs in its own container with its own Docker image, its own dependencies, its own resource limits. Your dbt task runs a slim Python image with just dbt-core. Your Spark submission task runs an image with PySpark and your JAR files. Your ML training task gets a GPU-enabled image with PyTorch. No dependency conflicts. No shared memory pressure. No crash-loop risk from a bad &lt;code>requirements.txt&lt;/code> — because each task manages its own dependencies independently.&lt;/p>
&lt;p>&lt;strong>Scale to zero.&lt;/strong> When nothing is running, you&amp;rsquo;re paying for the scheduler and webserver — maybe $50–80/month on a small EC2 instance or ECS task. When a hundred tasks fire simultaneously, Kubernetes spins up a hundred pods, each with exactly the resources that task needs. When they finish, those pods terminate and you stop paying. MWAA, by contrast, charges for workers 24/7 regardless of whether tasks are running.&lt;/p>
&lt;p>&lt;strong>No CeleryExecutor limitations.&lt;/strong> The KubernetesExecutor gives you everything MWAA&amp;rsquo;s CeleryExecutor can&amp;rsquo;t: custom Docker images per task, no 12-hour execution limit (Kubernetes pods can run indefinitely), pod-level resource requests and limits, and the ability to use Spot instances for worker nodes at 50–70% cost savings.&lt;/p>
&lt;p>&lt;strong>The cold-start trade-off.&lt;/strong> Spinning up a Kubernetes pod takes 15–30 seconds. For batch workloads that run hourly or daily, this is negligible. For real-time or sub-minute scheduling, it&amp;rsquo;s not ideal — but those workloads probably shouldn&amp;rsquo;t be in Airflow anyway.&lt;/p>
&lt;p>&lt;strong>What does this actually look like in practice?&lt;/strong> Your DAG file stays clean and declarative. Instead of using &lt;code>PythonOperator&lt;/code> to run transformations inside the Airflow worker, you use &lt;code>KubernetesPodOperator&lt;/code> pointing at a Docker image that contains your transformation logic and its dependencies:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>transform_sales &lt;span style="color:#f92672">=&lt;/span> KubernetesPodOperator(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> task_id&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#34;transform_sales_data&amp;#34;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> image&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#34;your-ecr-repo/sales-transform:v2.3&amp;#34;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> namespace&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#34;airflow-tasks&amp;#34;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> resources&lt;span style="color:#f92672">=&lt;/span>{&lt;span style="color:#e6db74">&amp;#34;request_cpu&amp;#34;&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;2&amp;#34;&lt;/span>, &lt;span style="color:#e6db74">&amp;#34;request_memory&amp;#34;&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;4Gi&amp;#34;&lt;/span>},
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> node_selector&lt;span style="color:#f92672">=&lt;/span>{&lt;span style="color:#e6db74">&amp;#34;lifecycle&amp;#34;&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;spot&amp;#34;&lt;/span>}, &lt;span style="color:#75715e"># 50-70% cost savings&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> is_delete_operator_pod&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#66d9ef">True&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Each task declares exactly what it needs. Your dbt task runs a slim image. Your Spark submission gets a heavier one. Your ML training task gets GPU resources. And because each pod is ephemeral, a dependency conflict in one task can never affect another — the isolation is complete.&lt;/p>
&lt;br>
&lt;p>So is this just a slower version of what MWAA does? No — it&amp;rsquo;s architecturally different in important ways.&lt;/p>
&lt;p>MWAA uses CeleryExecutor with always-on Fargate workers. Tasks run inside those workers, sharing the same Python environment, the same dependencies, the same memory. The workers exist whether tasks are queued or not. When you update &lt;code>requirements.txt&lt;/code>, every worker gets rebuilt (hence the 20–40 minute update times).&lt;/p>
&lt;p>The thin orchestrator pattern puts compute &lt;em>outside&lt;/em> the scheduler entirely. Airflow becomes a pure control plane. The data plane — where actual work happens — is ephemeral and independently deployable. You can update a task&amp;rsquo;s Docker image in seconds without touching the Airflow environment at all.&lt;/p>
&lt;p>The self-hosted overhead is real though. You&amp;rsquo;re managing the Airflow installation yourself: upgrades, database administration (or using Amazon RDS for the metadata database), security patching, and monitoring setup. Budget 0.25–0.5 FTE for a small-to-medium deployment. At $150K/year fully loaded, even 0.25 FTE equals $3,125/month in engineering time — which can exceed MWAA&amp;rsquo;s premium for teams that don&amp;rsquo;t already have Kubernetes expertise.&lt;/p>
&lt;p>This pattern works best when your team already runs Kubernetes for other workloads (so the infrastructure cost is shared), you need per-task dependency isolation, or you&amp;rsquo;re running compute-heavy tasks where Spot instance pricing makes a material difference.&lt;/p>
&lt;p>One team running 80 pipelines on self-hosted ECS reported $400–500/month total compute costs. A comparable MWAA setup would run $600–900/month. But the engineering time to maintain the self-hosted setup is the real cost comparison — and it varies enormously depending on team capability.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="when-mwaa-is--and-isnt--the-right-choice">When MWAA Is — and Isn&amp;rsquo;t — the Right Choice&lt;/h3>
&lt;br>
&lt;p>MWAA earns its place when three conditions align: &lt;strong>deep AWS commitment&lt;/strong> (your stack is S3, Redshift, Glue, EMR), &lt;strong>existing Airflow investment&lt;/strong> (DAGs, team expertise, operator familiarity), and a &lt;strong>preference for managed infrastructure&lt;/strong> over platform engineering. Teams running 50–500 DAGs on AWS with two or more data engineers and no multi-cloud requirements will find MWAA productive and cost-effective relative to the operational burden of self-hosting.&lt;/p>
&lt;p>Here&amp;rsquo;s where it falls short:&lt;/p>
&lt;p>&lt;strong>Intermittent or light workloads&lt;/strong> — unless you&amp;rsquo;re willing to accept MWAA Serverless&amp;rsquo;s operator restrictions, the always-on cost floor is too high. Step Functions or Prefect&amp;rsquo;s free tier are dramatically cheaper.&lt;/p>
&lt;p>&lt;strong>Multi-cloud requirements&lt;/strong> — MWAA is AWS-only. Astronomer or self-hosted Airflow provides portability across clouds.&lt;/p>
&lt;p>&lt;strong>Asset-centric observability&lt;/strong> — if data lineage, freshness tracking, and column-level quality monitoring are priorities, Dagster Cloud is purpose-built for this. Bolting these capabilities onto MWAA requires multiple additional tools and significant integration work.&lt;/p>
&lt;p>&lt;strong>KubernetesExecutor or per-task isolation&lt;/strong> — MWAA is locked to CeleryExecutor. Teams needing custom Docker images per task or pod-level resource isolation should look at Astronomer, the thin orchestrator pattern on self-hosted, or Dagster Cloud.&lt;/p>
&lt;p>&lt;strong>Rapid Airflow version adoption&lt;/strong> — MWAA lags upstream releases by 2–6 months. Astronomer typically offers day-zero support for new versions.&lt;/p>
&lt;p>&lt;strong>Cost sensitivity at scale with multiple environments&lt;/strong> — per-team MWAA environments compound quickly. Three teams each needing their own environment is $1,100+/month before any work happens. Astronomer&amp;rsquo;s scale-to-zero workers or self-hosted with Spot instances become significantly cheaper at this scale.&lt;/p>
&lt;br>
&lt;p>For teams genuinely evaluating from scratch, here&amp;rsquo;s the landscape at a glance:&lt;/p>
&lt;table>
 &lt;thead>
 &lt;tr>
 &lt;th>Factor&lt;/th>
 &lt;th>MWAA&lt;/th>
 &lt;th>Astronomer Astro&lt;/th>
 &lt;th>Dagster Cloud&lt;/th>
 &lt;th>Prefect Cloud&lt;/th>
 &lt;th>Step Functions&lt;/th>
 &lt;th>Self-Hosted + K8s&lt;/th>
 &lt;/tr>
 &lt;/thead>
 &lt;tbody>
 &lt;tr>
 &lt;td>Minimum monthly cost&lt;/td>
 &lt;td>~$470&lt;/td>
 &lt;td>~$260 (scale-to-zero)&lt;/td>
 &lt;td>$10&lt;/td>
 &lt;td>$0 (Hobby)&lt;/td>
 &lt;td>$0 (free tier)&lt;/td>
 &lt;td>~$80 (scheduler only)&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Heavy workload cost&lt;/td>
 &lt;td>$1,800–2,300&lt;/td>
 &lt;td>$1,200–2,000&lt;/td>
 &lt;td>Variable (credits)&lt;/td>
 &lt;td>$500+ plus compute&lt;/td>
 &lt;td>&amp;lt;$50&lt;/td>
 &lt;td>$400–600 plus eng. time&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Setup complexity&lt;/td>
 &lt;td>Medium&lt;/td>
 &lt;td>Low&lt;/td>
 &lt;td>Low&lt;/td>
 &lt;td>Low&lt;/td>
 &lt;td>Low&lt;/td>
 &lt;td>High&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Ongoing ops burden&lt;/td>
 &lt;td>Low&lt;/td>
 &lt;td>Low&lt;/td>
 &lt;td>Low&lt;/td>
 &lt;td>Low–Medium&lt;/td>
 &lt;td>None&lt;/td>
 &lt;td>High&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Airflow compatibility&lt;/td>
 &lt;td>Full&lt;/td>
 &lt;td>Full&lt;/td>
 &lt;td>None (different paradigm)&lt;/td>
 &lt;td>None (different paradigm)&lt;/td>
 &lt;td>N/A&lt;/td>
 &lt;td>Full&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>KubernetesExecutor&lt;/td>
 &lt;td>No (Celery only)&lt;/td>
 &lt;td>Yes&lt;/td>
 &lt;td>N/A&lt;/td>
 &lt;td>N/A&lt;/td>
 &lt;td>N/A&lt;/td>
 &lt;td>Yes&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Observability&lt;/td>
 &lt;td>CloudWatch only&lt;/td>
 &lt;td>Built-in Observe&lt;/td>
 &lt;td>Native asset tracking&lt;/td>
 &lt;td>Built-in dashboard&lt;/td>
 &lt;td>CloudWatch/X-Ray&lt;/td>
 &lt;td>Whatever you build&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Multi-cloud&lt;/td>
 &lt;td>No&lt;/td>
 &lt;td>Yes&lt;/td>
 &lt;td>Yes&lt;/td>
 &lt;td>Yes&lt;/td>
 &lt;td>No&lt;/td>
 &lt;td>Yes&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Ecosystem breadth&lt;/td>
 &lt;td>Largest (Airflow)&lt;/td>
 &lt;td>Largest (Airflow)&lt;/td>
 &lt;td>Growing&lt;/td>
 &lt;td>Moderate&lt;/td>
 &lt;td>AWS-native (220+ services)&lt;/td>
 &lt;td>Largest (Airflow)&lt;/td>
 &lt;/tr>
 &lt;/tbody>
&lt;/table>
&lt;p>Here&amp;rsquo;s how I&amp;rsquo;d frame the decision:&lt;/p>
&lt;p>&lt;strong>Choose MWAA&lt;/strong> if you&amp;rsquo;re on AWS, your team knows Airflow, you want AWS to handle the scheduler and metadata database, and your workloads are substantial enough to justify the cost floor. It&amp;rsquo;s a solid, mature service that has improved dramatically in 2024–2025.&lt;/p>
&lt;p>&lt;strong>Choose Astronomer Astro&lt;/strong> if developer experience, KubernetesExecutor support, and fast Airflow upgrades matter more than AWS-native billing integration. It&amp;rsquo;s MWAA but better in almost every operational dimension — at a comparable or lower cost for active workloads.&lt;/p>
&lt;p>&lt;strong>Choose Dagster Cloud&lt;/strong> if you&amp;rsquo;re starting fresh, prioritise observability and data quality, and your team is comfortable learning a new paradigm. The asset-centric model is genuinely superior for understanding the health of your data platform.&lt;/p>
&lt;p>&lt;strong>Choose Prefect Cloud&lt;/strong> if your team is Python-first, dislikes Airflow&amp;rsquo;s operator model, and values dynamic workflows with minimal boilerplate. The free tier makes experimentation trivial.&lt;/p>
&lt;p>&lt;strong>Choose the thin orchestrator pattern&lt;/strong> (self-hosted Airflow + Kubernetes/ECS compute) if you already run Kubernetes, need per-task isolation, want maximum cost control, and have the engineering capacity to maintain the Airflow installation.&lt;/p>
&lt;p>&lt;strong>Choose Step Functions&lt;/strong> if your workflows are linear, event-driven, AWS-native, and don&amp;rsquo;t require backfill or complex scheduling. For many data engineering use cases, Step Functions plus EventBridge is simpler, cheaper, and less to maintain than any Airflow deployment.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-20242025-updates-that-changed-the-equation">The 2024–2025 Updates That Changed the Equation&lt;/h3>
&lt;br>
&lt;p>MWAA&amp;rsquo;s pace of development accelerated dramatically in this period, and several updates addressed the community&amp;rsquo;s loudest complaints.&lt;/p>
&lt;p>&lt;strong>Environment class expansion&lt;/strong> in April 2024 (mw1.xlarge and mw1.2xlarge) and November 2024 (mw1.micro) gave teams both the headroom for heavy workloads and the affordable entry point for dev/test that had been missing since launch. The micro class at ~$37/month is particularly valuable — before its introduction, every dev environment cost the same $365+ as production.&lt;/p>
&lt;p>Airflow versions progressed from 2.8 through 2.9 and 2.10, with &lt;strong>Airflow 3.0 landing on MWAA in October 2025&lt;/strong>. The 3.0 release brought a redesigned React UI, event-driven scheduling via Assets (Airflow&amp;rsquo;s answer to Dagster&amp;rsquo;s Software-Defined Assets), a Task SDK with least-privilege execution, and Python 3.12 support. The Asset concept is significant — it means Airflow is converging toward the same asset-centric thinking that Dagster pioneered.&lt;/p>
&lt;p>&lt;strong>MWAA Serverless&lt;/strong> (November 2025) was the most consequential launch — true pay-per-execution pricing that directly addresses the always-on cost complaint. But as mentioned earlier, the limitations (no Airflow UI, YAML-only definitions, no PythonOperator) make it a different product for a different use case, not a cheaper version of MWAA.&lt;/p>
&lt;p>&lt;strong>Graceful updates&lt;/strong> (May 2025) eliminated the other major pain point: environments can now be updated without interrupting running tasks. The SigV4 REST API (October 2024) simplified programmatic access by supporting standard AWS credentials instead of login token management.&lt;/p>
&lt;p>Community sentiment sits at mixed-positive. One engineering team&amp;rsquo;s May 2025 evaluation captured the tension well: Airflow remains the industry standard with a proven track record, but the improvements in Airflow 3 are following features that Dagster and Prefect already have. Apache Airflow pulled 320 million downloads in 2024, dwarfing competitors — but Dagster&amp;rsquo;s growth trajectory and commit activity signal a genuine shift in the market.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="closing-thoughts">Closing Thoughts&lt;/h3>
&lt;br>
&lt;p>The orchestration landscape is converging. Airflow 3.0 adopted asset-aware scheduling — a concept Dagster pioneered. Prefect&amp;rsquo;s event-driven model influenced Airflow&amp;rsquo;s trigger and asset-watcher features. Dagster is building broader operator support. They&amp;rsquo;re all moving toward the same destination from different starting points.&lt;/p>
&lt;p>That convergence means the &amp;ldquo;right&amp;rdquo; choice increasingly depends not on feature parity but on team expertise, existing investment, and cloud strategy. And honestly? The worst decision you can make is choosing MWAA by default because you&amp;rsquo;re on AWS. Run the numbers. Evaluate the developer experience gaps. Try spinning up a Dagster Cloud or Prefect Cloud instance alongside your MWAA evaluation — both have free tiers, and the comparison will be illuminating.&lt;/p>
&lt;p>I think about that first debugging session on the inherited Step Functions setup often. Hours to find something that should have taken minutes. The platform we eventually chose didn&amp;rsquo;t change the underlying compute or the logic of the pipelines — it changed whether we could &lt;em>see&lt;/em> what was happening. Pick the orchestrator that gives your team that visibility. The one they&amp;rsquo;ll trust when something breaks at an inconvenient hour. That&amp;rsquo;s the one that wins.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Cloud Architecture</category><category>AWS</category><category>AWS</category><category>MWAA</category><category>Apache Airflow</category><category>Astronomer</category><category>Dagster</category><category>Prefect</category><category>Step Functions</category><category>Orchestration</category><category>Data Pipelines</category></item><item><title>Data Engineers: Just Do Code Reviews</title><link>https://ghostinthedata.info/posts/2026/2026-07-11-code-reviews-just-do-it/</link><pubDate>Sat, 11 Jul 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-07-11-code-reviews-just-do-it/</guid><author>Chris Hillman</author><description>The case for code reviews in software engineering was settled decades ago. Data engineering hasn't caught up — and the cost is measurable in hundreds of millions of dollars per incident.</description><content:encoded>&lt;p>In &lt;em>Code Complete&lt;/em>, Steve McConnell writes:&lt;/p>
&lt;blockquote>
&lt;p>Software testing alone has limited effectiveness — the average defect detection rate is only 25 percent for unit testing, 35 percent for function testing, and 45 percent for integration testing. In contrast, &lt;strong>the average effectiveness of design and code inspections are 55 and 60 percent&lt;/strong>.&lt;/p>&lt;/blockquote>
&lt;p>That was published in 1993. The evidence has only accumulated since. &lt;strong>Peer code review is the single most effective defect-removal technique available to software teams.&lt;/strong> Nothing else — not tests, not CI pipelines, not observability tooling — comes close on a per-hour basis.&lt;/p>
&lt;p>Software engineers have known this for fifty years. Data engineers mostly pretend it doesn&amp;rsquo;t apply to them.&lt;/p>
&lt;p>It does. And the stakes, if anything, are higher.&lt;/p>
&lt;hr>
&lt;h3 id="why-the-excuses-dont-survive-contact-with-the-data">Why the excuses don&amp;rsquo;t survive contact with the data&lt;/h3>
&lt;p>The objections are familiar. &lt;em>We&amp;rsquo;re not real software engineers. We&amp;rsquo;re too small a team. We need to move fast. Our tools don&amp;rsquo;t support it.&lt;/em> None of them hold up.&lt;/p>
&lt;p>The evidence on review effectiveness is overwhelming. Michael Fagan&amp;rsquo;s original 1976 IBM study found formal inspections caught 82% of errors before unit testing even began. NASA&amp;rsquo;s Software Engineering Laboratory found code reading detected roughly twice as many defects per hour as testing. Russell&amp;rsquo;s 1991 AT&amp;amp;T study found that each hour spent in review avoided 33 hours of downstream maintenance — making inspections up to 20× more cost-effective than testing per defect found.&lt;/p>
&lt;p>Modern numbers are just as clear. The SmartBear/Cisco study of 2,500 reviews across 3.2 million lines of code established that reviews conducted within sensible limits — under 400 lines per session, under 500 lines per hour, sessions capped at 60 minutes — yield 70–90% defect removal. Google&amp;rsquo;s case study of 9 million reviewed changes across 25,000 engineers proved the practice scales to 20,000 commits a day, with median review latency under four hours. The DORA 2023 &lt;em>State of DevOps&lt;/em> report, surveying nearly 3,000 practitioners, found that teams with faster code reviews had 50% better software delivery performance. Elite DORA performers ship with a 7× lower change failure rate than low performers.&lt;/p>
&lt;p>SmartBear&amp;rsquo;s &lt;em>State of Code Review&lt;/em> asked 800 respondents what most improves quality. Code review ranked first, with unit testing a distant second. Among teams satisfied with their software quality, 80% used tool-based peer review and 82% had documented review guidelines.&lt;/p>
&lt;p>This is not a close debate in software engineering. It shouldn&amp;rsquo;t be one in data engineering either.&lt;/p>
&lt;hr>
&lt;h3 id="data-pipelines-fail-differently--and-that-makes-reviews-more-important-not-less">Data pipelines fail differently — and that makes reviews more important, not less&lt;/h3>
&lt;p>Here&amp;rsquo;s the thing that makes the data engineering situation more precarious: application bugs are loud. They throw exceptions, crash processes, page on-call at 2am. The feedback loop that keeps application engineers honest — run it, watch it break — works quickly.&lt;/p>
&lt;p>Data bugs are quiet. The pipeline is green. Row counts look reasonable. The dashboard renders. And the number is wrong.&lt;/p>
&lt;p>Benn Stancil describes it precisely: silent data errors &amp;ldquo;don&amp;rsquo;t look suspicious and they trigger no warnings, in observability tools or in our manual spot checks. Instead, they often linger undetected, slowly and silently pushing the analytical assets that use them further and further from reality — until, during a board meeting…&amp;rdquo;&lt;/p>
&lt;p>Barr Moses of Monte Carlo has noted that time-to-detection for silent data errors can be measured in months, not minutes. Her 2023 survey of 200 data leaders found the average organisation deals with 67 data incidents per month, with average resolution time up 166% year over year to 15 hours per incident, and 68% of teams taking four or more hours just to detect a problem. In the 2024 follow-up, two-thirds of data leaders had experienced a data incident costing over $100,000 in the previous six months.&lt;/p>
&lt;p>The &amp;ldquo;run it and see if it breaks&amp;rdquo; safety net that keeps application code marginally honest doesn&amp;rsquo;t exist for pipelines. The only reliable substitute is another human reading the code before it runs.&lt;/p>
&lt;hr>
&lt;h3 id="the-public-case-studies-have-already-priced-in-what-happens-when-you-skip-it">The public case studies have already priced in what happens when you skip it&lt;/h3>
&lt;p>These are not theoretical risks.&lt;/p>
&lt;p>&lt;strong>Unity Software&lt;/strong> told investors in Q1 2022 that ingesting bad training data from a large customer cost it roughly $110 million in 2022 revenue, erasing approximately $5 billion in market capitalisation in a day. A data pipeline integrity problem, undetected until it had been compounding for quarters.&lt;/p>
&lt;p>&lt;strong>Citigroup&lt;/strong> paid a $400 million OCC fine in 2020 for long-standing failure to establish effective risk management and data governance programs — followed by another $135.6 million in 2024 for insufficient progress on remediation. Data quality and governance failures, sustained over years.&lt;/p>
&lt;p>&lt;strong>Equifax&lt;/strong> sent incorrect credit scores to lenders for three weeks in 2022, shifting scores by 25 points or more for around 300,000 consumers. The downstream decisions those scores informed — loan approvals, interest rates, credit limits — cannot be unwound.&lt;/p>
&lt;p>&lt;strong>Public Health England&lt;/strong> lost roughly 16,000 positive COVID-19 test results in October 2020 because an .xls pipeline hit Excel&amp;rsquo;s row limit. Those contacts were not traced.&lt;/p>
&lt;p>Every one of these is, at root, a data-pipeline change that no sufficiently rigorous second set of eyes reviewed before it reached production. The review wouldn&amp;rsquo;t have guaranteed they caught it. The absence of a review guaranteed they had no chance to.&lt;/p>
&lt;hr>
&lt;h3 id="the-broader-cost-is-already-priced-in--you-just-dont-see-the-bill">The broader cost is already priced in — you just don&amp;rsquo;t see the bill&lt;/h3>
&lt;p>Gartner&amp;rsquo;s running estimate is that poor data quality costs the average organisation $12.9 million per year. Thomas Redman, writing in &lt;em>Harvard Business Review&lt;/em>, put the drag on the US economy at $3.1 trillion annually. His most uncomfortable finding, from measuring 75 executives&amp;rsquo; own data: only 3% of company data meets basic quality standards, and 47% of newly-created records contain at least one critical error.&lt;/p>
&lt;p>The practitioner surveys agree. Great Expectations&amp;rsquo; 2022 study found 77% of data practitioners report quality issues, with 91% saying those issues hurt company performance. dbt Labs&amp;rsquo; &lt;em>State of Analytics Engineering 2024&lt;/em> reported 57% of data professionals cite poor data quality as their top challenge — up from 41% in 2022. The problem is getting worse. Monte Carlo&amp;rsquo;s telemetry puts it at roughly one incident per fifteen tables per year in production. Forrester&amp;rsquo;s 2023 &lt;em>Data Culture and Literacy Survey&lt;/em> found more than a quarter of respondents lose over $5 million a year to poor data quality, with 7% losing more than $25 million.&lt;/p>
&lt;p>The worst number: Monte Carlo&amp;rsquo;s 2022 data suggests data engineers spend roughly 40% of their working week — two full days — firefighting quality issues. Every hour of that is an hour a peer review before merge might have prevented.&lt;/p>
&lt;hr>
&lt;h3 id="what-the-data-community-has-actually-said-about-this">What the data community has actually said about this&lt;/h3>
&lt;p>Michael Kaminsky at Locally Optimistic is clear: &amp;ldquo;Code review for analytics is often substantively different from code review for software engineering because reviewers need to check the business logic and the analytical methods as well as the code.&amp;rdquo; Different, meaning harder and more consequential — not optional.&lt;/p>
&lt;p>Tristan Handy&amp;rsquo;s argument that mart models should be treated as stable interfaces, with versions and deprecation windows like APIs, is now the foundation of dbt&amp;rsquo;s contracts feature. It is also the right frame for what a reviewer is protecting: not just the code, but the downstream trust placed in it.&lt;/p>
&lt;p>Emilie Schario&amp;rsquo;s observation remains accurate: &amp;ldquo;Data is behind software development when it comes to learning and implementing the best practices of DevOps.&amp;rdquo; The review is how you close the gap.&lt;/p>
&lt;p>And Maxime Beauchemin, who built Airflow and Superset, has documented the cultural cost of data engineering&amp;rsquo;s second-class status. When the team is small, when velocity is the priority, when &amp;ldquo;we&amp;rsquo;re not really software engineers&amp;rdquo; is the ambient assumption, reviews are the first thing to go. And the work becomes fragile, tribal, and unreviewable — not because of the tools, but because of the habits.&lt;/p>
&lt;p>Datafold&amp;rsquo;s write-up of analytics-engineer review culture names the failure mode precisely: &amp;ldquo;Faced with the prospect of move fast and break stuff vs. move slow, carefully review, and maybe break less stuff, most reviewers will default to the first option, and &amp;lsquo;LGTM&amp;rsquo; their way through the day.&amp;rdquo;&lt;/p>
&lt;p>That is not a review culture. That is a rubber-stamp culture with extra steps.&lt;/p>
&lt;hr>
&lt;h3 id="what-a-data-engineering-review-should-actually-cover">What a data engineering review should actually cover&lt;/h3>
&lt;p>The community consensus — distilled from dbt Labs&amp;rsquo; PR guidance, the GitLab Data Team Handbook, Locally Optimistic, Datafold, and the dbt Discourse — converges on a specific checklist. One focused on the failure modes unique to data, not imported wholesale from software engineering.&lt;/p>
&lt;table>
 &lt;thead>
 &lt;tr>
 &lt;th>Area&lt;/th>
 &lt;th>What the reviewer checks&lt;/th>
 &lt;/tr>
 &lt;/thead>
 &lt;tbody>
 &lt;tr>
 &lt;td>&lt;strong>SQL logic and grain&lt;/strong>&lt;/td>
 &lt;td>No silent fan-outs from joins; stated grain matches actual grain; CTEs are readable; row counts make sense against production&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>dbt design&lt;/strong>&lt;/td>
 &lt;td>&lt;code>ref()&lt;/code> used throughout; sources declared; correct materialisation type; incremental models have &lt;code>unique_key&lt;/code> and a sane predicate&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>Schema and breaking changes&lt;/strong>&lt;/td>
 &lt;td>Contracts enforced for public models; version bumps with a deprecation window; data-diff output reviewed&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>Tests&lt;/strong>&lt;/td>
 &lt;td>New models carry &lt;code>unique&lt;/code> and &lt;code>not_null&lt;/code> at the grain; grain-changing transforms have data tests; business rules encoded as tests, not comments&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>Docs and metadata&lt;/strong>&lt;/td>
 &lt;td>Model and column descriptions present; source freshness configured; PII flags set&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>Idempotency and performance&lt;/strong>&lt;/td>
 &lt;td>Re-runs produce identical results; no pinned &lt;code>CURRENT_TIMESTAMP&lt;/code>; partition and cluster keys considered; warehouse cost estimated&lt;/td>
 &lt;/tr>
 &lt;/tbody>
&lt;/table>
&lt;p>This is not a burdensome list. It&amp;rsquo;s the checklist a thorough engineer would run through mentally anyway — just written down and shared, so someone else can run through it too.&lt;/p>
&lt;hr>
&lt;h3 id="just-do-it">Just do it&lt;/h3>
&lt;p>The case for code review in software engineering was settled by Fagan, IBM, NASA, Cisco, Microsoft, Google, and DORA — across decades and millions of commits. Reading code finds roughly 70% of defects before they ship. Teams that review fast deliver 50% better. Silent errors cost enterprises tens of millions per year and mid-market teams their credibility with the business.&lt;/p>
&lt;p>The case for code review in data engineering is not weaker than in software engineering. It&amp;rsquo;s stronger, because pipelines fail silently, errors compound downstream before anyone notices, and the financial consequences show up on earnings calls rather than error logs.&lt;/p>
&lt;p>The practitioners who have thought hardest about this — Kaminsky, Stancil, Moses, Sanderson, Beauchemin, Handy — all land in the same place. Every SQL change, every dbt model, every DAG, every schema migration deserves a second set of eyes before it touches production. Not because data engineers are software engineers, but because the cost of pretending the practice doesn&amp;rsquo;t apply has become impossible to afford.&lt;/p>
&lt;p>The tools exist. The checklists exist. The statistics exist. The only thing missing is the decision to stop skipping it.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Data Quality</category><category>Career Development</category><category>Code Review</category><category>Data Quality</category><category>Data Engineering</category><category>SQL</category><category>dbt</category><category>Engineering Culture</category><category>Best Practices</category></item><item><title>You Can't Incentivise a Pipeline That Doesn't Break</title><link>https://ghostinthedata.info/posts/2026/2026-07-04-incentive-pay/</link><pubDate>Sat, 04 Jul 2026 08:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-07-04-incentive-pay/</guid><author>Chris Hillman</author><description>Incentive pay is designed for a world where output is countable. In data engineering, the most valuable work is invisible. Attaching money to metrics doesn't reveal that work — it buries it.</description><content:encoded>&lt;p>I worked alongside a data engineer who was, by every formal measure, the best performer on the team. He was also quietly destroying the platform.&lt;/p>
&lt;p>Not maliciously. He was optimising for the thing being measured. His work shipped fast because he skipped the edge case analysis. He closed tickets at first resolution without ever checking whether the underlying pattern would recur. He didn&amp;rsquo;t review anyone else&amp;rsquo;s PRs (not his KPIs, so why would he?).&lt;/p>
&lt;p>One of his teammates was the person actually holding the platform together. Quieter, slower to release, never at the top of any metric. She questioned everything, dug into the data, and worked with stakeholders to write elaborate testing scenarios. She was the one who caught a new source system sending nulls where the schema expected integers, and had been doing so for three weeks. She paired with the juniors. She wrote the runbook nobody had written.&lt;/p>
&lt;p>Her performance review that year was average. His was excellent.&lt;/p>
&lt;p>If you&amp;rsquo;ve worked in data engineering for any length of time, you&amp;rsquo;ve seen a version of this. If you&amp;rsquo;ve moved into leadership, you may have been the one signing off on it.&lt;/p>
&lt;p>That&amp;rsquo;s the incentive pay problem. Not that it doesn&amp;rsquo;t motivate people. It motivates them to do the wrong things.&lt;/p>
&lt;hr>
&lt;h2 id="the-measurement-trap">The Measurement Trap&lt;/h2>
&lt;br>
&lt;p>There&amp;rsquo;s a name for this: Goodhart&amp;rsquo;s Law. When a measure becomes a target, it ceases to be a good measure. The moment you attach money to a metric, you&amp;rsquo;ve changed what people optimise for, and nearly every metric you can pull from the tooling is the wrong target.&lt;/p>
&lt;p>Pipelines deployed rewards volume and punishes care. An engineer who ships fifteen small, well-tested, well-documented pipelines is worth more than one who ships thirty fragile ones. No spreadsheet you build will ever agree.&lt;/p>
&lt;p>Tickets closed incentivises decomposition. Want a higher count? Chop big problems into small ones and resolve at the surface. The underlying pattern that keeps generating tickets isn&amp;rsquo;t your problem. That&amp;rsquo;s a different ticket, a different quarter, someone else&amp;rsquo;s metric.&lt;/p>
&lt;p>Incident count punishes the team for things outside their control. An upstream system starts sending malformed payloads at 2am and your bonus goes down. (The vendor who caused it hits their SLA and gets paid in full, naturally.)&lt;/p>
&lt;p>I&amp;rsquo;ve watched other leads volunteer for these measures because they&amp;rsquo;re easy to extract from Jira and GitHub. Easy to extract is not the same as meaningful. In my experience it&amp;rsquo;s usually the opposite.&lt;/p>
&lt;hr>
&lt;h2 id="the-invisible-work-problem">The Invisible Work Problem&lt;/h2>
&lt;br>
&lt;p>The deeper issue is that the most valuable work a data engineer does often leaves no trace.&lt;/p>
&lt;p>The schema that doesn&amp;rsquo;t need refactoring in six months because someone thought carefully about how it would evolve. The data contract that caught a breaking change. Good engineering actively suppresses the numbers that look like productivity, because fewer incidents means fewer tickets, and fewer tickets means a lower &amp;ldquo;work completed&amp;rdquo; count.&lt;/p>
&lt;p>Then there&amp;rsquo;s the social work. The engineer who reviews PRs with care, who explains why and not just what, who flags when a colleague&amp;rsquo;s approach will cause trouble downstream, is compressing the learning curve for the whole team. They prevent incidents that never appear in any log, because they never happened.&lt;/p>
&lt;p>I wrote about this from a different angle in &lt;a href="https://ghostinthedata.info/posts/2025/2025-12-08-invisible-pr/" target="_blank" rel="noopener">The Invisible PR You&amp;rsquo;re Building Right Now&lt;/a> — the reputation built through small interactions that compound over years. Incentive pay is the structural version of the same idea. Most schemes don&amp;rsquo;t just miss this work, they penalise it, because a careful PR review is time not spent on your own ticket count.&lt;/p>
&lt;hr>
&lt;h2 id="the-lopsided-ledger">The Lopsided Ledger&lt;/h2>
&lt;br>
&lt;p>Even a perfect set of metrics — which you can&amp;rsquo;t build, but suppose — runs into some uncomfortable arithmetic. A strong review doesn&amp;rsquo;t make a good engineer better; at most it confirms what they already believed. A review that misses what someone actually did makes them recalibrate, and next quarter they do the work that gets measured instead of the work that matters. The downside dwarfs the upside. And since most people believe they&amp;rsquo;re above average and most systems grade on a curve, disappointment is the default output. You&amp;rsquo;re spending management time and team morale on a mechanism that runs at a loss.&lt;/p>
&lt;hr>
&lt;h2 id="so-what-actually-works">So What Actually Works&lt;/h2>
&lt;br>
&lt;p>I&amp;rsquo;m not arguing that data engineers don&amp;rsquo;t care about money. They do, and they should.&lt;/p>
&lt;p>Dan Pink covered this territory in &lt;em>Drive&lt;/em>, and his central finding is facinating: if-then rewards work well for mechanical tasks and reliably degrade performance on cognitive ones. Data engineering is nothing but cognitive work. His answer, and mine after years of watching teams, is to pay people well enough that money stops being the conversation — then use the levers that actually move things.&lt;/p>
&lt;p>Autonomy. Engineers who are told what outcome is needed, and trusted to work out how, consistently produce better work than those handed a specification. The person who understands the problem deeply enough to design the solution will always see things the specification missed. Give people the problem, not the method.&lt;/p>
&lt;p>Mastery. The most engaged engineers on my teams are the ones getting meaningfully better at something. Not completing training modules — deepening their craft. Conference time, learning budgets that actually get spent, a problem slightly beyond their current reach with a safety net underneath.&lt;/p>
&lt;p>Purpose. Data engineers rarely see the downstream consequence of what they build. Closing that loop matters enormously. When an engineer learns that the student outcome data they built powers the retention intervention that changed whether someone finished their degree, no quarterly bonus comes close.&lt;/p>
&lt;p>And recognition, the specific and timely kind. Not plaques, not a Teams message with too many emojis. The manager who says quietly in a 1:1: &amp;ldquo;The way you caught that schema drift before it hit the mart prevented an incident nobody will ever know about, and I noticed.&amp;rdquo; That costs nothing, and it does more for morale than any incentive scheme I&amp;rsquo;ve witnessed.&lt;/p>
&lt;hr>
&lt;h2 id="the-pipeline-that-doesnt-break">The Pipeline That Doesn&amp;rsquo;t Break&lt;/h2>
&lt;br>
&lt;p>The best data engineering work is characterised by what doesn&amp;rsquo;t happen.&lt;/p>
&lt;p>The data that arrives without error. The migration that completes cleanly. The breaking change caught in review. The junior who is now autonomous on a system they couldn&amp;rsquo;t have touched six months ago.&lt;/p>
&lt;p>None of this is countable. None of it generates a closed ticket or increments a deployment counter. Which means it will always be invisible to any incentive scheme complex enough to automate and simple enough to implement.&lt;/p>
&lt;p>There is an honest cost on the other side. Drop the metrics and you lose the comfort of a ranking you can defend in a calibration meeting. Judging invisible work means knowing your team&amp;rsquo;s work well enough to see it, and that&amp;rsquo;s harder than reading a dashboard.&lt;/p>
&lt;p>Pay people fairly. Pay them enough that salary isn&amp;rsquo;t what they&amp;rsquo;re thinking about while debugging a nightly job at 7am. Then build the conditions where the work itself is worth doing.&lt;/p>
&lt;p>The engineers who hold your platform together aren&amp;rsquo;t doing it for the bonus. And the moment you start paying them as if they were, you&amp;rsquo;ll start losing the ones who weren&amp;rsquo;t.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>leadership</category><category>team</category><category>Leadership</category><category>Team Culture</category><category>Performance Reviews</category><category>Data Engineering</category><category>Management</category><category>Incentives</category><category>Motivation</category></item><item><title>Your Team Already Has Patterns. They Just Don't Know It.</title><link>https://ghostinthedata.info/posts/2026/2026-06-27-pattern-bank/</link><pubDate>Sat, 27 Jun 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-06-27-pattern-bank/</guid><author>Chris Hillman</author><description>A pattern bank gives your data engineering team a shared vocabulary for how work gets done — and the confidence to estimate it. Here's how to build one collaboratively, without turning it into a bureaucratic exercise.</description><content:encoded>&lt;p>When I started a new role, one of the first things I did was try to understand how data moved through the system. Not the dashboards, not the data models — the pipes. Where did things come from? How did they get in? What happened to them along the way?&lt;/p>
&lt;p>There were somewhere between twenty and thirty source systems feeding the platform. Not a massive number, but enough to tell a story when you looked at the ingestion layer all at once. What I found was that all the pipelines had originated from two base templates. A sensible starting point. The kind of thing a small team puts in place early to stop complete chaos.&lt;/p>
&lt;p>But over time, as new sources were onboarded by different engineers under different pressures, the implementations had drifted from those templates in ways that were hard to track. Some pipelines were applying data transformations before data even landed in Snowflake — so the raw layer wasn&amp;rsquo;t actually raw. Some were enforcing data contracts at the landing stage, before the data had been typed, which created problems downstream when contract violations turned up in places the pipeline wasn&amp;rsquo;t designed to handle them. And there was no consistent approach to manifest file checking — verifying that the file you received matched what the source said it sent — which is exactly the kind of gap that sits quietly until a source delivers a partial file on a bad day and nobody notices until a stakeholder does.&lt;/p>
&lt;p>None of this was catastrophic. I want to be clear about that. I&amp;rsquo;ve seen far worse — environments where the ingestion layer is a genuine free-for-all with no common ancestry at all. By comparison, this was a B+. The bones were solid. There was a lineage to the templates you could trace, and the engineers who&amp;rsquo;d built on them had generally been trying to do the right thing. But there was drift, and where there&amp;rsquo;s drift there are gaps, and gaps have a way of becoming expensive at the worst possible moments.&lt;/p>
&lt;p>The flip side was interesting. The transformation layer and the serving patterns were genuinely robust. Whoever had shaped those had been thinking in terms of reusable approaches from the beginning. SCD2 loads were SCD2 loads. Reporting tables followed a consistent structure. That part of the platform had a grammar to it, even if nobody had written it down.&lt;/p>
&lt;p>The contrast between the two layers was instructive. And it surfaced a question I keep coming back to: what would it look like if the whole team had a shared vocabulary for the work they were doing — not just in transformation, but across the whole pipeline lifecycle? Not a rigid rulebook. Not a standards document that nobody reads. Just a common language, agreed on together, that made the implicit explicit.&lt;/p>
&lt;p>That&amp;rsquo;s what a pattern bank is. And this article is about how to build one.&lt;/p>
&lt;hr>
&lt;h3 id="the-problem-that-nobody-names">The problem that nobody names&lt;/h3>
&lt;/br>
&lt;p>Before we get to the solution, it&amp;rsquo;s worth naming what actually breaks when there&amp;rsquo;s no shared pattern vocabulary on a data engineering team.&lt;/p>
&lt;p>The most obvious symptom is inconsistency. When every engineer approaches recurring problems from first principles, you end up with a codebase that looks like it was built by ten different people — because it was. Code reviews become debates about form rather than substance. A new team member joins and spends their first month just trying to understand why the SFTP ingestion from System A looks nothing like the SFTP ingestion from System B, even though they&amp;rsquo;re structurally the same thing. The answer — &amp;ldquo;that&amp;rsquo;s just how it grew&amp;rdquo; — isn&amp;rsquo;t satisfying, and it doesn&amp;rsquo;t help them build.&lt;/p>
&lt;p>But inconsistency isn&amp;rsquo;t the thing that hurts most. The thing that hurts most is estimation.&lt;/p>
&lt;p>Oe of the early challenges was quoting projects to internal customers — duration, cost, team size. The kind of thing stakeholders reasonably want to know before they commit. And without a clear picture of which parts of a solution were known quantities versus genuinely novel territory, those conversations were uncomfortable. You could make an educated guess, but you couldn&amp;rsquo;t anchor it to anything concrete. Each project felt like a bespoke exercise, even when large parts of it were things the team had done before.&lt;/p>
&lt;p>The problem has a name in project planning circles: unknown unknowns. The things you don&amp;rsquo;t know you don&amp;rsquo;t know. The hidden surprises that don&amp;rsquo;t appear in any upfront estimate because you haven&amp;rsquo;t recognised them yet. Pattern reuse directly attacks this. When you&amp;rsquo;ve built the same ingestion type before, you know roughly how long it takes, where the edge cases live, and what tends to go wrong. You&amp;rsquo;ve converted some of the unknown unknowns into known quantities — and that changes everything about how you plan.&lt;/p>
&lt;p>There was one more piece that crystallised why the pattern bank mattered. The technical architects were responsible for designing how source systems should serve data to the platform. They sat upstream of the data team, making decisions about formats, frequency, delivery mechanisms. But they didn&amp;rsquo;t have a clear picture of what happened once data arrived — what the ingestion layer expected, what would make the data engineer&amp;rsquo;s life easy versus hard, what patterns the team was already working with.&lt;/p>
&lt;p>The pattern bank, as it developed, became a communication artefact as much as an engineering one. It made the platform legible to people who sat outside it. Here is how we ingest data. Here is what we need from you when you&amp;rsquo;re designing a new source feed. Here is why the way you&amp;rsquo;re proposing to serve this data is going to create problems downstream.&lt;/p>
&lt;p>That transparency was worth as much as the internal consistency.&lt;/p>
&lt;hr>
&lt;h3 id="what-a-pattern-bank-actually-is">What a pattern bank actually is&lt;/h3>
&lt;/br>
&lt;p>The simplest definition: a pattern bank is a catalogue of reusable solution approaches that your team draws from when designing new work.&lt;/p>
&lt;p>It&amp;rsquo;s not a code library — though patterns might eventually generate template code. It&amp;rsquo;s not a runbook — those are step-by-step guides for specific incidents. It&amp;rsquo;s not a playbook — those describe how to handle particular scenarios. A pattern bank operates at a higher level of abstraction than all of those. It describes the &lt;em>shape&lt;/em> of a solution, not the &lt;em>substance&lt;/em> of a specific implementation.&lt;/p>
&lt;p>Think of it like a recipe catalogue. When you&amp;rsquo;re planning a dinner party, you choose recipes based on what you&amp;rsquo;re cooking for and what ingredients you have. You don&amp;rsquo;t invent a new cooking method each time — you draw from a repertoire. The pattern bank is the repertoire. Each new project is the dinner.&lt;/p>
&lt;p>For a data engineering team, the pattern bank tends to organise naturally into three layers that follow how data moves through a platform:&lt;/p>
&lt;p>&lt;strong>Ingestion patterns&lt;/strong> cover how data gets from a source into your environment. This might be a direct file transfer from SFTP, an API pull, an event stream, a database replication — the mechanics differ, but each is a recognisable type. At the ingestion layer, the goal is usually to get data as close to your transformation environment as possible in its raw form, before worrying about types or structure. Staging — the step where you impose data types and light transformations — often lives here too, and it&amp;rsquo;s worth treating it as a distinct sub-pattern.&lt;/p>
&lt;p>&lt;strong>Transformation patterns&lt;/strong> cover what happens to data once it&amp;rsquo;s in. SCD1 for current-state overwrite, SCD2 for full history with row versioning, transactional loads for append-only fact tables, reference table management, bi-temporal designs for systems that need to track both valid time and transaction time. These patterns are where most of the design complexity lives, and they&amp;rsquo;re often the most stable part of a mature pattern bank because the underlying approaches have been well-understood for decades.&lt;/p>
&lt;p>&lt;strong>Delivery patterns&lt;/strong> cover how processed data gets to consumers. A flat file export to a third-party system, a reporting aggregate table consumed by a BI tool, a real-time serving layer for an operational dashboard — each is a pattern with its own shape, its own testing considerations, its own failure modes.&lt;/p>
&lt;p>The power of this structure is in what it enables at design time. When a new project lands, you&amp;rsquo;re no longer starting with a blank page. You&amp;rsquo;re asking: which ingestion pattern fits this source? Which transformation pattern fits this business requirement? Which delivery pattern fits this consumer? You pick from the catalogue, combine them into a solution design, and — critically — you know what you&amp;rsquo;re working with because you&amp;rsquo;ve worked with it before.&lt;/p>
&lt;p>&lt;img src="https://ghostinthedata.info/posts/2026/2026-06-27-pattern-bank/images/pattern-bank-stage-menu.svg" alt="A solution design is one pattern from each stage, with shared tails factored into sub-patterns">&lt;/p>
&lt;p>&lt;em>One recipe, one pattern from each stage. The amber path is a single solution design; the grey fan into the shared sub-pattern is the factoring at work.&lt;/em>&lt;/p>
&lt;p>One instance of an approach is a one-off solution. Two instances is a coincidence. Three instances is a pattern worth documenting. That&amp;rsquo;s the threshold I&amp;rsquo;d suggest for adding something to the bank — build it twice, recognise it on the third, name it and write it down.&lt;/p>
&lt;hr>
&lt;h3 id="what-it-should-not-be">What it should not be&lt;/h3>
&lt;/br>
&lt;p>The counter-example that keeps me honest about what a pattern bank should &lt;em>not&lt;/em> be.&lt;/p>
&lt;p>I&amp;rsquo;ve seen pattern banks that went too far. Not too far in terms of coverage — too far in terms of specificity. Every combination of source type and transformation type and delivery type had its own entry. Every edge case was enumerated. Every variant had its own named pattern. The catalogue was exhaustive in the worst sense of the word.&lt;/p>
&lt;p>The result was a pattern bank that engineers couldn&amp;rsquo;t use without first spending an hour navigating it. The patterns were so granular that applying one felt like following a script rather than making a design decision. Worse — they were so specific that anything slightly different from the documented pattern triggered a governance conversation about whether to create a new entry or adapt an existing one. The bank had become bureaucracy wearing an engineering costume.&lt;/p>
&lt;p>A pattern bank should describe the shape of a solution clearly enough that a mid-level engineer can use it to make a design decision without asking five follow-up questions. If your patterns require a full implementation guide to be useful, they&amp;rsquo;re not patterns — they&amp;rsquo;re procedures. Keep them abstract enough to cover genuine variants. Keep them concrete enough to actually guide work. The test is whether someone can look at a pattern and understand &lt;em>what kind of problem it solves&lt;/em> within about thirty seconds.&lt;/p>
&lt;p>The three-layer structure helps with this. When you keep ingestion patterns, transformation patterns, and delivery patterns separate, each pattern is already scoped to a defined part of the problem. That&amp;rsquo;s enough structure. Within each pattern, describe what problem it solves, when to use it, when &lt;em>not&lt;/em> to use it, and roughly what it involves. That&amp;rsquo;s the minimum viable pattern documentation. A short note on known effort range doesn&amp;rsquo;t hurt either.&lt;/p>
&lt;p>Resist the urge to enumerate every combination. The recipe catalogue analogy holds: a good recipe book teaches you to cook, not to follow instructions. The patterns should develop your team&amp;rsquo;s design intuition, not replace it.&lt;/p>
&lt;p>There&amp;rsquo;s also a structural cure for the combinatorial urge, and it took me longer than it should have to see it: factor the convergence points out. When several patterns funnel into the same mechanism — say, five different ingestion types that all end in the same landing-and-staging steps — that shared tail isn&amp;rsquo;t a reason to write five longer entries, and it&amp;rsquo;s certainly not a reason to write twenty-five combined ones. It&amp;rsquo;s a sub-pattern. Give it one entry of its own and have the parent patterns reference it.&lt;/p>
&lt;p>On the platform I described at the start, every ingestion pattern — file drops, API pulls, database replication, direct pushes — converged on a single landing mechanism into the warehouse. Once we pulled that out as its own sub-pattern, each parent entry lost about a third of its length and gained a crisp ending: &amp;ldquo;hands over to the landing sub-pattern.&amp;rdquo; And when the landing mechanism later needed a change, it was one edit instead of five. Factoring is how you keep the bank small while the platform grows.&lt;/p>
&lt;hr>
&lt;h3 id="you-cannot-build-this-alone">You cannot build this alone&lt;/h3>
&lt;/br>
&lt;p>Most writing about pattern banks covers what they should contain and then leaves you to figure out how to get your team to actually care about it.&lt;/p>
&lt;p>The truth is that a pattern bank built by one person — even a thoughtful, experienced person — will struggle to gain traction. Not because the patterns are wrong, but because nobody else owns them. When something doesn&amp;rsquo;t quite fit a pattern from the bank, the engineer&amp;rsquo;s instinct is to work around it rather than raise the question, because the bank feels like someone else&amp;rsquo;s thing. The patterns were handed down; they weren&amp;rsquo;t grown.&lt;/p>
&lt;p>The irony is that the teams who need a pattern bank most urgently — teams with inconsistent approaches, estimation problems, high bus-factor risk — are also the teams most likely to resist a top-down version of one. Engineers who&amp;rsquo;ve been doing the work their way for two years don&amp;rsquo;t want to be told their approach is being standardised out of existence.&lt;/p>
&lt;p>So you have to build it differently.&lt;/p>
&lt;p>The starting point is an audit, not a design. Before you introduce a pattern bank as a concept, spend time with your team&amp;rsquo;s existing work. Look at what&amp;rsquo;s been built over the last six to twelve months. Look for the shapes in it. Which approaches keep appearing? Where did engineers independently converge on similar solutions? Where did they diverge on problems that are structurally the same?&lt;/p>
&lt;p>This is the part where you&amp;rsquo;re listening, not speaking. You&amp;rsquo;re mapping the territory.&lt;/p>
&lt;p>When you bring the team together for the first conversation about patterns, the framing matters enormously. The worst version of that conversation starts with: &amp;ldquo;I&amp;rsquo;ve been looking at how we do things, and I think we should standardise on some approaches.&amp;rdquo; Even if that&amp;rsquo;s true and well-intentioned, it positions you as the author and everyone else as the audience.&lt;/p>
&lt;p>The better version starts with a question: &amp;ldquo;What do we keep building over and over?&amp;rdquo; Give people a chance to name the recurring shapes in their own work. You&amp;rsquo;ll find that engineers who&amp;rsquo;ve been on the team for a while can articulate these patterns clearly — they just haven&amp;rsquo;t had a reason to before. When they name a pattern, write it down. Give it back to them as documentation. That&amp;rsquo;s their knowledge, made visible.&lt;/p>
&lt;p>When you name what already exists rather than imposing what should exist, the dynamic shifts. You&amp;rsquo;re not standardising their work. You&amp;rsquo;re recognising it.&lt;/p>
&lt;p>This is the buy-in mechanism. The process of building the pattern bank &lt;em>together&lt;/em> is the adoption mechanism. When an engineer contributed to the definition of a pattern, they&amp;rsquo;ll defend it, refine it, and reach for it naturally in design conversations. When they received it from above, they&amp;rsquo;ll comply with it when they remember to — and quietly work around it when something doesn&amp;rsquo;t quite fit.&lt;/p>
&lt;hr>
&lt;h3 id="making-the-first-session-work">Making the first session work&lt;/h3>
&lt;/br>
&lt;p>The first time you sit down as a team to build the pattern bank, keep it light. You&amp;rsquo;re not trying to produce a finished document. You&amp;rsquo;re trying to start a conversation that will take months to reach maturity.&lt;/p>
&lt;p>A good opening question: &amp;ldquo;If you were explaining to a new team member what kinds of pipelines we build, how would you describe them?&amp;rdquo; This is deliberately informal. You&amp;rsquo;re not asking for a taxonomy — you&amp;rsquo;re asking for the words people already use.&lt;/p>
&lt;p>Let the conversation happen. Take notes. You&amp;rsquo;ll start to hear natural groupings emerge. Someone will say &amp;ldquo;well, most of our ingestion is basically file drops or API calls.&amp;rdquo; Someone else will add &amp;ldquo;plus the two database replications, but those work the same way as the file drops once they&amp;rsquo;re in landing.&amp;rdquo; That&amp;rsquo;s a pattern taking shape.&lt;/p>
&lt;p>When objections come up — &amp;ldquo;but the System A feed is completely different because&amp;hellip;&amp;rdquo; — treat them as contributions, not resistance. Either the system A feed is genuinely a different pattern (name it separately), or the difference is a variant within the existing pattern (document it as a variant). Both outcomes are useful. The engineer who raised the objection has just enriched the catalogue.&lt;/p>
&lt;p>There&amp;rsquo;s an important interpersonal dynamic to watch in these early sessions. Engineers who&amp;rsquo;ve been on the team a long time have often built strong opinions about how certain things should be done — and those opinions are usually right, because they&amp;rsquo;re grounded in hard experience. If you&amp;rsquo;re newer to the team, or coming in with a perspective shaped by a different environment, the temptation is to steer the conversation toward what you think the patterns should look like. Resist that. The first job is to surface what exists, not to redesign it.&lt;/p>
&lt;p>Even if you can see clearly that a particular approach has flaws — that two of the team&amp;rsquo;s existing ingestion patterns are more similar than they appear and probably belong under one umbrella — raise it as a question, not a conclusion. &amp;ldquo;I&amp;rsquo;m noticing these two look structurally similar — do you see them as the same pattern, or are there meaningful differences I&amp;rsquo;m missing?&amp;rdquo; That&amp;rsquo;s an invitation to examine, not an instruction to comply. And the answer might surprise you. The engineer who built both of them might have a very clear reason why they&amp;rsquo;re distinct that isn&amp;rsquo;t obvious from the outside.&lt;/p>
&lt;p>The sceptics in the room deserve more than a paragraph, because for a long time I misread what was actually happening for them. The standard advice — don&amp;rsquo;t argue them into participation, ask them what they&amp;rsquo;d add or change, give them an editing role rather than an audience role — is right as far as it goes. But it treats scepticism as a mood to be managed, and it&amp;rsquo;s usually something more legitimate than that.&lt;/p>
&lt;p>The engineers most resistant to a pattern bank are very often the ones whose expertise &lt;em>is&lt;/em> the current way of working. The person who can untangle the gnarliest legacy pipeline from memory, who knows every undocumented quirk of every source feed, who gets pulled into every incident because nobody else holds the map — that person has real, earned status, and it rests on knowledge that lives in their head. A pattern bank threatens to convert that hard-won knowledge into a page anyone can read. That&amp;rsquo;s a genuine loss, and you cannot argue someone out of feeling it.&lt;/p>
&lt;p>What works is role transfer, not persuasion. The new way of working needs heroes too: someone has to lead the audit of existing work, someone has to own the pattern standards, someone has to write the entries for the genuinely tricky patterns that nobody else fully understands. The people who hold the deepest knowledge of the current state should get first claim on those roles. You&amp;rsquo;re not asking them to surrender their expertise — you&amp;rsquo;re asking them to encode it, with their name on it.&lt;/p>
&lt;p>And there&amp;rsquo;s one question that does more work than any amount of advocacy: &amp;ldquo;what would you need to see before you&amp;rsquo;d trust this?&amp;rdquo; Ask it sincerely and write the answers down, because the answers &lt;em>are&lt;/em> the acceptance criteria for the pattern bank — co-authored by the people most likely to find its weaknesses. The sceptic who told you exactly what would convince them has just committed, in public, to being convincible. The person most likely to undermine a pattern bank built without them is the same person who becomes its most vocal defender when they&amp;rsquo;ve shaped it.&lt;/p>
&lt;p>At the end of the session, you should have a rough list of candidate patterns and a few volunteers to draft the first write-ups. Keep the documentation minimal — one page per pattern is plenty to start. The goal is a living document, not a finished artefact.&lt;/p>
&lt;p>The question of where to keep the pattern bank is worth getting right early. It should live somewhere the team already goes — a Teams wiki, a shared OneNote, a Confluence space, a structured markdown repo. Wherever your existing documentation lives is the right place. Don&amp;rsquo;t create a new system that requires a new habit. The pattern bank should be the path of least resistance in a design conversation, not a side trip.&lt;/p>
&lt;p>Don&amp;rsquo;t wait until the bank is &amp;ldquo;complete&amp;rdquo; to start using it. Start referencing patterns in design conversations almost immediately. When a new project comes in, run the question out loud: &amp;ldquo;Which ingestion pattern are we looking at here? Which transformation pattern?&amp;rdquo; Even before the bank is formally documented, the vocabulary starts to take hold. The naming matters as much as the documentation — maybe more. Once a team has words for what it&amp;rsquo;s doing, it starts to think in those words.&lt;/p>
&lt;hr>
&lt;h3 id="what-to-put-in-a-pattern-entry">What to put in a pattern entry&lt;/h3>
&lt;/br>
&lt;p>The minimum useful documentation for a pattern:&lt;/p>
&lt;p>&lt;strong>Name.&lt;/strong> Short, memorable, specific to your team&amp;rsquo;s context. &amp;ldquo;SFTP file ingestion&amp;rdquo; is fine. &amp;ldquo;Direct file transfer pattern&amp;rdquo; is fine. Resist the urge to make it sound architectural.&lt;/p>
&lt;p>&lt;strong>What problem it solves.&lt;/strong> One sentence. What kind of situation does this pattern address?&lt;/p>
&lt;p>&lt;strong>When to use it.&lt;/strong> What conditions make this the right choice?&lt;/p>
&lt;p>&lt;strong>When not to use it.&lt;/strong> This is the part most pattern documentation skips, and it&amp;rsquo;s often the most valuable. What situations look like they&amp;rsquo;d fit this pattern but actually don&amp;rsquo;t?&lt;/p>
&lt;p>&lt;strong>The basic shape.&lt;/strong> Three to six steps describing the high-level approach. Not a full implementation guide. Think &amp;ldquo;what are the stages?&amp;rdquo; not &amp;ldquo;what does the code look like?&amp;rdquo;&lt;/p>
&lt;p>&lt;strong>Known variants.&lt;/strong> What changes in different situations? An SFTP pull from an external system and an S3-to-S3 transfer might share a parent pattern but have meaningfully different considerations. Document the variants without creating a separate entry for each.&lt;/p>
&lt;p>&lt;strong>What it hands over, and what it needs to start.&lt;/strong> Patterns don&amp;rsquo;t run in isolation — each one ends where another begins, and the boundary is where work stalls. Two short lists fix most of that: the things this pattern produces that the next one depends on (the handover), and the things that must already exist before work on this pattern can sensibly start (the pickup criteria). If a transformation pattern needs the source data landed, typed, and described before it can begin, say so in the entry — because that single line is the difference between an engineer starting tomorrow and an engineer discovering, two days in, that they&amp;rsquo;re blocked.&lt;/p>
&lt;p>&lt;strong>How it&amp;rsquo;s validated.&lt;/strong> What kind of testing closes this pattern out, and with what kind of data? A useful rule of thumb that took us a while to articulate: fabricated rows belong in unit tests and nowhere else; real production history is the default for business validation; and if you genuinely need a synthetic scenario, create it through the source system&amp;rsquo;s own test environment so it arrives via the real pipeline — never hand-inserted into the warehouse. A pattern entry doesn&amp;rsquo;t need the whole standard. It just needs to say which of those applies, so validation is part of the design conversation rather than an afterthought at the end.&lt;/p>
&lt;p>&lt;strong>Effort signal.&lt;/strong> Is this a well-understood pattern with a reasonable effort history? Or does it have significant unknown elements? Even a simple &amp;ldquo;known / known with caveats / novel&amp;rdquo; classification is enough. It gives you something to point to when you&amp;rsquo;re estimating.&lt;/p>
&lt;p>To make this concrete, here&amp;rsquo;s roughly what a pattern entry might look like for a direct file ingestion pattern — pared back, readable, just enough to guide a design decision:&lt;/p>
&lt;hr>
&lt;p>&lt;strong>Pattern: Direct File Ingestion&lt;/strong>&lt;/p>
&lt;p>&lt;em>What it solves:&lt;/em> Landing structured or semi-structured files from an external source into the data platform&amp;rsquo;s landing zone.&lt;/p>
&lt;p>&lt;em>Use when:&lt;/em> Source delivers files on a schedule (push or pull), and transformation happens separately once the file is confirmed complete.&lt;/p>
&lt;p>&lt;em>Don&amp;rsquo;t use when:&lt;/em> Source delivers events in real-time that need immediate processing — that&amp;rsquo;s a streaming pattern. Also not appropriate when source volume is large enough that full file delivery creates latency problems.&lt;/p>
&lt;p>&lt;em>Basic shape:&lt;/em>&lt;/p>
&lt;ol>
&lt;li>File arrives in agreed drop location (SFTP, S3, or internal file transfer)&lt;/li>
&lt;li>Arrival trigger fires (event-driven or scheduled poll)&lt;/li>
&lt;li>File is validated for completeness (presence check, row count if available)&lt;/li>
&lt;li>File is copied to landing zone in raw form — no type conversion at this stage&lt;/li>
&lt;li>Staging job applies data types, basic cleansing, and audit columns&lt;/li>
&lt;li>Downstream transformation pattern picks up from staging&lt;/li>
&lt;/ol>
&lt;p>&lt;em>Known variants:&lt;/em> SFTP pull (we poll the source), S3 push (source lands in our bucket), internal transfer from another system on the platform. Error handling and retry logic differ slightly for each, but the core shape is the same.&lt;/p>
&lt;p>&lt;em>Hands over:&lt;/em> A landed, typed staging table; a manifest reconciliation result; a note of any rejected or quarantined rows. &lt;em>Needs to start:&lt;/em> Agreed drop location and credentials; a sample file; a named contact at the source who can answer format questions.&lt;/p>
&lt;p>&lt;em>Validated by:&lt;/em> Completeness check against the manifest; row counts reconciled to source; staging types asserted against the agreed contract.&lt;/p>
&lt;p>&lt;em>Effort signal:&lt;/em> Well-known. Delivery estimate should be grounded in actuals from previous builds. Flag if source has no delivery receipt mechanism — that adds complexity and uncertainty.&lt;/p>
&lt;hr>
&lt;p>That&amp;rsquo;s one page. Maybe two if the variant notes are detailed. It&amp;rsquo;s not comprehensive documentation — it&amp;rsquo;s a design guide. The engineer using it brings the implementation knowledge; the pattern provides the frame.&lt;/p>
&lt;p>The test I mentioned earlier is worth applying to every entry: can a mid-level engineer use this to make a design decision within about thirty seconds of reading it? If the answer is no, trim it or elevate the abstraction level. If someone needs a wall of text to understand what kind of problem the pattern addresses, the problem hasn&amp;rsquo;t been stated clearly enough.&lt;/p>
&lt;hr>
&lt;h3 id="patterns-in-use-from-blank-page-to-solution-design">Patterns in use: from blank page to solution design&lt;/h3>
&lt;/br>
&lt;p>The recipe analogy only works if you can actually see someone cooking. So here&amp;rsquo;s what using a pattern bank looks like in practice, when a project lands on your desk.&lt;/p>
&lt;p>Say your team is onboarding a new source system — a third-party vendor who delivers customer transaction data daily via SFTP. The commercial team wants it available in the reporting environment within forty-eight hours of the file landing. There are some business rules around transaction categorisation that need to be applied, and the output needs to land in an existing aggregate table that feeds a Power BI dashboard.&lt;/p>
&lt;p>Without a pattern bank, this is a design conversation that starts with &amp;ldquo;so how do we want to approach this?&amp;rdquo; With a pattern bank, it&amp;rsquo;s a pattern-matching exercise.&lt;/p>
&lt;p>&lt;strong>Ingestion:&lt;/strong> The source is delivering files via SFTP on a schedule. That&amp;rsquo;s the direct file ingestion pattern. You&amp;rsquo;ve done it before; you know the shape. The main variant question is whether you&amp;rsquo;re pulling from their server or they&amp;rsquo;re pushing to yours — that changes the authentication setup but not the core pattern. You note that their SFTP has had reliability issues in the past (the commercial team mentioned it), so you flag the retry handling variant in your notes.&lt;/p>
&lt;p>&lt;strong>Transformation:&lt;/strong> The business rules around transaction categorisation sound like a reference table lookup — you&amp;rsquo;re mapping raw values from the source to agreed internal categories. That&amp;rsquo;s a known pattern. The historical record requirement is the key question: does the business need to see how a transaction was categorised at the time it was processed, even if the categorisation rules change later? If yes, that&amp;rsquo;s a bi-temporal concern and it changes the transformation pattern significantly. If no — if current categorisation is all that matters — it&amp;rsquo;s a much simpler SCD1 or SCD2 load depending on whether you need row history.&lt;/p>
&lt;p>&lt;strong>Delivery:&lt;/strong> The output goes into an existing aggregate table feeding a Power BI dashboard. Depending on how that table is structured, you&amp;rsquo;re either appending new rows, overwriting a date partition, or recalculating an aggregate. Each is a recognisable delivery pattern with known characteristics.&lt;/p>
&lt;p>By the time that design conversation is twenty minutes in, you have a solution sketch: direct file ingestion (SFTP pull variant), SCD1 transformation with reference table lookup, aggregate table delivery. Three patterns, all known, all with effort history. The only real design question that needs time is the bi-temporal one — you need to clarify the business requirement before you can confirm the transformation pattern.&lt;/p>
&lt;p>And before the estimate goes anywhere, you walk the boundaries: what does ingestion hand to transformation, what does transformation hand to delivery, and who is waiting on whom at each point. Which brings me to the part of this picture I got wrong for the longest time.&lt;/p>
&lt;hr>
&lt;h3 id="the-seams-between-patterns">The seams between patterns&lt;/h3>
&lt;/br>
&lt;p>Everything I&amp;rsquo;ve described so far treats patterns as blocks: pick one per stage, snap them together, estimate from history. That&amp;rsquo;s true, and it&amp;rsquo;s also where the picture quietly lies to you — because in practice, the blocks are not where delivery goes wrong. The boundaries between them are.&lt;/p>
&lt;p>I learned this when we took our pattern bank a step further and turned it into a delivery plan: each pattern decomposed into roughly day-sized pieces of work, each piece with an exit criterion. The decomposition surfaced something the catalogue view had hidden. The estimates &lt;em>inside&lt;/em> patterns were fine — we had effort history, the day-sized pieces held up. The overruns lived &lt;em>between&lt;/em> patterns: the ingestion work finished on Tuesday and the transformation work started the following Thursday, and nobody could quite say where the week went.&lt;/p>
&lt;p>The week went into the seam. A review that sat in someone&amp;rsquo;s queue. A handover that turned out to be missing the one thing the next engineer needed, triggering a conversation, then a clarification, then a small rework. A wait for a batch window, or for another team, or for a sign-off from someone who didn&amp;rsquo;t know they were on the critical path.&lt;/p>
&lt;p>This gives you a distinction worth building into how you plan: effort versus elapsed time. Effort lives inside patterns — it&amp;rsquo;s the day-sized, estimable work the bank already describes. Elapsed time accumulates at seams, and it follows completely different rules. You can&amp;rsquo;t reduce a review queue by working harder, and you can&amp;rsquo;t estimate a sign-off from effort history. If your plans only count effort, the seams are invisible right up until the deadline isn&amp;rsquo;t.&lt;/p>
&lt;p>&lt;img src="https://ghostinthedata.info/posts/2026/2026-06-27-pattern-bank/images/pattern-bank-seams.svg" alt="Patterns are blocks of estimable effort; elapsed time accumulates at the seams between them">&lt;/p>
&lt;p>&lt;em>The same work, on a calendar. The solid segments are effort — the patterns. The gaps are the seams.&lt;/em>&lt;/p>
&lt;p>So once the patterns are named, name the seams too. Every place where work passes between patterns — or between people — gets the same lightweight treatment a pattern gets: what crosses the boundary (the handover pack), and what the receiving side needs before they can start (the pickup criteria). The handover from ingestion to transformation, for instance, might be: the staging table, the agreed contract it conforms to, a note of known quirks, and confirmation the data is actually flowing. Write it down once, and every future handover at that seam stops being a negotiation.&lt;/p>
&lt;p>The quality bar for a handover is a test I now apply to everything: a piece of work is done when a different engineer could pick it up tomorrow with no conversation. Not &amp;ldquo;done pending a chat,&amp;rdquo; not &amp;ldquo;done but ask me about the weird bit.&amp;rdquo; No conversation. It sounds strict, and it is — but every conversation a handover requires is elapsed time hiding in plain sight, and elapsed time at seams is precisely the thing your effort-based estimates can&amp;rsquo;t see.&lt;/p>
&lt;p>The pattern bank describes the blocks. The seam map describes the gaps between the blocks. You need both, and the second one is the one nobody writes down.&lt;/p>
&lt;hr>
&lt;h3 id="the-estimation-conversation">The estimation conversation&lt;/h3>
&lt;/br>
&lt;p>There&amp;rsquo;s a distinction in planning that doesn&amp;rsquo;t get talked about enough: the difference between things you don&amp;rsquo;t know, and things you don&amp;rsquo;t know you don&amp;rsquo;t know.&lt;/p>
&lt;p>The first kind — known unknowns — you can account for in estimates. You know they&amp;rsquo;re there, you can add buffer, you can plan a spike to resolve them early.&lt;/p>
&lt;p>The second kind — unknown unknowns — are the ones that blow up timelines. They surface mid-project, when you&amp;rsquo;re already committed to a delivery date. They&amp;rsquo;re the integration behaviour you didn&amp;rsquo;t anticipate, the edge case in the source system nobody mentioned, the schema change that came through without notice.&lt;/p>
&lt;p>Pattern reuse doesn&amp;rsquo;t eliminate unknown unknowns. But it converts some of them into known unknowns, and some known unknowns into known quantities. When you&amp;rsquo;ve built the same ingestion type four times, you know where the surprises usually come from. You&amp;rsquo;ve already met most of the ways that type of pipeline can misbehave.&lt;/p>
&lt;p>&lt;img src="https://ghostinthedata.info/posts/2026/2026-06-27-pattern-bank/images/pattern-bank-estimation-shift.svg" alt="Pattern reuse converts unknown unknowns into known unknowns, and known unknowns into known quantities">&lt;/p>
&lt;p>&lt;em>Each repetition of a pattern moves surprises to the right — out of the category that blows up timelines and into the category you can plan.&lt;/em>&lt;/p>
&lt;p>That conversion is where a pattern bank earns its keep. It&amp;rsquo;s not a tidiness exercise. It changes the quality of your estimates, and it changes how you talk about risk with internal customers.&lt;/p>
&lt;p>At the organisation I mentioned earlier, once we had a clearer picture of which project components were known patterns versus genuinely novel territory, the project conversations changed character. Instead of quoting a number and hoping for the best, we could say: &amp;ldquo;This solution uses three patterns from our standard set — the effort estimate for those is grounded in what we&amp;rsquo;ve delivered before. This fourth component is new for us — we&amp;rsquo;re treating it as discovery work and we&amp;rsquo;ll re-estimate once we&amp;rsquo;ve built a proof of concept.&amp;rdquo; That&amp;rsquo;s a different kind of conversation. It&amp;rsquo;s a more honest one, and paradoxically, it inspires more confidence.&lt;/p>
&lt;p>There&amp;rsquo;s a maturation point worth knowing about in advance. Once a pattern has genuine effort history — three or four builds behind it — you can decompose it into day-sized pieces of work, each with its own exit criterion. That&amp;rsquo;s the moment the bank stops being a catalogue and starts being a delivery playbook: a new project isn&amp;rsquo;t just &amp;ldquo;three known patterns,&amp;rdquo; it&amp;rsquo;s a sequence of named, day-sized stories you can lay against a calendar. And when you do lay it against a calendar, estimate the seams separately, in elapsed time rather than effort — because the question at a seam isn&amp;rsquo;t &amp;ldquo;how long will this take to do&amp;rdquo; but &amp;ldquo;how long will this take to happen.&amp;rdquo;&lt;/p>
&lt;p>Stakeholders aren&amp;rsquo;t usually uncomfortable with uncertainty. They&amp;rsquo;re uncomfortable with surprises. Pattern banking reduces surprises by making the known quantities explicit — which in turn makes the uncertain parts easier to name, plan around, and manage.&lt;/p>
&lt;hr>
&lt;h3 id="from-documentation-to-machinery">From documentation to machinery&lt;/h3>
&lt;/br>
&lt;p>Remember the opening of this article: a team with two sensible base templates, and twenty implementations that had drifted from them in ways nobody could track. Here&amp;rsquo;s the uncomfortable implication for everything I&amp;rsquo;ve said so far — a documented pattern bank is still just documentation, and documentation drifts. The patterns describe what the team agreed to do; nothing stops the codebase from quietly doing something else, one reasonable-seeming exception at a time. The bank I&amp;rsquo;ve described prevents the team from &lt;em>forgetting&lt;/em> the patterns. It doesn&amp;rsquo;t prevent them from &lt;em>departing&lt;/em> from them.&lt;/p>
&lt;p>So it&amp;rsquo;s worth knowing that a pattern bank has a maturity path, and documentation is the middle of it, not the end.&lt;/p>
&lt;p>&lt;img src="https://ghostinthedata.info/posts/2026/2026-06-27-pattern-bank/images/pattern-bank-maturity.svg" alt="A pattern bank matures from vocabulary to documentation to machinery">&lt;/p>
&lt;p>The first stage is vocabulary: the team has names for the work, spoken but not written. This is further than most teams ever get, and it&amp;rsquo;s where the adoption battle is won. The second stage is documentation: the one-page entries, the seam map, the effort signals. The third stage is machinery: the pattern&amp;rsquo;s invariants are enforced by your delivery tooling, so that violating them doesn&amp;rsquo;t produce a governance conversation — it produces a failed build.&lt;/p>
&lt;p>Not everything in a pattern can or should be mechanised. But some pattern elements are invariants — things that must hold for the platform to stay trustworthy — and invariants are exactly the things that drift erodes first. A sequencing rule is a good example. We had one that said, in effect, data ships before code: the ingestion side of a new source goes to production first, and transformation work doesn&amp;rsquo;t start until the data it depends on actually exists where production builds can see it. As documentation, that rule held about as well as documented rules ever do. Then we wired it into CI — the transformation build simply fails if the upstream data isn&amp;rsquo;t there — and the rule stopped being a rule. It became a property of the system. The pattern cannot drift on that invariant, because the machinery won&amp;rsquo;t let it.&lt;/p>
&lt;p>Encoding invariants has a sharp corollary that I didn&amp;rsquo;t anticipate: optional elements and mechanical gates don&amp;rsquo;t mix. We&amp;rsquo;d had data contracts in our patterns as &amp;ldquo;preferred but optional&amp;rdquo; — encouraged, written up, mostly adopted. The moment we wanted CI to validate new sources against their contract, &amp;ldquo;optional&amp;rdquo; stopped being coherent. A gate can&amp;rsquo;t check a thing that may or may not exist. So the gate forced the decision we&amp;rsquo;d been politely deferring: contracts became required for new sources, with existing ones grandfathered. That&amp;rsquo;s the general lesson — every time you automate the enforcement of a pattern element, you&amp;rsquo;ll be forced to decide whether the things it depends on are truly required. The machinery is honest in a way the documentation never has to be.&lt;/p>
&lt;p>Don&amp;rsquo;t read this as &amp;ldquo;mechanise everything.&amp;rdquo; Pick the invariants whose violation is expensive — the raw layer staying raw, the sequencing between stages, the contract a source promised — and let the rest stay as guidance. The thirty-second test still applies to the documentation; the machinery is just there to hold the lines that matter most while humans exercise judgement on everything else.&lt;/p>
&lt;hr>
&lt;h3 id="a-fourth-job-scoping-change">A fourth job: scoping change&lt;/h3>
&lt;/br>
&lt;p>I&amp;rsquo;ve described three jobs for a pattern bank so far: a shared design vocabulary, grounded estimation, and legibility to the people around the team. There&amp;rsquo;s a fourth, and I only discovered it later, when the bank was already in place and we needed it for something it was never designed to do.&lt;/p>
&lt;p>We were proposing a significant change to how the team delivered — the kind of change that touches process people have used for years, with passionate, technically senior advocates on multiple sides. The instinctive reaction to a proposal like that is that &lt;em>everything&lt;/em> is changing: the way I build, the way I test, the way my work gets to production, the way I&amp;rsquo;m on the hook when it breaks. When change feels total, people defend totally.&lt;/p>
&lt;p>The pattern bank changed the shape of that conversation entirely. Because the team&amp;rsquo;s whole way of working was laid out as a catalogue of named patterns, we could put the proposal against it and show precisely where the change landed: of the entire bank, exactly one pattern&amp;rsquo;s promotion steps were being rewritten, plus one new gate at a seam between two stages. Every ingestion pattern: untouched. Every delivery pattern: untouched. The transformation logic itself: untouched. We could literally point at the diagram and say — the workshop is deciding the contents of one box.&lt;/p>
&lt;p>&lt;img src="https://ghostinthedata.info/posts/2026/2026-06-27-pattern-bank/images/pattern-bank-change-scope.svg" alt="Scoping a process change against the pattern bank: one box changes, everything else is untouched">&lt;/p>
&lt;p>&lt;em>The most useful slide in the whole proposal wasn&amp;rsquo;t about the change. It was about everything that wasn&amp;rsquo;t changing.&lt;/em>&lt;/p>
&lt;p>That framing did more to lower the temperature than any argument about the merits of the change itself. Not because it dodged the hard conversation — the contents of that one box still had to be debated, properly and at length — but because it bounded the conversation. People could engage with the actual proposal instead of defending against an imagined one.&lt;/p>
&lt;p>This is, I think, a structural property rather than a one-off trick. Without a pattern bank, a process change has no edges: nobody can say with confidence what it touches and what it doesn&amp;rsquo;t, so everyone reasonably assumes it touches them. With a pattern bank, a change is a diff. You can enumerate exactly which entries are modified, which seams gain or lose a gate, and which patterns are provably unaffected. The same catalogue that scopes new work scopes change to the work itself.&lt;/p>
&lt;p>If you&amp;rsquo;re a lead who expects to steer your team through a significant shift in the next year or two — a new deployment model, a platform migration, a rework of how releases happen — that alone might justify building the bank now. It&amp;rsquo;s much easier to show people that only one box is changing if the boxes already exist.&lt;/p>
&lt;hr>
&lt;h3 id="keeping-it-alive-without-making-it-a-burden">Keeping it alive without making it a burden&lt;/h3>
&lt;/br>
&lt;p>The biggest risk to a pattern bank isn&amp;rsquo;t that it will be rejected. It&amp;rsquo;s that it will be adopted once and then quietly ignored as the team moves on to delivery pressures.&lt;/p>
&lt;p>The solution is governance that&amp;rsquo;s light enough to actually happen.&lt;/p>
&lt;p>Pattern review doesn&amp;rsquo;t need its own meeting. Fold it into your existing team rhythms — a retrospective, a sprint review, a regular team catch-up. Set aside fifteen minutes every month or two to ask: have any of our recently built solutions introduced a new pattern we should document? Has anything we&amp;rsquo;ve built repeatedly over the past quarter highlighted that one of our existing patterns needs refinement?&lt;/p>
&lt;p>The trigger for adding a new pattern remains the same: you&amp;rsquo;ve built the same thing three times. Before that, document it as a one-off or a variant. After that, give it its own entry. The three-instance rule keeps the bank from growing with speculative patterns that haven&amp;rsquo;t proven their recurring value.&lt;/p>
&lt;p>Retiring patterns is just as important as adding them. A pattern that hasn&amp;rsquo;t been used in twelve months is probably no longer part of how your team works. Don&amp;rsquo;t delete it — move it to a legacy section. Old pipelines built on deprecated patterns still exist, and future engineers will want context when they encounter them.&lt;/p>
&lt;p>The most common way a pattern bank dies is through drift: the documented patterns diverge from what the team actually does, and nobody updates the documentation because the overhead isn&amp;rsquo;t worth it. The way to avoid this is to make update rituals small. One sentence changed, one variant added, one effort note updated after a project completes. It takes five minutes if it&amp;rsquo;s part of the project close-out conversation. It takes months if it&amp;rsquo;s treated as a separate documentation effort.&lt;/p>
&lt;p>One small artefact worth keeping alongside the bank: a glossary. A dozen or so terms of art, one line each — what &amp;ldquo;landing&amp;rdquo; means here, what &amp;ldquo;staging&amp;rdquo; means here, what &amp;ldquo;the contract&amp;rdquo; refers to. It sounds trivial until you&amp;rsquo;ve watched two documents written three months apart quietly disagree about what a word means, or a design review burn twenty minutes discovering that two people were using &amp;ldquo;raw&amp;rdquo; differently. The patterns give the team nouns for solutions; the glossary keeps the rest of the vocabulary honest. It costs a page, and it makes everything else the team writes — pattern entries, design docs, handover notes — interoperable.&lt;/p>
&lt;p>One practical approach: put the pattern bank in a place the team already touches regularly. A Teams wiki, a Confluence space, a shared document with a clear structure — wherever the team already goes for reference material. Don&amp;rsquo;t create a dedicated system that requires a new habit to use. The path of least resistance should lead to the pattern bank, not away from it.&lt;/p>
&lt;hr>
&lt;h3 id="what-it-looks-like-when-its-working">What it looks like when it&amp;rsquo;s working&lt;/h3>
&lt;/br>
&lt;p>The sign that a pattern bank has taken hold isn&amp;rsquo;t that engineers consult it before starting every project. It&amp;rsquo;s that the vocabulary from it enters the team&amp;rsquo;s natural speech.&lt;/p>
&lt;p>You&amp;rsquo;ll hear it in design conversations: &amp;ldquo;this looks like an SCD2 load with a late-arriving records wrinkle&amp;rdquo; rather than &amp;ldquo;so this is kind of like what we did for System B but slightly different.&amp;rdquo; You&amp;rsquo;ll hear it in planning: &amp;ldquo;the ingestion here is a known pattern, it&amp;rsquo;s the delivery piece that&amp;rsquo;s new territory.&amp;rdquo; You&amp;rsquo;ll hear it at the seams: &amp;ldquo;what&amp;rsquo;s in the handover pack for this one — could someone pick it up tomorrow without talking to you?&amp;rdquo; You&amp;rsquo;ll hear it when a new team member asks how things work, and an experienced engineer can answer them in fifteen minutes using pattern language rather than two weeks of reading old code.&lt;/p>
&lt;p>You&amp;rsquo;ll also hear it in conversations with people outside the data team — architects, project managers, internal customers — who start to understand the platform not as a black box but as a system with legible parts. That transparency matters. When the architects designing how source systems should serve data understand what an ingestion pattern expects, they make better design decisions. When an internal customer understands the difference between a project that uses known patterns and one that introduces a new one, they understand why the effort estimates are different.&lt;/p>
&lt;p>The pattern bank doesn&amp;rsquo;t make the work easier. It makes the work legible. And when the work is legible — to the engineers doing it, to the people managing it, and to the people upstream of it — everything else gets a little simpler.&lt;/p>
&lt;p>Going back to that organisation where I started: the ingestion layer that had twenty different approaches to twenty similar problems didn&amp;rsquo;t get rebuilt. That wasn&amp;rsquo;t the point. But over time, as new source systems were onboarded, they were designed against the patterns the team had agreed on. The sprawl stopped sprawling. The conversations about effort estimates became more grounded. The invariants that mattered most stopped relying on memory and started failing builds instead. And when the day came that we needed to change how the team delivered — not just what it delivered — the bank turned a frightening proposal into a bounded one.&lt;/p>
&lt;p>When a new engineer joined the team, there was something to hand them that explained not just what the platform did, but why things were built the way they were.&lt;/p>
&lt;p>That&amp;rsquo;s what the pattern bank was, at its core. Not a governance document. Not a bureaucratic exercise. A way of saying: here is what we know, here is how we think about this, and here is how we&amp;rsquo;ve agreed to do it together.&lt;/p>
&lt;p>Start there. Build it with your team. Refine it every time you learn something new.&lt;/p>
&lt;p>&lt;/br>&lt;/br>&lt;/p></content:encoded><category>Data Engineering</category><category>Leadership</category><category>Career Development</category><category>Data Engineering</category><category>Pattern Bank</category><category>Team Culture</category><category>Project Management</category><category>Solution Design</category><category>Engineering Leadership</category><category>Estimation</category></item><item><title>Keep Moving</title><link>https://ghostinthedata.info/posts/2026/2026-06-20-fire-and-motion/</link><pubDate>Tue, 23 Jun 2026 08:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-06-20-fire-and-motion/</guid><author>Chris Hillman</author><description>Why forward momentum — not planning, not tooling, not the perfect architecture — is the only thing that actually ships a data platform.</description><content:encoded>&lt;p>Some days I open my laptop and by 5pm I genuinely cannot tell you what I did.&lt;/p>
&lt;p>Not because it was complicated. Not because there were emergencies. The stand-up happened. A few Teams messages were sent. A ticket was groomed. A document was &amp;ldquo;reviewed&amp;rdquo;. A meeting was attended where everyone agreed something was important and then the meeting ended and nothing changed.&lt;/p>
&lt;p>And then somehow it was evening and the pipeline I meant to fix was exactly as broken as it was in the morning.&lt;/p>
&lt;p>I used to think this was a time management problem. I&amp;rsquo;ve since come to believe it&amp;rsquo;s a physics problem.&lt;/p>
&lt;hr>
&lt;h2 id="an-object-at-rest">An object at rest&lt;/h2>
&lt;p>Actual productive work — the kind that moves a platform forward — probably occupies about two or three hours of a given workday. Maybe four if you&amp;rsquo;re lucky. The rest is coordination, context-switching, Teams, and the ritual of &lt;em>almost&lt;/em> starting.&lt;/p>
&lt;p>I&amp;rsquo;ve seen senior engineers arrive late and leave early and outship everyone else on the team. I&amp;rsquo;ve seen people in meetings from 9 to 5 who are technically &amp;ldquo;at work&amp;rdquo; but haven&amp;rsquo;t touched a pipeline in three weeks. Hours in office — or in Jira — are a spectacularly poor proxy for output.&lt;/p>
&lt;p>The hard part isn&amp;rsquo;t the work itself. Once you&amp;rsquo;re in it — once the editor is open, the query is running, the dbt model is taking shape — it flows. The painful part is the gap between &lt;em>intending&lt;/em> to start and &lt;em>actually&lt;/em> starting. That gap has a gravitational pull all its own.&lt;/p>
&lt;p>I have a version of this routine I&amp;rsquo;m not proud of: open laptop, check email, check Teams, open the ticket I meant to work on, re-read my own notes from last week, decide I need more context, open Confluence, fall into a 20-minute rabbit hole of documentation I wrote six months ago, circle back to Teams, check if anyone replied to that thread, think about getting coffee, get coffee, and then, somewhere around 11am, finally open the editor and actually write the thing that takes forty-five minutes.&lt;/p>
&lt;p>The work was never the problem. It was everything I let stand between me and starting.&lt;/p>
&lt;p>This isn&amp;rsquo;t a character flaw. It&amp;rsquo;s inertia. Physics applies to engineers too.&lt;/p>
&lt;hr>
&lt;h2 id="the-only-way-to-build-a-platform">The only way to build a platform&lt;/h2>
&lt;p>There&amp;rsquo;s a way I&amp;rsquo;ve come to think about what separates data platforms that actually mature from the ones that are permanently &amp;ldquo;almost there&amp;rdquo;.&lt;/p>
&lt;p>It&amp;rsquo;s not the architecture. It&amp;rsquo;s not the tooling. It&amp;rsquo;s not the quality of the backlog or the rigour of the sprint ceremonies or the neatness of the Confluence hierarchy.&lt;/p>
&lt;p>It&amp;rsquo;s whether the team is moving forward — every day, in some small way — or whether it&amp;rsquo;s standing still.&lt;/p>
&lt;p>That sounds obvious until you try to do it. Because there are a hundred forces, every single day, designed to stop you from moving.&lt;/p>
&lt;p>Some of them are internal. The engineer who won&amp;rsquo;t write a single line until the data model is &amp;ldquo;finalised&amp;rdquo;. The team that spends three sprints documenting the current state before touching anything. The migration project that requires a complete source system audit before the first table lands in Snowflake. All reasonable-sounding activities. All forms of not moving.&lt;/p>
&lt;p>Some of them are external. That&amp;rsquo;s where it gets more interesting.&lt;/p>
&lt;hr>
&lt;h2 id="cover-fire">Cover fire&lt;/h2>
&lt;p>The data tooling industry has discovered something effective: if you can keep teams busy reacting to you, they don&amp;rsquo;t have time to ship.&lt;/p>
&lt;p>Not consciously, probably. But the effect is the same.&lt;/p>
&lt;p>Joel Spolsky wrote about this pattern more than twenty years ago in &lt;a href="https://www.joelonsoftware.com/2002/01/06/fire-and-motion/" target="_blank" rel="noopener">Fire and Motion&lt;/a>. In infantry terms: you fire at the enemy not because you expect to hit much, but because it forces them to keep their heads down while you advance. In software terms, the big players generate a constant barrage of new technologies and standards so that everyone else burns their days keeping up instead of building. He was writing about Microsoft in 2002. The names have changed. The strategy hasn&amp;rsquo;t.&lt;/p>
&lt;p>Every six months there is a new architectural paradigm your team is supposed to evaluate. The lakehouse. The semantic layer. Data contracts. The medallion architecture (deprecated). The medallion architecture (rehabilitated). Streaming-first. The modern data stack (thriving). The modern data stack (dead). AI-native pipelines. Agents that write your dbt models for you.&lt;/p>
&lt;p>The vendor ecosystem has learned to generate enough noise that a conscientious engineering team can spend most of its time just &lt;em>keeping up&lt;/em>: reading the announcements, watching the conference talks, debating whether to adopt the new pattern, setting up a proof-of-concept, discovering the PoC is more complex than advertised, abandoning it, and then starting the cycle again when the next announcement drops.&lt;/p>
&lt;p>Meanwhile your users still can&amp;rsquo;t get a reliable enrolment count. The Oracle migration is in month fourteen. The dashboard that takes eleven seconds to load still takes eleven seconds to load.&lt;/p>
&lt;p>The Snowflake vs Databricks wars, dbt Fusion vs dbt Core, Iceberg vs Delta vs everything else: some of it is genuinely important and you should stay informed. But a lot of it is just noise that benefits the people generating it, not the people consuming it. While your team is deep in an evaluation of a new ingestion paradigm, the basic pipelines you already have aren&amp;rsquo;t getting more reliable. The stakeholders who&amp;rsquo;ve been asking for a new data source for four months are still waiting.&lt;/p>
&lt;p>They don&amp;rsquo;t care which open table format you&amp;rsquo;re using. They care whether their data is there in the morning.&lt;/p>
&lt;p>When a sales engineer from a major vendor asks whether your platform supports the latest hot acronym — and they will, and it will be in a room full of your stakeholders — that question is doing work. It&amp;rsquo;s designed to make you feel behind. To make you wonder if you should be rebuilding rather than shipping. To redirect your energy toward something that benefits them.&lt;/p>
&lt;p>The right response is rarely to immediately build the thing they&amp;rsquo;re asking about. The right response is usually to keep moving on the things that actually matter to the people you serve.&lt;/p>
&lt;hr>
&lt;h2 id="what-moving-forward-looks-like">What moving forward looks like&lt;/h2>
&lt;p>I want to be specific about this, because &amp;ldquo;keep moving&amp;rdquo; is easy to say and genuinely hard to operationalise.&lt;/p>
&lt;p>It doesn&amp;rsquo;t mean rushing. It doesn&amp;rsquo;t mean skipping the thinking or writing code you&amp;rsquo;ll hate in six months. It means that every sprint, every week, every day — something is better than it was. A pipeline is more reliable. A model is documented. A test passes that didn&amp;rsquo;t pass before. A stakeholder gets a data source they couldn&amp;rsquo;t access last month.&lt;/p>
&lt;p>Small, consistent, compounding. The platform that looks dramatically better in twelve months is almost always the one that had someone quietly improving things every single day, not the one that attempted a big-bang rewrite.&lt;/p>
&lt;p>On the personal side: the most useful thing I&amp;rsquo;ve found is to remove the decision of whether to start. If you have to decide to begin every morning, you&amp;rsquo;re giving inertia a fighting chance. The editor should be open before you check Teams. The model you&amp;rsquo;re working on should be the first tab, not the fifth. Pair programming works partly because it eliminates optionality. You can&amp;rsquo;t not start when someone is sitting there waiting.&lt;/p>
&lt;p>On the team side: the most dangerous status update is &amp;ldquo;working on it&amp;rdquo;. It sounds like movement. It often isn&amp;rsquo;t. The rituals that actually correlate with forward progress are the boring ones: small PRs merged frequently, tickets moving from in-progress to done in days not weeks, users getting things they asked for.&lt;/p>
&lt;p>On the strategic side: have an honest answer for what you&amp;rsquo;re choosing not to evaluate right now, and why. &amp;ldquo;We&amp;rsquo;re not looking at that this quarter because we&amp;rsquo;re migrating the Oracle GL tables and that&amp;rsquo;s more valuable&amp;rdquo; is a complete sentence. You don&amp;rsquo;t need to defend it at length.&lt;/p>
&lt;hr>
&lt;h2 id="the-platform-that-keeps-moving-wins">The platform that keeps moving wins&lt;/h2>
&lt;p>I don&amp;rsquo;t think data engineering is particularly hard compared to what else people do with computers. The craft is learnable. The patterns are well documented. The tooling is genuinely better than it was five years ago.&lt;/p>
&lt;p>What&amp;rsquo;s hard is sustaining forward motion through the accumulated weight of meetings and tool evaluations and perfectly reasonable reasons not to ship today.&lt;/p>
&lt;p>The teams I&amp;rsquo;ve seen build platforms that people actually trust and use are almost never the ones with the most sophisticated architecture. They&amp;rsquo;re the ones that kept moving — through organisational changes, through vendor noise, through the days when it felt like nothing was getting done.&lt;/p>
&lt;p>There&amp;rsquo;s a cost to this approach, and it&amp;rsquo;s worth naming. If you treat most of the noise as noise, occasionally you&amp;rsquo;ll be wrong. Something you dismissed will turn out to be the real thing, and you&amp;rsquo;ll arrive eighteen months after the early adopters, reading their write-ups instead of writing your own. That&amp;rsquo;s the trade, and I&amp;rsquo;d take it. Being late to a paradigm is recoverable. A platform nobody trusts is not.&lt;/p>
&lt;p>The job, at every level, is to move forward a little every day. Fix the pipeline. Write the test. Ship the model. Even when it&amp;rsquo;s small. Even when it&amp;rsquo;s imperfect.&lt;/p>
&lt;p>The platform gets better a little at a time, or not at all.&lt;/p>
&lt;p>Somehow, launch the editor.&lt;/p></content:encoded><category>Data Engineering</category><category>Leadership</category><category>Productivity</category><category>Data Engineering</category><category>Platform Strategy</category><category>Team Leadership</category><category>Technical Strategy</category></item><item><title>Ghost Skills: Teaching AI Agents to Think Like Data Engineers</title><link>https://ghostinthedata.info/posts/2026/2026-06-14-ai-ghost-skills-ai-agents/</link><pubDate>Sun, 14 Jun 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-06-14-ai-ghost-skills-ai-agents/</guid><author>Chris Hillman</author><description>A new open-source skills repo that gives AI coding agents the methodology they keep forgetting — grain, keys, SCD2, incident comms, and more. Tool-agnostic, composable, and opinionated.</description><content:encoded>&lt;p>Another week, another skills repo on the GitHub trending page. I know. There are roughly seventeen of them now, all promising to turn your AI coding agent from a confident intern into a slightly-less-confident intern. Most of them are great. Most of them are also built by solo devs, for solo devs, on solo-dev codebases that fit comfortably in a context window.&lt;/p>
&lt;p>Which is fine, if that&amp;rsquo;s your world.&lt;/p>
&lt;p>Less fine if your world involves a Snowflake warehouse with four tables that &lt;em>could&lt;/em> be the source of truth for &amp;ldquo;customer&amp;rdquo;, an SCD2 someone half-built in 2021 and quietly walked away from, and a dbt project where &lt;code>stg_users_final_v3_actually_use_this&lt;/code> is, somehow, the one you&amp;rsquo;re meant to use. (Don&amp;rsquo;t laugh. You&amp;rsquo;ve seen worse.)&lt;/p>
&lt;p>I&amp;rsquo;ve been using AI agents day-to-day for a while now — mostly as a peer reviewer, a second pair of eyes, a thinking partner that doesn&amp;rsquo;t sigh when I ask it to look at the same query three times. They&amp;rsquo;re genuinely useful.&lt;/p>
&lt;p>So I built something &lt;a href="https://github.com/ghostinthedata-info/skills" target="_blank" rel="noopener">Ghost Skills&lt;/a> — a collection of data-engineering methodology skills for AI coding agents. Before I get into what&amp;rsquo;s in it, let me explain why I bothered.&lt;/p>
&lt;br>
&lt;hr>
&lt;h3 id="the-problem-isnt-the-model-its-that-data-work-has-rules-nobody-told-it-about">The problem isn&amp;rsquo;t the model. It&amp;rsquo;s that data work has rules nobody told it about.&lt;/h3>
&lt;br>
&lt;p>Your AI agent is, by default, an extremely bright graduate on their first day. It knows Python. It knows SQL. It can write a beautifully-commented &lt;code>for&lt;/code> loop and explain window functions to you in five different ways. It will happily generate a fact table, a dbt model, an Airflow DAG, all of it looking very professional.&lt;/p>
&lt;p>What it doesn&amp;rsquo;t know is everything that actually matters.&lt;/p>
&lt;p>It doesn&amp;rsquo;t know that the &lt;code>customer_id&lt;/code> in that source isn&amp;rsquo;t really unique because of the 2022 CRM migration. It doesn&amp;rsquo;t know that Roger pushed a release on Friday afternoon (because of course he did) and the platform&amp;rsquo;s been wheezing through a seven-year backfill ever since. It doesn&amp;rsquo;t know that this dimension changes slowly and the audit team will lose their minds if you flatten the history. It doesn&amp;rsquo;t know that the business definition of &amp;ldquo;active customer&amp;rdquo; changed in March, but only in the marketing data mart, and only sometimes.&lt;/p>
&lt;p>These aren&amp;rsquo;t model intelligence problems. The model is plenty smart. They&amp;rsquo;re &lt;em>context&lt;/em> problems — and the context is exactly the stuff that takes a data engineer two years to learn and roughly thirty minutes to forget when they leave the company.&lt;/p>
&lt;p>What agents are missing isn&amp;rsquo;t capability. It&amp;rsquo;s methodology. It&amp;rsquo;s the durable craft of data engineering — the part that doesn&amp;rsquo;t change when you swap Snowflake for BigQuery, or dbt for SQLMesh, or your warehouse for a &amp;ldquo;lakehouse&amp;rdquo; that&amp;rsquo;s somehow priced like a warehouse anyway.&lt;/p>
&lt;p>That&amp;rsquo;s the gap this is attempting to fill.&lt;/p>
&lt;br>
&lt;hr>
&lt;h3 id="skills-briefly-for-anyone-who-hasnt-drunk-this-particular-kool-aid-yet">Skills, briefly, for anyone who hasn&amp;rsquo;t drunk this particular Kool-Aid yet&lt;/h3>
&lt;br>
&lt;p>If you&amp;rsquo;ve not been living inside Claude Code or Cursor for the last six months: a &lt;em>skill&lt;/em> is a folder of instructions you give an AI coding agent. &amp;ldquo;When you build a fact table, declare the grain first.&amp;rdquo; &amp;ldquo;When you write tests, follow this severity model.&amp;rdquo; &amp;ldquo;When something breaks, here&amp;rsquo;s the comms workflow.&amp;rdquo; You version it, you commit it, you keep it in the repo. The agent reads it before doing the thing, and re-reads it next session because agents have the long-term memory of a goldfish.&lt;/p>
&lt;p>It&amp;rsquo;s a deceptively powerful pattern. Instead of pasting the same prompt every session (and forgetting half of it, and writing it slightly differently this time), the standard lives in the repo. New team member joins? They get the skills. Agent updates next month? It still reads the skills.&lt;/p>
&lt;p>The catch — and we&amp;rsquo;ll come back to this — is that skills tell an agent &lt;em>how&lt;/em> you do something. They don&amp;rsquo;t tell it &lt;em>what&lt;/em> you&amp;rsquo;ve built, &lt;em>why&lt;/em> you built it that way, or &lt;em>which&lt;/em> of your seven utility tables is the one that&amp;rsquo;s still maintained.&lt;/p>
&lt;p>For solo devs, that&amp;rsquo;s fine. There is no &amp;ldquo;seven utility tables&amp;rdquo;. For data teams at any kind of scale, it&amp;rsquo;s a real limit. More on that at the end.&lt;/p>
&lt;br>
&lt;hr>
&lt;h3 id="introducing-ghost-skills">Introducing Ghost Skills&lt;/h3>
&lt;br>
&lt;p>Right. The repo: &lt;a href="https://github.com/ghostinthedata-info/skills" target="_blank" rel="noopener">github.com/ghostinthedata-info/skills&lt;/a>. Built mostly for myself, polished up for sharing.&lt;/p>
&lt;p>30-second setup:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>npx skills@latest add ghostinthedata-info/skills
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Pick the skills you want, pick your agent (Claude Code, Codex, Cursor — whatever you&amp;rsquo;re using), and run &lt;code>/setup-ghost-skills&lt;/code>. It&amp;rsquo;ll ask you three questions: warehouse dialect, transform tooling, and how your domain docs are laid out. Then it writes itself a configuration block into your &lt;code>CLAUDE.md&lt;/code> or &lt;code>AGENTS.md&lt;/code> and gets out of your way.&lt;/p>
&lt;p>The agent reads it next session. And the one after that. And the one after that. The same standards, every time, without you re-explaining them.&lt;/p>
&lt;p>That&amp;rsquo;s it. That&amp;rsquo;s the whole thing.&lt;/p>
&lt;br>
&lt;hr>
&lt;h3 id="whats-in-the-catalogue">What&amp;rsquo;s in the catalogue&lt;/h3>
&lt;br>
&lt;p>The skills group into four areas, mapped to roughly how data work actually happens.&lt;/p>
&lt;p>&lt;strong>Discovery&lt;/strong> is everything you should do before you build anything, and frequently don&amp;rsquo;t. &lt;code>profile-data&lt;/code> runs the baseline checks on a new dataset — row counts, cardinality, null analysis, key uniqueness — so the agent isn&amp;rsquo;t generating models against assumptions that fall apart in week two. &lt;code>gather-requirements&lt;/code> pins down grain, sources, consumers, freshness SLAs, and acceptance criteria one question at a time, instead of letting the agent guess. &lt;code>refine-context&lt;/code> stress-tests a plan against your documented domain model and writes the decisions back to &lt;code>CONTEXT.md&lt;/code> and ADRs as they crystallise.&lt;/p>
&lt;p>&lt;strong>Modeling&lt;/strong> is the methodology your agent should already be applying and usually isn&amp;rsquo;t. &lt;code>dimensional-modeling&lt;/code> walks the four-step process. &lt;code>fact-table-design&lt;/code> forces grain declaration and measure classification before a single line of SQL gets generated. &lt;code>keys&lt;/code> covers business, natural, surrogate, composite, and durable keys — and the anti-patterns that bite you eighteen months later when the source system gets re-platformed. &lt;code>slowly-changing-dimensions&lt;/code> covers SCD types 0 through 7, including my Healing Tables approach to deterministic, path-independent SCD2.&lt;/p>
&lt;p>&lt;strong>Quality&lt;/strong> is where defensive engineering lives. &lt;code>test-data&lt;/code> produces a test plan you can actually defend in a code review — uniqueness, referential integrity, nulls, accepted values, freshness, volume variance — and maps each test to a severity level. &lt;code>performance-tuning&lt;/code> follows the measure-first philosophy: find the critical path, then fix partition pruning, incremental processing, and the phantom dependencies that are silently doubling your costs. &lt;code>spark-performance&lt;/code> handles the distributed case — partition counts, shuffle minimisation, skew handling, broadcasting small tables.&lt;/p>
&lt;p>&lt;strong>Operations&lt;/strong> is what should happen when things go pop. &lt;code>incident-comms&lt;/code> gives the agent the severity classification, notification workflow, update cadence, and post-incident review template — the &amp;ldquo;Don&amp;rsquo;t Go Dark&amp;rdquo; pattern in skill form. &lt;code>pipeline-design&lt;/code> encodes idempotency, reproducibility, and defensive engineering as defaults instead of afterthoughts. &lt;code>data-as-a-product&lt;/code> brings data mesh thinking into the agent — domain ownership, discoverability, SLAs, federated governance. &lt;code>data-security-classification&lt;/code> covers the four A&amp;rsquo;s, PII/PHI handling, and least privilege.&lt;/p>
&lt;p>There&amp;rsquo;s a setup skill (&lt;code>setup-ghost-skills&lt;/code>) that runs once per repo and handles the configuration. Everything else, you opt into.&lt;/p>
&lt;br>
&lt;hr>
&lt;h3 id="why-tool-agnostic-even-though-tool-specific-would-sell-better">Why tool-agnostic, even though tool-specific would sell better&lt;/h3>
&lt;br>
&lt;p>The temptation when building something like this is to specialise. Snowflake-only. dbt-only. Make every skill assume your exact stack and produce immediately runnable code.&lt;/p>
&lt;p>I deliberately didn&amp;rsquo;t do that, and the tradeoff is worth naming honestly. Generic skills are less immediately useful than highly specific ones. A skill that knows your exact dbt project structure, your column naming conventions, your team&amp;rsquo;s snake_case-vs-camelCase argument that&amp;rsquo;s been going on since 2022 — that&amp;rsquo;s more powerful, day one.&lt;/p>
&lt;p>But it&amp;rsquo;s also yours. It&amp;rsquo;s not shareable. And the bits that are genuinely universal — &lt;em>declare the grain before you build the fact&lt;/em>, &lt;em>a business key isn&amp;rsquo;t a real business key until you&amp;rsquo;ve checked it&amp;rsquo;s unique&lt;/em>, &lt;em>test before publish&lt;/em> — those don&amp;rsquo;t change between platforms. They didn&amp;rsquo;t change when we moved from Teradata to Snowflake. They won&amp;rsquo;t change when we move from Snowflake to whatever everyone&amp;rsquo;s furiously rebranding next year.&lt;/p>
&lt;p>Ghost Skills is the base layer. Your repo&amp;rsquo;s &lt;code>CONTEXT.md&lt;/code> is your specialisation layer. They&amp;rsquo;re meant to work together, not replace each other. Fork it. Add a &lt;code>snowflake-cost-optimisation&lt;/code> skill on top. Open a PR if you want to share it back. (Or don&amp;rsquo;t. Do what you like.)&lt;/p>
&lt;br>
&lt;hr>
&lt;h3 id="skills-help-they-dont-solve-everything-lets-not-pretend">Skills help. They don&amp;rsquo;t solve everything. Let&amp;rsquo;s not pretend.&lt;/h3>
&lt;br>
&lt;p>The honest thing to say at this point is that skills are not a silver bullet, and anyone telling you otherwise is selling you a course.&lt;/p>
&lt;p>Skills tell an agent &lt;em>how&lt;/em> to do things. They don&amp;rsquo;t tell it &lt;em>what&amp;rsquo;s there&lt;/em> or &lt;em>why&lt;/em>. An agent that knows the Kimball four-step process will still make the wrong grain declaration if it doesn&amp;rsquo;t understand the business process it&amp;rsquo;s modelling. A skill can encode what SCD2 is; it can&amp;rsquo;t replace a conversation with the stakeholder who actually needs the history. A skill can encode the Write-Audit-Publish pattern; it can&amp;rsquo;t tell the agent which of your fourteen &amp;ldquo;users&amp;rdquo; tables is the one to run it against.&lt;/p>
&lt;p>The other thing skills can&amp;rsquo;t do is fix a stale skill. The conventions you encode today are the conventions you knew today. If your team&amp;rsquo;s standards evolve — and they should — your skills need to evolve with them. A stale skill is worse than no skill, because it gives the agent a confident wrong answer instead of asking a question.&lt;/p>
&lt;p>So treat them as a starting point. Use them. Fork them. Argue with them in a code review. Update them when your standards shift. The agent will keep reading whatever&amp;rsquo;s in the repo, faithfully, every session — which means the skills are only ever as good as the last person who maintained them.&lt;/p>
&lt;p>The repo is at &lt;a href="https://github.com/ghostinthedata-info/skills" target="_blank" rel="noopener">github.com/ghostinthedata-info/skills&lt;/a>. Fork it for your cloud-specific layer. Open issues if something&amp;rsquo;s wrong. Open a PR if you&amp;rsquo;ve got methodology worth encoding and you don&amp;rsquo;t want to keep it to yourself.&lt;/p>
&lt;p>And if you build something better on top of it, please tell me.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>AI Tools</category><category>Tools</category><category>AI Agents</category><category>Claude Code</category><category>dbt</category><category>Data Modeling</category><category>Dimensional Modeling</category><category>Data Quality</category><category>Open Source</category><category>Snowflake</category></item><item><title>The Competitive Moat That AI Can't Replicate</title><link>https://ghostinthedata.info/posts/2026/2026-06-13-human-connection-moat/</link><pubDate>Sat, 13 Jun 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-06-13-human-connection-moat/</guid><author>Chris Hillman</author><description>In a world racing to automate every interaction, the organisations that invest in genuine human connection are quietly building something no algorithm can copy.</description><content:encoded>&lt;h2 id="the-restaurant-that-refused-to-take-bookings-online">The Restaurant That Refused to Take Bookings Online&lt;/h2>
&lt;br>
&lt;p>Let me tell you a story about a restaurant owner who became obsessed with human connection.&lt;/p>
&lt;p>He didn&amp;rsquo;t want people booking online. He wanted them to call. He wanted the ritual of a human voice, the small exchange about an anniversary or a first date, the warmth of being recognised. His team thought he was losing his mind. Online bookings were standard. Everyone did it. Why make customers work harder?&lt;/p>
&lt;p>One of the managers finally asked him a pointed question: &lt;em>Have you ever actually made a reservation with us?&lt;/em>&lt;/p>
&lt;p>He hadn&amp;rsquo;t. So he did.&lt;/p>
&lt;p>He called the restaurant. He was put on hold for thirty minutes. When someone finally answered, they were apologetic but firm — the restaurant was fully booked. No warmth. No conversation. Just a long wait and a closed door.&lt;/p>
&lt;p>In trying to humanise the process, he&amp;rsquo;d made it worse.&lt;/p>
&lt;p>Here&amp;rsquo;s where most stories would end with &amp;ldquo;so they moved everything online and fired the reservation staff.&amp;rdquo; That&amp;rsquo;s not what happened. They did move to online booking, but they kept the entire reservation team and repurposed them. These people now spent their days learning about the customers coming in that night. Who was celebrating a birthday? Who was on a first date? What had a regular not finished on their plate six months ago?&lt;/p>
&lt;p>The team became mini-concierges. Every guest walked in to find someone who knew them — not in a creepy, surveillance-state way, but in the way a good friend remembers what you&amp;rsquo;re going through. The technology handled the transactional layer so the humans could focus on the relational layer.&lt;/p>
&lt;p>I&amp;rsquo;m telling you this story because it captures something I think most organisations have forgotten in the rush to automate everything: &lt;strong>the work of connection isn&amp;rsquo;t overhead. It&amp;rsquo;s the point.&lt;/strong> And as AI swallows the transactional layer of every industry at once, that relational layer is about to become the only point of difference left.&lt;/p>
&lt;br>
&lt;br>
&lt;hr>
&lt;h2 id="trust-is-built-one-marble-at-a-time">Trust Is Built One Marble at a Time&lt;/h2>
&lt;br>
&lt;p>Brené Brown tells a story about her daughter Ellen that I realate to. Ellen came home from third grade devastated. She&amp;rsquo;d shared something vulnerable with a friend at recess, and by the time class started again, half the school knew. She told her mum, through tears, that she would never trust anyone again.&lt;/p>
&lt;p>Brown&amp;rsquo;s first instinct, honestly, was to agree. Tell no one anything, ever. But she caught herself and offered a different idea — the marble jar. There&amp;rsquo;s a jar in Ellen&amp;rsquo;s classroom where good choices add marbles and bad ones remove them. When the jar fills up, the class celebrates. Trust, Brown told her, works the same way. You share the hard things with people who have filled up your marble jar over time. Marble by marble. Small moment by small moment.&lt;/p>
&lt;p>Ellen immediately named two friends whose jars were full.&lt;/p>
&lt;p>What stopped Brown cold was &lt;em>what counted as marbles&lt;/em>. One friend had scooted over at the lunch table to make &amp;ldquo;half a heinie seat&amp;rdquo; when Ellen had nowhere to sit. Another had remembered the names of Ellen&amp;rsquo;s grandparents at a soccer game. Brown couldn&amp;rsquo;t believe it. Surely trust required grander gestures than this.&lt;/p>
&lt;p>She went back to her research data and found exactly the opposite. The single highest-ranked trust-building behaviour in her entire dataset? &lt;strong>People who attended funerals.&lt;/strong> Not the people who made big speeches or grand gestures. The people who showed up in small moments that mattered.&lt;/p>
&lt;p>Trust isn&amp;rsquo;t a transaction. You can&amp;rsquo;t purchase it, automate it, or accelerate it with a clever marketing campaign. It accumulates in tiny, unremarkable moments that look like nothing from the outside — and everything to the person on the receiving end.&lt;/p>
&lt;p>And here&amp;rsquo;s what I find most interesting about the marble jar: once it&amp;rsquo;s full, it&amp;rsquo;s incredibly hard to empty. You can weather bad days, miscommunications, even genuine mistakes. But if you&amp;rsquo;ve never bothered to fill it, a single bad interaction is all it takes. There&amp;rsquo;s nothing to cushion the fall.&lt;/p>
&lt;p>Every organisation has customers whose jars are full and customers whose jars are empty. The ones with full jars forgive outages, laugh off a late delivery, stay through a price increase. The ones with empty jars churn the first time anything goes wrong. Most organisations don&amp;rsquo;t realise which is which until it&amp;rsquo;s too late.&lt;/p>
&lt;br>
&lt;br>
&lt;hr>
&lt;h2 id="hospitality-is-a-dialogue-service-is-a-monologue">Hospitality Is a Dialogue, Service Is a Monologue&lt;/h2>
&lt;br>
&lt;p>Danny Meyer, the restaurateur behind Union Square Hospitality Group and Shake Shack, draws a great distinction. Service is the technical delivery of a product — the food arrives hot, the bill is accurate, the room is clean. &lt;strong>Hospitality is how the delivery of that product makes its recipient feel.&lt;/strong>&lt;/p>
&lt;p>Service is a monologue. You decide your standards and you deliver against them. Hospitality is a dialogue. You notice, you adjust, you respond.&lt;/p>
&lt;p>Meyer puts it even more sharply: hospitality exists when you believe the other person is on your side. It&amp;rsquo;s present when something happens &lt;em>for&lt;/em> you. It&amp;rsquo;s absent when something happens &lt;em>to&lt;/em> you. Those two little prepositions contain the whole difference. You feel it instantly when it&amp;rsquo;s there. You feel it even more instantly when it&amp;rsquo;s not.&lt;/p>
&lt;p>Meyer isn&amp;rsquo;t anti-technology, by the way. But he draws a firm line: the technology he&amp;rsquo;s interested in is technology that &lt;em>reinforces&lt;/em> hospitality, not technology that replaces it. He gives sommeliers Apple Watches for real-time information. He doesn&amp;rsquo;t give them to waiters, because waiters need to maintain eye contact with guests.&lt;/p>
&lt;p>Thats super important. A restaurateur whose restaurants have earned basically every award that exists decided his waiters should not be looking at screens because &lt;strong>eye contact is more important than information&lt;/strong>. When was the last time you saw a decision like that made inside a large organisation?&lt;/p>
&lt;p>The best story I know in this vein is from Will Guidara, who ran Eleven Madison Park when it was named the best restaurant in the world. One night, Guidara overheard a table of European food tourists lamenting that they&amp;rsquo;d eaten at Per Se, Le Bernardin, and Daniel — but they were flying home the next day and had never tried a proper New York City hot dog.&lt;/p>
&lt;p>Guidara ran outside, bought a $2 hot dog from a street cart, convinced his chef to plate it with &amp;ldquo;swooshes of ketchup and a quenelle of relish,&amp;rdquo; and had it served as a surprise mid-course. Every person at the table said it was the highlight of not just the meal, but their entire trip to New York.&lt;/p>
&lt;p>A $2 hot dog in a room full of the most expensive food in America. That&amp;rsquo;s the gap between service and hospitality. You cannot design an algorithm that eavesdrops on dinner conversation and dispatches someone to buy a street hot dog, because the person on the receiving end would immediately sense the machinery of it. The magic is precisely that a human heard, a human decided, a human cared.&lt;/p>
&lt;p>Guidara was so moved by what that moment revealed that he created a dedicated role called the &lt;strong>Dreamweaver&lt;/strong> — someone whose entire job was to make these moments happen. Researching guests ahead of visits, building cheat sheets so staff could greet people by name, listening for conversational cues during dinner. This is exactly what the restaurant in the opening did. The technology handles the boring parts so the humans can do the thing only humans can do.&lt;/p>
&lt;br>
&lt;br>
&lt;hr>
&lt;h2 id="the-things-we-destroy-because-we-cant-measure-them">The Things We Destroy Because We Can&amp;rsquo;t Measure Them&lt;/h2>
&lt;br>
&lt;p>Here&amp;rsquo;s where I get frustrated, and this is the part of the article I&amp;rsquo;ve rewritten three times.&lt;/p>
&lt;p>I&amp;rsquo;ve spent a lot of time in and around banking over the past few years, and I&amp;rsquo;ve watched a particular kind of decision get made over and over. It always starts the same way. Someone pulls a report showing branch transaction volumes declining. Someone else pulls a report showing the cost per square metre of the branch network. A third person models what happens if you close the bottom quartile. The numbers are clean. The deck practically writes itself. And the decision gets made in a room where nobody has ever stood behind the counter and watched what actually happens on a Tuesday morning at half past ten.&lt;/p>
&lt;p>What you see when you actually stand there is not transactions. You see a retired schoolteacher who comes in every week not because she can&amp;rsquo;t use the app — she can — but because the teller knows her by name and asks about her grandson. You see a small business owner dropping off the weekly takings and mentioning, almost offhandedly, that he&amp;rsquo;s thinking about expanding into the next suburb, and the branch manager quietly filing that away because she knows a commercial lender who should probably call him. You see the elderly man whose wife just died, standing at the counter trying to change the joint account, and the staff member who recognises what&amp;rsquo;s actually happening and gently walks him through it for forty-five minutes without once looking at a clock.&lt;/p>
&lt;p>None of that shows up in the transaction volume report. None of it shows up in the cost-per-square-metre model. But it&amp;rsquo;s where the loyalty comes from. It&amp;rsquo;s the reason the schoolteacher&amp;rsquo;s kids bank with the same bank and her business-owner neighbour refinanced there and the grieving husband never even considered moving his accounts when the rate went up.&lt;/p>
&lt;p>When the branch closes, all of that evaporates. Not dramatically. Just quietly. The schoolteacher switches to the credit union down the road because they still have tellers. The business owner realises nobody at the new centralised lending hub remembers him, and starts taking meetings with competitors. The grieving husband&amp;rsquo;s children, watching their dad struggle with the phone banking menu, consolidate everything with a different bank when they inherit. Three years later the spreadsheet shows a modest increase in customer churn and nobody connects it to the branch closure, because the causal chain is too long and the data was never collected in the first place.&lt;/p>
&lt;p>The broader pattern is consistent everywhere I look. Australian banks have closed something like 2,500 branches since 2017. The justification was always the same: digital adoption, customer preference, efficiency. But when whistleblowers from the Finance Sector Union testified at the Senate inquiry, a darker story emerged. Branch staff had been performance-managed to &lt;em>suppress&lt;/em> in-branch activity — pushed to redirect customers to ATMs, pressured to sign them up for digital banking. The data that justified the closures had been partly manufactured by the closures themselves.&lt;/p>
&lt;p>And what the banks never properly measured was any of this. Every one of those moments is share of wallet. Every one of them is retention. And none of it appeared on the spreadsheet that closed the branch.&lt;/p>
&lt;p>The tragic irony is what came next. The same banks that couldn&amp;rsquo;t justify keeping branches open are now spending billions on AI personalisation engines designed to replicate the exact relationships they dismantled. JPMorgan Chase budgeted $18 billion for technology spending in 2025. Bank of America&amp;rsquo;s AI assistant has surpassed two billion client interactions. And yet only about a quarter of consumers say their bank provides tailored financial advice. They destroyed the organic version and are now paying orders of magnitude more trying to rebuild a synthetic one.&lt;/p>
&lt;p>There&amp;rsquo;s a concept I keep coming back to called the &lt;strong>McNamara Fallacy&lt;/strong>, named after Robert McNamara, the US Defense Secretary who tried to run the Vietnam War like a business. McNamara was brilliant with metrics. A general once told him he needed to add an &amp;ldquo;x-factor&amp;rdquo; to his measurement list — the feelings of the rural Vietnamese people. McNamara wrote it down, asked what it meant, and then sarcastically told the general he couldn&amp;rsquo;t measure it, so he erased it.&lt;/p>
&lt;p>Daniel Yankelovich described the fallacy as a four-step descent. First, you measure whatever can be easily measured. Second, you disregard what can&amp;rsquo;t be easily measured. Third, you presume that what can&amp;rsquo;t be measured easily isn&amp;rsquo;t important. And finally — this is the fatal step — &lt;strong>you presume that what can&amp;rsquo;t be easily measured doesn&amp;rsquo;t exist.&lt;/strong> Yankelovich called this last step &amp;ldquo;suicide.&amp;rdquo;&lt;/p>
&lt;p>Every organisation I&amp;rsquo;ve worked with has done this at some level. Dashboards capture the measurable. But the most valuable things — trust, loyalty, the feeling a customer gets when someone remembers their name, the way an employee feels when their manager notices they&amp;rsquo;re having a hard week — live in the unmeasured space between the data points. And when we optimise away the things we can&amp;rsquo;t count, we destroy the foundation that made the things we &lt;em>can&lt;/em> count actually work.&lt;/p>
&lt;p>If you work in data, as most people reading this do. The reports that justified those branch closures were technically accurate. Well-modelled, properly sourced, beautifully visualised to requirements. Someone like us built every one of them.&lt;/p>
&lt;p>Wells Fargo is the textbook case. Leadership set a target of eight accounts per customer. &amp;ldquo;Eight is Great&amp;rdquo; became the mantra. Over fourteen years, employees created &lt;strong>3.5 million fraudulent accounts&lt;/strong> to hit the number. They optimised the metric so aggressively they destroyed the thing — trust — that made banking relationships worth anything in the first place. The fake accounts generated about two million dollars in fees. The fallout cost them more than three billion.&lt;/p>
&lt;p>Charles Goodhart&amp;rsquo;s law captures this beautifully: &lt;strong>&amp;ldquo;When a measure becomes a target, it ceases to be a good measure.&amp;rdquo;&lt;/strong> The moment you start optimising for NPS, NPS stops telling you anything true about your customer relationships. The moment you start optimising for branch foot traffic, foot traffic stops telling you anything true about community banking. You end up managing the shadow and losing the substance.&lt;/p>
&lt;br>
&lt;br>
&lt;hr>
&lt;h2 id="why-the-apple-store-is-always-full-and-the-carrier-store-is-always-empty">Why the Apple Store Is Always Full and the Carrier Store Is Always Empty&lt;/h2>
&lt;br>
&lt;p>Walk into any shopping centre and do a little experiment. Find the Apple Store. It&amp;rsquo;ll be packed. People learning photography, kids in coding workshops, someone at the Genius Bar getting help with an old laptop, others just hanging out because the store feels like a place you want to be. Then walk to the Telstra store, or Optus, or whichever carrier is nearest. You&amp;rsquo;ll find two or three bored staff members, maybe one customer, and the unmistakable atmosphere of a waiting room at a tax office.&lt;/p>
&lt;p>They sell, to a remarkable degree, the same products. iPhones, accessories, plans. So why does one feel like a community space and the other feel like a place you&amp;rsquo;re only in because you have to be?&lt;/p>
&lt;p>The answer starts with Ron Johnson, who built Apple retail alongside Steve Jobs in 2000. Johnson made a decision most retailers would never make: a store needed to be much more than a place to acquire merchandise. It needed to help people enrich their lives. Any website could transact. A store had to do something a website couldn&amp;rsquo;t.&lt;/p>
&lt;p>He and Jobs asked themselves a question I still think about: &lt;em>what would the Four Seasons do?&lt;/em> The Four Seasons doesn&amp;rsquo;t have cashiers. It has a concierge. So Apple introduced concierge greeters. The Four Seasons has a bar — a place where you can sit, ask questions, and get real help from a knowledgeable person. Apple created the Genius Bar. It dispensed advice instead of alcohol, but the idea was the same: a place to be cared for, not a place to be processed.&lt;/p>
&lt;p>Then they did something that, from a commercial perspective, looks insane. &lt;strong>They removed commissions.&lt;/strong> Apple Store employees earn zero percent on sales. No upsell quotas, no commission targets, no pressure scripts. Johnson&amp;rsquo;s framing was that Apple wanted to reach your heart instead of your wallet — and staff whose paycheque depends on closing a sale cannot be fully on the customer&amp;rsquo;s side.&lt;/p>
&lt;p>Compare this to how a carrier store operates. Employees have sales targets. They&amp;rsquo;re incentivised to push you toward plans and devices that pay them more. Former carrier staff in Australia and the US tell the same story: focus on customer service and you get coached on pushing harder instead. The misalignment between what the employee needs and what the customer needs sits in the air. Customers feel it as distrust, even if they can&amp;rsquo;t name it.&lt;/p>
&lt;p>Apple Stores generate roughly $5,500 in revenue per square foot — about twice the next-highest retailer on earth. They get over a million customer visits a day. But only about &lt;strong>one in a hundred visitors actually buys something.&lt;/strong> Ninety-nine percent walk in, spend time, and walk out without making a purchase. And that&amp;rsquo;s the whole point. Apple built the most profitable retail operation in history by investing in the ninety-nine people who aren&amp;rsquo;t buying today, because every visit is a marble in the jar.&lt;/p>
&lt;p>The carrier stores are trying to convert everyone who walks in. They fail. Apple is trying to make everyone who walks in feel good. They succeed, and they also happen to sell more per square foot than any retailer in history. The causation is not subtle.&lt;/p>
&lt;br>
&lt;br>
&lt;hr>
&lt;h2 id="ai-raises-the-floor-humans-raise-the-ceiling">AI Raises the Floor. Humans Raise the Ceiling.&lt;/h2>
&lt;br>
&lt;p>This isn&amp;rsquo;t an anti-AI article. I work in data. I use AI every day. I think it&amp;rsquo;s genuinely transformative and most organisations should be using it more aggressively than they currently are.&lt;/p>
&lt;p>But the conversation about AI has got the direction wrong. The question isn&amp;rsquo;t whether AI will replace human connection. The question is what AI should free us &lt;em>up to do&lt;/em>. If the answer is &amp;ldquo;more automation, more optimisation, more throughput&amp;rdquo; — we&amp;rsquo;ve missed the plot entirely.&lt;/p>
&lt;p>There&amp;rsquo;s a concept emerging in the marketing and design literature that I find genuinely useful: &lt;strong>AI raises the floor, but it lowers the ceiling.&lt;/strong> A Fortune 500 study of customer-support agents found that AI boosted the average agent&amp;rsquo;s productivity by 14%, but boosted the least-skilled agents by 34%. That&amp;rsquo;s the floor rising. But when researchers tested AI on creative writing, they found something unsettling. Individual AI-assisted stories were rated higher than purely human stories. Read a whole collection of them together, though, and they were rated &lt;em>lower&lt;/em>. They all felt the same. The ceiling dropped because the variance dropped.&lt;/p>
&lt;p>Here&amp;rsquo;s what that means for organisations competing on anything other than price. If every company uses the same AI tools to optimise the same metrics, the optimisation itself becomes commodity. Everyone has the same chatbot, the same personalisation engine, the same churn predictor. The floor rises for everyone simultaneously, which means relative advantage disappears. The only remaining differentiator — the actual moat — becomes the irreducibly human: genuine empathy in a crisis, a person who actually cares, a brand that feels run by real people who listen.&lt;/p>
&lt;p>John Naisbitt, the futurist who coined &lt;strong>&amp;ldquo;high-tech, high-touch,&amp;rdquo;&lt;/strong> made an observation back in 1999 that feels almost prophetic now: the two biggest markets in America are consumer technology and &lt;em>escape from consumer technology&lt;/em>. We&amp;rsquo;re in the same dynamic today, just at a higher frequency. The more AI we get, the more valuable the unmediated human encounter becomes.&lt;/p>
&lt;p>I&amp;rsquo;m seeing this play out in hospitality right now. Mid-range hotels are racing toward automation — kiosk check-in, chatbot concierges, app-based everything. But &lt;strong>luxury hotels are moving in the opposite direction.&lt;/strong> They&amp;rsquo;re bringing back butlers, personalised greetings, tailored in-person experiences. In an automated world, human service becomes the luxury.&lt;/p>
&lt;p>Simon Sinek put it well: AI is a tool, not a replacement. The struggles and imperfections in human interaction are often where the most meaningful growth and strongest bonds happen. Those imperfections are not bugs to be optimised away. They&amp;rsquo;re the texture of real connection.&lt;/p>
&lt;br>
&lt;br>
&lt;hr>
&lt;h2 id="your-employees-are-the-moat-the-compounding-is-invisible">Your Employees Are the Moat. The Compounding Is Invisible.&lt;/h2>
&lt;br>
&lt;p>There&amp;rsquo;s a piece of this conversation that most leaders miss, and it&amp;rsquo;s the part I care about most. You cannot deliver genuine human connection to customers if you haven&amp;rsquo;t first delivered it to your employees. The service-profit chain — a Harvard Business School framework that&amp;rsquo;s been around for decades — is unambiguous: internal service quality drives employee satisfaction, which drives retention, which drives external service value, which drives customer loyalty, which drives revenue. It&amp;rsquo;s a chain. You cannot cheat the order.&lt;/p>
&lt;p>Herb Kelleher ran Southwest Airlines on this principle for decades. Employees come first; treat them right and they treat your customers right, and the customers come back. The things you can&amp;rsquo;t buy, he said, are dedication, devotion, loyalty — the feeling of participating in a crusade. Competitors can buy the physical stuff. They cannot buy the feeling. The results: number one in US Department of Transportation customer service rankings twenty-six times in thirty-four years, and forty-seven consecutive years of profitability in an industry famous for burning money.&lt;/p>
&lt;p>His successor as president, Colleen Barrett, started at Southwest as his legal secretary. She said he treated her as an equal from the day they met, and he treated every employee the same way. That&amp;rsquo;s not a culture program. That&amp;rsquo;s a human being who decided to care about other human beings and then built a company around the decision.&lt;/p>
&lt;p>This matters for the AI conversation because there&amp;rsquo;s a temptation, especially under cost pressure, to look at AI and think &amp;ldquo;great, we can run leaner.&amp;rdquo; That&amp;rsquo;s the trap. Use AI to squeeze more out of fewer people and you&amp;rsquo;ll hit short-term margins while quietly draining the moat. Use it to free people up to do the things only people can do — notice, listen, care, remember — and you&amp;rsquo;ll build something that compounds for years.&lt;/p>
&lt;p>And compounding is exactly the right word: &lt;strong>trust compounds, but the compounding is invisible for a long time.&lt;/strong> James Clear&amp;rsquo;s ice cube analogy applies. Heat a frozen room from twenty-five degrees to thirty-one (Fahrenheit — stay with me) and nothing appears to happen. Nothing, nothing, nothing — then at thirty-two degrees, everything happens at once. All the action is at the threshold, and all of it is invisible until it isn&amp;rsquo;t. Most organisations quit before the threshold arrives, because the accounting cycle is quarterly and the compounding cycle is years.&lt;/p>
&lt;p>Seth Godin says: &lt;em>&amp;ldquo;You must build trust before you need it. Building trust right when you want to make a sale is just too late.&amp;rdquo;&lt;/em> You get there drip by drip by drip, until people would miss you if you were gone.&lt;/p>
&lt;br>
&lt;br>
&lt;hr>
&lt;h2 id="so-what-do-you-actually-do-with-this">So What Do You Actually Do With This?&lt;/h2>
&lt;br>
&lt;p>&lt;strong>Ask yourself which of your customer-facing moments are dialogue and which are monologue.&lt;/strong> If every interaction is something you&amp;rsquo;ve decided to deliver &lt;em>to&lt;/em> the customer rather than &lt;em>for&lt;/em> them, you don&amp;rsquo;t have hospitality. You have service. Service is fine. Service is table stakes. But it&amp;rsquo;s not a moat.&lt;/p>
&lt;p>&lt;strong>Ask yourself what your employees would say if someone asked them whether they were cared for.&lt;/strong> Not in a survey. In an honest conversation over coffee. That answer is the ceiling on what your customers will ever feel, no matter how much you spend on customer experience.&lt;/p>
&lt;p>&lt;strong>Ask yourself which of your metrics you&amp;rsquo;re managing and which ones you&amp;rsquo;re serving.&lt;/strong> If your teams are optimising for NPS, you&amp;rsquo;ve lost NPS as a signal. The moment a metric becomes a target, it stops telling you the truth. You need to look at the behaviour underneath the number — the actual quality of the human moments — and measure those, even if the measurement is messy.&lt;/p>
&lt;p>&lt;strong>Ask yourself where you&amp;rsquo;re using technology to replace human moments rather than amplify them.&lt;/strong> The restaurant in the opening didn&amp;rsquo;t remove the reservation staff. They redeployed them to do the thing only humans can do. That&amp;rsquo;s the move. Automate the transaction. Invest the savings in the relationship.&lt;/p>
&lt;p>Remember that the dashboards capture the measurable but the valuable things live in the space between the measurements. The best data work I&amp;rsquo;ve ever seen hasn&amp;rsquo;t been about building better dashboards. It&amp;rsquo;s been about giving humans better context so they can have better conversations with other humans. That&amp;rsquo;s the quiet revolution most data teams are missing. We keep building tools to replace the human layer when we should be building tools to enrich it.&lt;/p>
&lt;p>The moat isn&amp;rsquo;t your data pipeline. It isn&amp;rsquo;t your AI model. It isn&amp;rsquo;t your CRM. The moat is what your people do with the insights — the handshake, the eye contact, the moment of genuine care that no algorithm can fake and no competitor can copy. You can&amp;rsquo;t buy it and you can&amp;rsquo;t rush it. You can only show up, drip by drip, marble by marble, for longer than everyone else is willing to.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Leadership</category><category>Personal Development</category><category>Career Development</category><category>Human Connection</category><category>Trust</category><category>Customer Experience</category><category>Leadership</category><category>AI</category><category>Employee Experience</category><category>Organisational Culture</category></item><item><title>SQL Tells You What. Comments Tell You Why.</title><link>https://ghostinthedata.info/posts/2026/2026-06-06-code-tells-what-comments-why/</link><pubDate>Sat, 06 Jun 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-06-06-code-tells-what-comments-why/</guid><author>Chris Hillman</author><description>SQL is a declarative language — it tells you what the query does, never why. Here's why that distinction matters more in data engineering than anywhere else.</description><content:encoded>&lt;p>The best SQL doesn&amp;rsquo;t need comments. Write meaningful CTE names, descriptive aliases, clear column labels — and a skilled reader will follow your logic without a single annotation. That&amp;rsquo;s the right instinct.&lt;/p>
&lt;p>It&amp;rsquo;s also only half right.&lt;/p>
&lt;p>SQL is a declarative language. You&amp;rsquo;re not writing &lt;em>how&lt;/em> the database retrieves your data; you&amp;rsquo;re writing &lt;em>what&lt;/em> you want. That&amp;rsquo;s a useful distinction, because &amp;ldquo;what&amp;rdquo; and &amp;ldquo;why&amp;rdquo; are very different questions, and SQL can answer exactly one of them.&lt;/p>
&lt;p>A query can be perfectly, elegantly readable and still be completely opaque about the reason it exists. The name of your CTE, no matter how well chosen, cannot explain the business decision that gave birth to it.&lt;/p>
&lt;hr>
&lt;h3 id="the-self-documenting-ceiling">The self-documenting ceiling&lt;/h3>
&lt;p>Consider this query:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(order_total) &lt;span style="color:#66d9ef">AS&lt;/span> lifetime_value,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> DATEDIFF(&lt;span style="color:#66d9ef">day&lt;/span>, &lt;span style="color:#66d9ef">MIN&lt;/span>(order_date), &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> customer_age_days
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> orders
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_date &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#66d9ef">day&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">90&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> status_code &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">IN&lt;/span> (&lt;span style="color:#ae81ff">1&lt;/span>, &lt;span style="color:#ae81ff">2&lt;/span>, &lt;span style="color:#ae81ff">9&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_id
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Clean SQL. No ambiguity about &lt;em>what&lt;/em> it does. But try answering these questions from the code alone:&lt;/p>
&lt;p>Why 90 days? Why not 60, or 180, or the full history? Is 90 the standard customer lifetime window, a finance reporting period, the SLA in a merchant agreement, or the number someone picked in a meeting four years ago and nobody has questioned since?&lt;/p>
&lt;p>What are status codes 1, 2, and 9? You might guess from context — pending, draft, cancelled — but you&amp;rsquo;re guessing. More importantly: why are they excluded? Is this a business rule about what counts as &amp;ldquo;real&amp;rdquo; revenue, or a data quality workaround because those status codes appear when an upstream webhook fires twice?&lt;/p>
&lt;p>These are not trivial questions. The 90-day cutoff defines what &amp;ldquo;active customer&amp;rdquo; means across your entire reporting layer. The excluded status codes determine what gets counted as revenue. Change either without understanding the original intent, and you&amp;rsquo;re not refactoring — you&amp;rsquo;re silently redefining business logic that stakeholders are relying on.&lt;/p>
&lt;p>Better SQL helps, but only so far:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(order_total) &lt;span style="color:#66d9ef">AS&lt;/span> lifetime_value,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> DATEDIFF(&lt;span style="color:#66d9ef">day&lt;/span>, &lt;span style="color:#66d9ef">MIN&lt;/span>(order_date), &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> customer_age_days
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> orders
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_date &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#66d9ef">day&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">90&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> status_code &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">IN&lt;/span> (
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#ae81ff">1&lt;/span>, &lt;span style="color:#75715e">-- pending
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> &lt;span style="color:#ae81ff">2&lt;/span>, &lt;span style="color:#75715e">-- draft
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> &lt;span style="color:#ae81ff">9&lt;/span> &lt;span style="color:#75715e">-- cancelled
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> )
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_id
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>This is better. You&amp;rsquo;ve replaced the mystery with labels. But inline labels aren&amp;rsquo;t explanations — they&amp;rsquo;re just naming the what more explicitly. You still don&amp;rsquo;t know why.&lt;/p>
&lt;hr>
&lt;h3 id="what-sql-literally-cannot-tell-you">What SQL literally cannot tell you&lt;/h3>
&lt;p>SQL cannot explain why the program was written the way it was. It cannot discuss the reasons certain alternative approaches were taken. It cannot tell you that the business definition changed, that the upstream system is broken, or that this logic exists specifically to handle a problem that was supposed to be fixed in Q2 and wasn&amp;rsquo;t.&lt;/p>
&lt;p>Here&amp;rsquo;s the kind of context that belongs in a comment:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Customer lifetime window: 90-day lookback aligns with the SLA in the Merchant Agreement (v3.2).
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Reviewed and confirmed with Finance, Dec 2024. Don&amp;#39;t change without checking with that team first.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- status_codes 1, 2, 9 excluded per the order lifecycle defined in the legacy Salesforce migration.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- These codes appear as artefacts when the Salesforce sync fires before the order is confirmed.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- The upstream fix was deprioritised (JIRA: DATA-3841). Removing this filter will cause double-counting.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(order_total) &lt;span style="color:#66d9ef">AS&lt;/span> lifetime_value,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> DATEDIFF(&lt;span style="color:#66d9ef">day&lt;/span>, &lt;span style="color:#66d9ef">MIN&lt;/span>(order_date), &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> customer_age_days
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> orders
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_date &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#66d9ef">day&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">90&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> status_code &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">IN&lt;/span> (&lt;span style="color:#ae81ff">1&lt;/span>, &lt;span style="color:#ae81ff">2&lt;/span>, &lt;span style="color:#ae81ff">9&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_id
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>That comment block is five lines. It contains: the business rationale, the stakeholder who owns the decision, the date it was confirmed, the upstream system causing the problem, the known issue reference, and the consequence of removing the filter. None of that information is recoverable from the SQL. All of it is load-bearing.&lt;/p>
&lt;hr>
&lt;h3 id="the-deduplication-that-nobody-remembers-adding">The deduplication that nobody remembers adding&lt;/h3>
&lt;p>Have you seen this line before:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>QUALIFY ROW_NUMBER() OVER (
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> PARTITION &lt;span style="color:#66d9ef">BY&lt;/span> customer_id, order_id
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span> ingested_at &lt;span style="color:#66d9ef">DESC&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>) &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">1&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The SQL is completely transparent about what it does: take the most recently ingested row for each customer/order combination. What it cannot tell you is why duplicates exist in the first place, whether the source of those duplicates has been fixed, whether this deduplication logic is still necessary, or whether &amp;ldquo;most recently ingested&amp;rdquo; is actually the right tiebreaker for your use case.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Deduplicating on customer_id + order_id to handle duplicate webhook events from Eftpos.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Eftpos retries failed webhooks, which can fire the same order_created event multiple times.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Taking the latest ingested_at gives us the most recent event state (status, amount, etc.)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Revisit if we adopt Eftpos&amp;#39;s idempotency keys at the integration layer — this dedup may become redundant.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Background: this logic was added after the October 2024 incident (post-mortem in Confluence).
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>QUALIFY ROW_NUMBER() OVER (
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> PARTITION &lt;span style="color:#66d9ef">BY&lt;/span> customer_id, order_id
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span> ingested_at &lt;span style="color:#66d9ef">DESC&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>) &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">1&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The second version tells the reader everything they need to know to maintain this safely: the upstream cause, the logic behind the tiebreaker, when to reconsider it, and where to find the history. Without that comment, every future engineer who touches this code has to reverse-engineer context that no longer exists anywhere.&lt;/p>
&lt;hr>
&lt;h3 id="when-two-models-define-the-same-word-differently">When two models define the same word differently&lt;/h3>
&lt;p>This one doesn&amp;rsquo;t generate an error. It generates confusion at the worst possible moment.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- IMPORTANT: &amp;#39;churned&amp;#39; here means &amp;gt;90 days since last order.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- This differs from mkt_customer_segments, which uses &amp;gt;60 days for winback campaign targeting.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Both definitions are intentional. Customer Success uses 90 days because cohort analysis
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- showed that 30% of &amp;#34;60-day churned&amp;#34; customers reorder organically within the next 30 days
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- and shouldn&amp;#39;t receive a discount. Marketing uses the stricter threshold to maximise winback opportunities.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- The discrepancy is documented and known. Do not &amp;#34;fix&amp;#34; one to match the other.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">CASE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHEN&lt;/span> days_since_last_order &lt;span style="color:#f92672">&amp;lt;=&lt;/span> &lt;span style="color:#ae81ff">30&lt;/span> &lt;span style="color:#66d9ef">THEN&lt;/span> &lt;span style="color:#e6db74">&amp;#39;active&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHEN&lt;/span> days_since_last_order &lt;span style="color:#f92672">&amp;lt;=&lt;/span> &lt;span style="color:#ae81ff">90&lt;/span> &lt;span style="color:#66d9ef">THEN&lt;/span> &lt;span style="color:#e6db74">&amp;#39;at_risk&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">ELSE&lt;/span> &lt;span style="color:#e6db74">&amp;#39;churned&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">END&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> customer_segment
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Without that comment, someone eventually &amp;ldquo;cleans up&amp;rdquo; one of the two definitions because having different churn thresholds in two models looks like a bug. They&amp;rsquo;re wrong, and they&amp;rsquo;re about to break a winback campaign that&amp;rsquo;s been performing well for six months.&lt;/p>
&lt;p>The SQL cannot warn them. The comment can.&lt;/p>
&lt;hr>
&lt;h3 id="dbt-descriptions-arent-exempt">dbt descriptions aren&amp;rsquo;t exempt&lt;/h3>
&lt;p>dbt gives you a proper home for the &amp;ldquo;why&amp;rdquo;: model descriptions and column-level descriptions in your YAML.&lt;/p>
&lt;p>Bad:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">models&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">fct_orders&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">description&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;The orders fact table. Contains order data.&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">columns&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">customer_segment&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">description&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;The customer segment based on days since last order.&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>That description restates what the model name already says. It is a comment-shaped void.&lt;/p>
&lt;p>Good:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">models&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">fct_orders&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">description&lt;/span>: &amp;gt;&lt;span style="color:#e6db74">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Order-level fact table used as the single source of truth for revenue reporting.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Revenue excludes test accounts (prefixed &amp;#39;INTERNAL_&amp;#39;) and gift card redemptions,
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> per the Finance revenue recognition policy agreed January 2025.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> See Confluence: Revenue Recognition Standards v3.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Do not use this model for marketing attribution — use mkt_attributed_revenue instead,
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> which applies different channel logic.&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">columns&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">customer_segment&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">description&lt;/span>: &amp;gt;&lt;span style="color:#e6db74">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Lifecycle segment defined by the Customer Success team, Q1 2025.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Thresholds (30/90 days) were derived from cohort analysis showing a median
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> reorder window of 28 days. Intentionally differs from the Marketing segment
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> definition in mkt_customer_segments, which uses a 60-day churn threshold
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> for winback campaign purposes. Both are correct for their respective use cases.&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>That description does something a column name never can: it tells you who owns the definition, when it was set, what analysis underpins it, and — critically — where the definition intentionally diverges from a similar field elsewhere in the warehouse. That last part is the thing that saves someone three hours of confused Teams messages.&lt;/p>
&lt;hr>
&lt;h3 id="the-archaeology-problem">The archaeology problem&lt;/h3>
&lt;p>Data teams inherit codebases. It happens constantly, and it will keep happening. Engineers leave. Consultants finish their contracts. Reorgs shuffle ownership. What gets left behind is the SQL, and whatever context wasn&amp;rsquo;t written down is gone.&lt;/p>
&lt;p>Good SQL survives the handover. It&amp;rsquo;s readable, consistent, and correct. But &amp;ldquo;readable&amp;rdquo; in the sense of mechanically parseable is not the same as &amp;ldquo;understandable&amp;rdquo; in the sense of knowing what business problem the code was solving, whose decision it was, and what you&amp;rsquo;d need to change if that decision changed.&lt;/p>
&lt;p>Comments are how you write for the engineer who inherits this in two years. They&amp;rsquo;re how you write for yourself in six months. They&amp;rsquo;re how you ensure that a working pipeline stays working through the context collapse that happens every time anyone leaves a team.&lt;/p>
&lt;p>SQL can only tell you what. Try not to shortchange the people reading it in either direction.&lt;/p></content:encoded><category>Data Engineering</category><category>Data Quality</category><category>SQL</category><category>dbt</category><category>Documentation</category><category>Data Quality</category><category>Code Comments</category><category>Data Pipelines</category><category>Best Practices</category></item><item><title>Don't Go Dark: Visibility Is a Data Engineering Skill</title><link>https://ghostinthedata.info/posts/2026/2026-05-23-dont-go-dark/</link><pubDate>Sat, 23 May 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-05-23-dont-go-dark/</guid><author>Chris Hillman</author><description>Jeff Atwood wrote 'Don't Go Dark' for software engineers in 2008. The advice didn't reach us. Here's what it means for data engineers navigating long migrations, invisible pipelines, and distributed teams.</description><content:encoded>&lt;p>There&amp;rsquo;s a specific kind of silence in data engineering that I&amp;rsquo;ve learned to fear.&lt;/p>
&lt;p>Not the silence of a system that&amp;rsquo;s working well. Not the comfortable quiet of a team in flow. I mean the silence of a project that&amp;rsquo;s been running for three weeks and you still can&amp;rsquo;t point to a single visible thing it has produced. The kind of silence where, if your manager stopped you in the hallway and asked &amp;ldquo;how&amp;rsquo;s that migration going?&amp;rdquo;, you&amp;rsquo;d say &amp;ldquo;fine&amp;rdquo; because saying anything more accurate would require explaining things you haven&amp;rsquo;t fully articulated yet — even to yourself.&lt;/p>
&lt;p>I watched a talented engineer do this once. Let&amp;rsquo;s call her Mia. She&amp;rsquo;d been handed a significant refactor: six months of accumulated technical debt in a set of dbt models that sat at the centre of our revenue reporting. The kind of work that is genuinely unglamorous, invisible by nature, and deeply important. She approached it the way a lot of good engineers approach hard problems — she went quiet and started digging.&lt;/p>
&lt;p>For three weeks, she showed up, she worked hard, and she said almost nothing. Her pull requests were sparse. Her standups were &amp;ldquo;still investigating.&amp;rdquo; Her Teams messages were infrequent. She wasn&amp;rsquo;t slacking off. She was grinding through some of the most complex lineage work I&amp;rsquo;d ever seen someone untangle.&lt;/p>
&lt;p>Then a finance stakeholder messaged me — not her, me. Someone had noticed a spike in our Snowflake credit usage. They were nervous. They&amp;rsquo;d been burned by silent changes before.&lt;/p>
&lt;p>I didn&amp;rsquo;t have an answer. Mia had one, but nobody knew to ask her.&lt;/p>
&lt;p>I went to her and said, &amp;ldquo;walk me through what you&amp;rsquo;ve been doing.&amp;rdquo; What she showed me was extraordinary. She&amp;rsquo;d traced dependency chains through twelve models. She&amp;rsquo;d found three bugs that had been silently wrong for months. She&amp;rsquo;d drafted a migration plan that would cut query time in half. It was genuinely impressive work — and it had been completely invisible for twenty-one days.&lt;/p>
&lt;p>That&amp;rsquo;s the moment I understood something that Jeff Atwood wrote in 2008, in a post called &lt;a href="https://blog.codinghorror.com/dont-go-dark/" target="_blank" rel="noopener">Don&amp;rsquo;t Go Dark&lt;/a>, that I wish someone had made me read before I managed my first data team. Atwood&amp;rsquo;s essay was aimed at software engineers. It quoted a Microsoft engineering rule that said, in essence: &lt;em>three weeks is going dark.&lt;/em> The rule wasn&amp;rsquo;t about laziness. It wasn&amp;rsquo;t about bad engineers. It was about a failure mode so common in technical work that it needed a name.&lt;/p>
&lt;p>Mia wasn&amp;rsquo;t lazy. She was going dark.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="what-atwood-actually-meant">What Atwood Actually Meant&lt;/h3>
&lt;p>Atwood&amp;rsquo;s post is short — shorter than most people expect from the weight it carries. He builds on two sources: a software engineering principle from Jim McCarthy&amp;rsquo;s &lt;em>Dynamics of Software Development&lt;/em>, and an observation from open-source contributor Ben Collins-Sussman. Together they identify a pattern that every engineering lead will recognise the moment they read it.&lt;/p>
&lt;p>The pattern is this: programmers don&amp;rsquo;t want to show work in progress. They want to disappear into the problem, solve it completely, and emerge holding something finished. Collins-Sussman pinned it to fear — not arrogance — arguing that most developers treat coding as an arena for personal heroics rather than an inherently social activity. They&amp;rsquo;ll share code once it&amp;rsquo;s polished. They&amp;rsquo;ll share failures once they&amp;rsquo;re resolved. But the messy in-between, the half-formed thinking, the models with todos and the branches with four commits and no description? That stays private.&lt;/p>
&lt;p>McCarthy&amp;rsquo;s rule was direct: manage your tasks in short enough increments that you always have a visible deliverable at the end of each one. Three weeks without something to show is too long. The early warning of slippage — one day, discovered this week — is worth ten times more than six months of slippage discovered at deadline.&lt;/p>
&lt;p>Atwood&amp;rsquo;s additional observation is the one that has aged best: agile development structures made going dark mechanically difficult. If you&amp;rsquo;re running sprints, you&amp;rsquo;re producing something reviewable every two weeks regardless. The iteration boundary forces surfacing. Joel Spolsky, Atwood&amp;rsquo;s co-founder at Stack Overflow and another writer in this intellectual lineage, made the same argument about daily builds a few years earlier: a team that builds its software every day has, at minimum, a daily signal that things are either working or not. The build is the forcing function. It converts invisible progress into visible evidence.&lt;/p>
&lt;p>Open source projects, Atwood observed, face the highest risk of going dark — because there&amp;rsquo;s no manager, no sprint ceremony, no daily build process unless someone chooses to implement one. The discipline has to be entirely internal. The only currency that matters on an open source project is visible, reviewable work.&lt;/p>
&lt;p>Data engineering, as a discipline, has no such forcing function. Our work doesn&amp;rsquo;t end at a sprint boundary with a demo. Our pipelines run on schedules that have nothing to do with how long the underlying thinking took. Our models transform quietly, in the background, producing outputs that are visible only when something breaks — or worse, when something breaks &lt;em>subtly&lt;/em> and doesn&amp;rsquo;t look broken at all.&lt;/p>
&lt;p>That&amp;rsquo;s the gap Atwood&amp;rsquo;s advice never closed. He fixed going dark for software teams. Nobody translated it for us.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="why-data-engineering-goes-dark-more-than-anyone-else">Why Data Engineering Goes Dark More Than Anyone Else&lt;/h3>
&lt;p>There&amp;rsquo;s a structural reason data work is so prone to invisibility.&lt;/p>
&lt;p>When a front-end engineer ships a feature, there&amp;rsquo;s an artifact: a button, a page, a flow you can click through. When a backend engineer merges a PR, there&amp;rsquo;s a diff: before and after, visible in the tool, reviewable by anyone. The work and the evidence of work are coupled. They arrive together.&lt;/p>
&lt;p>Data work decouples them. Benn Stancil, writing about the data industry, put it well when he described the fuzzy boundary between production and everything else in data teams — there&amp;rsquo;s often no clean staging/prod line, no clear &amp;ldquo;shipped&amp;rdquo; moment, no obvious artifact to point to. A dbt model runs. A dashboard refreshes. A pipeline completes. Did anything change? Did anything improve? Is this better than it was last week? The answers require knowing the context that lives in the engineer&amp;rsquo;s head, which is precisely the context that isn&amp;rsquo;t written down.&lt;/p>
&lt;p>This gets worse the longer the project. Migrations are the canonical example. A warehouse migration — say, moving from Redshift to Snowflake, or rebuilding a legacy reporting layer in dbt — might take three months of continuous work before a stakeholder sees a single dashboard change. During those three months, the data engineer is doing real, difficult, valuable work: tracing lineage, resolving naming conflicts, negotiating data contracts with upstream producers, writing tests that have never existed before. All of it is invisible. All of it looks like nothing to anyone watching from the outside.&lt;/p>
&lt;p>There&amp;rsquo;s also the problem that Locally Optimistic describes as the distinction between linear and circular projects. Linear projects have a known path and a clear endpoint. Circular projects — most of the genuinely interesting data work — have outcomes that depend heavily on what you discover along the way. The best-laid migration plan doesn&amp;rsquo;t survive contact with a production schema that hasn&amp;rsquo;t been properly documented since 2019. Exploration takes time that doesn&amp;rsquo;t map cleanly to sprint cards or progress percentages.&lt;/p>
&lt;p>When you combine those two things — structural invisibility and genuinely unpredictable progress — you get a discipline that almost naturally produces going-dark behaviour, even among engineers who are working diligently and in good faith. Mia wasn&amp;rsquo;t hiding. The nature of her work was hiding it for her.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-ic-chapter-staying-visible-without-performing-busyness">The IC Chapter: Staying Visible Without Performing Busyness&lt;/h3>
&lt;p>Staying visible doesn&amp;rsquo;t mean performing busyness. It doesn&amp;rsquo;t mean posting standup updates about tasks you haven&amp;rsquo;t started. It doesn&amp;rsquo;t mean narrating every dead end like it&amp;rsquo;s a breakthrough. That&amp;rsquo;s noise, and experienced leads learn to tune it out.&lt;/p>
&lt;p>What it does mean is producing &lt;em>artifacts&lt;/em> — real, durable, reviewable traces of your thinking — on a cadence that keeps your work legible to the people around you. The goal is to make silence meaningful. When there&amp;rsquo;s nothing from you for three days, the people you work with should know that&amp;rsquo;s fine because you&amp;rsquo;ve pre-declared it and pre-committed to what you&amp;rsquo;ll surface when you come up for air. When you&amp;rsquo;re grinding through something genuinely hard, the work itself should leave evidence that can be reviewed asynchronously, without requiring a meeting to explain it.&lt;/p>
&lt;p>&lt;strong>Draft PRs as the lowest-friction visibility tool you already have.&lt;/strong>&lt;/p>
&lt;p>Open a pull request on day one of a significant piece of work. Not when it&amp;rsquo;s ready for review — when you have a scaffold and an intent. GitHub&amp;rsquo;s draft PR feature exists exactly for this: your code can&amp;rsquo;t be merged, reviewers aren&amp;rsquo;t auto-requested, and everyone in the team can see that the work exists, what it&amp;rsquo;s targeting, and roughly how it&amp;rsquo;s going. Update it with a comment as you progress. &amp;ldquo;Blocked on upstream source delay — will resume Thursday.&amp;rdquo; &amp;ldquo;Discovered three extra joins needed — revised estimate to end of week.&amp;rdquo; &amp;ldquo;Models passing, running full test suite now.&amp;rdquo;&lt;/p>
&lt;p>This pattern turns the PR into the canonical artifact of the work — a living document that replaces scattered messages, status check-ins, and &amp;ldquo;hey, how&amp;rsquo;s that thing going?&amp;rdquo; conversations. The PR description, written well, tells the story of what changed, why it changed, and how it was tested. Six months from now, when a model breaks at 2am and the engineer who built it is on leave, that description is the thing that saves the on-call.&lt;/p>
&lt;p>&lt;strong>Commit messages as team communication.&lt;/strong>&lt;/p>
&lt;p>There&amp;rsquo;s a widely-cited piece by Chris Beams on how to write a good git commit message that&amp;rsquo;s been floating around engineering circles since 2014. The core argument is simple: the subject line tells you &lt;em>what&lt;/em> changed; the body tells you &lt;em>why&lt;/em>. A commit message that says &lt;code>fix bug&lt;/code> costs the next engineer ten minutes of archaeology. A commit message that says &lt;code>Fix null handling in fct_orders when currency column is missing&lt;/code> and then explains that a vendor renamed a field and the defensive &lt;code>.get()&lt;/code> was silently returning zero for three months — that commit message is worth something.&lt;/p>
&lt;p>The Conventional Commits specification formalises this further: &lt;code>feat:&lt;/code>, &lt;code>fix:&lt;/code>, &lt;code>refactor:&lt;/code>, &lt;code>docs:&lt;/code>, &lt;code>chore:&lt;/code>, each with optional scope. It sounds bureaucratic until you try to understand a six-month-old change at 9pm during an incident, and someone had the discipline to write &lt;code>refactor(revenue): consolidate gross margin calculation into shared macro&lt;/code> in the subject line. Then you understand immediately.&lt;/p>
&lt;p>&lt;strong>dbt documentation as a visibility discipline.&lt;/strong>&lt;/p>
&lt;p>Every model in a dbt project can carry descriptions at the model level and at the column level, written in the same &lt;code>schema.yml&lt;/code> file that defines your tests. Most teams don&amp;rsquo;t write them. Or they write them at launch and never update them. Or they live in a separate Confluence page that nobody visits.&lt;/p>
&lt;p>Here&amp;rsquo;s the thing: &lt;code>schema.yml&lt;/code> descriptions review in the same PR as the code they describe. They can&amp;rsquo;t silently drift from the implementation because they&amp;rsquo;re changed in the same commit. A model description that says &amp;ldquo;Revenue aggregated by market, excluding internal transfers — see ADR-004 for the exclusion logic&amp;rdquo; is a piece of team communication that persists across every engineer who will ever touch that model.&lt;/p>
&lt;p>Go further. Use dbt Exposures to document which dashboards and downstream artifacts depend on your models. A handful of YAML lines, committing which report or ML model consumes each table, converts &amp;ldquo;will this refactor break anything?&amp;rdquo; from a Slack archaeology project into something you can run a query against. When a finance stakeholder asks if their dashboard will be affected by a model change, you can show them the dependency chain. That&amp;rsquo;s a conversation that previously required guesswork. Now it requires a terminal command.&lt;/p>
&lt;p>&lt;strong>Architecture Decision Records for the decisions people will question.&lt;/strong>&lt;/p>
&lt;p>ADRs — short markdown files committed to &lt;code>docs/adr/&lt;/code> in your repository — document the decisions that will otherwise live only in one engineer&amp;rsquo;s memory. Why did you choose Snowflake over BigQuery? Why does the revenue model exclude returns that arrived after 30 days? Why is this model incremental rather than table? These questions come up six months later, in a meeting, when the person who made the decision has moved on or forgotten the reasoning. An ADR is three paragraphs and a status field. It takes fifteen minutes to write. It has saved me more than one painful meeting where the alternative was reconstructing a decision from Slack history.&lt;/p>
&lt;p>&lt;strong>Async standups that actually work.&lt;/strong>&lt;/p>
&lt;p>The daily standup ritual made sense when software teams sat together and needed a quick synchronising pulse. For data teams — often distributed across timezones, frequently doing work that looks the same from the outside regardless of whether it&amp;rsquo;s going well or terribly — the fifteen-minute morning video call has become a source of collective fiction. Everyone says &amp;ldquo;still on X, should be done soon&amp;rdquo; and nobody learns anything.&lt;/p>
&lt;p>Async standup tools like Geekbot (for Slack) or equivalent bots for Teams send each person the standup questions by DM on a schedule and post responses to a shared channel. The format that consistently performs best isn&amp;rsquo;t the classic three questions — it&amp;rsquo;s a single question: &lt;em>&amp;ldquo;What would you like the team to know today?&amp;rdquo;&lt;/em> That question has enough latitude to surface a blocker, a discovery, a risk, or a win. It invites signal rather than mandating a ritual. Responses posted to a shared channel are searchable, linkable, and skimmable — a week of them gives a manager more information than five daily standups would.&lt;/p>
&lt;p>The written standup format also decouples availability from communication. A data engineer in Melbourne doesn&amp;rsquo;t need to be online at the same time as her lead in London for him to understand what&amp;rsquo;s happening. The post is there when he arrives. She&amp;rsquo;s not interrupted in her maker&amp;rsquo;s block to attend a meeting she&amp;rsquo;ll spend eight minutes waiting through before she has sixty seconds to talk.&lt;/p>
&lt;p>One discipline worth enforcing: when you name a blocker in a written standup, link to it. Don&amp;rsquo;t write &amp;ldquo;blocked on upstream data delay&amp;rdquo; — write &amp;ldquo;blocked on upstream &lt;code>raw_orders&lt;/code> freshness, opened Issue #247 to track.&amp;rdquo; The standup post becomes navigable rather than descriptive. Anyone who wants to understand the blocker can follow the link. Anyone who just needs the pulse can read the post and move on.&lt;/p>
&lt;p>&lt;strong>Data diffs in CI — making the invisible visible before it ships.&lt;/strong>&lt;/p>
&lt;p>The traditional CI check for a dbt project tells you whether models compile and whether your schema-level tests pass. Unique keys, not-null constraints, accepted values — these are important. They catch structural problems before they reach production.&lt;/p>
&lt;p>What they don&amp;rsquo;t catch is when a model that was returning &lt;code>65.00&lt;/code> starts returning &lt;code>65.10&lt;/code>, or when a currency denomination silently shifts from whole dollars to cents because an upstream source changed its representation without changing its schema. Those changes are invisible to schema tests. They&amp;rsquo;re invisible to the log output. They&amp;rsquo;re invisible right up until a finance stakeholder notices that last month&amp;rsquo;s revenue looks 100× higher than expected in a board deck.&lt;/p>
&lt;p>Data diff tooling — whether through Datafold integrated into GitHub CI, or the open-source &lt;code>audit_helper&lt;/code> package in dbt, or a homegrown comparison query — runs value-level comparisons between the current state of a model and a reference point. Merge a PR that refactors how gross margin is calculated, and the CI comment shows you a row-level diff of what changed. Every affected row. Every affected column. Before the code reaches production.&lt;/p>
&lt;p>This changes the nature of what a PR review means. Instead of reviewing whether the SQL is logically correct — a genuinely hard thing to assess from reading code alone — reviewers can look at the data that would result. A refactor that should be semantically identical shows zero diff. One that inadvertently changes three percent of records in a specific market shows exactly which records and by how much. The artifact created by data diff in CI is one of the most powerful anti-going-dark tools available to a data team, because it makes the consequence of code changes legible to anyone who can read a table — not just the engineers who wrote the SQL.&lt;/p>
&lt;p>&lt;strong>Async documentation that travels with the work.&lt;/strong>&lt;/p>
&lt;p>One pattern that consistently separates data teams that are legible from those that aren&amp;rsquo;t: Loom or short screen recordings attached to PR descriptions on significant changes. Not for every fix. Not as a replacement for the written description. But for a model redesign, a migration milestone, or a new data product, a two-minute screen recording walking through the dbt DAG and explaining the decisions is worth more than three paragraphs of prose that nobody will read.&lt;/p>
&lt;p>The video doesn&amp;rsquo;t have to be polished. It has to exist. It converts work that is legible only to the person who did it into work that is legible to a finance analyst, a product manager, or a new data engineer joining in six months. It replaces the meeting where you&amp;rsquo;d have explained this anyway, except now it&amp;rsquo;s available asynchronously, can be rewatched, and is linked from the artifact it describes.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-distributed-team-problem">The Distributed Team Problem&lt;/h3>
&lt;p>Remote and distributed data teams face a compounded version of the going-dark problem. The structural invisibility of data work — no natural demo format, no clean shipped/not-shipped boundary — gets worse when your team spans timezones and your primary communication channel is a text-based messaging app with a scroll-back limit.&lt;/p>
&lt;p>The teams that navigate this best tend to share a principle: optimize for knowledge retrieval, not knowledge transfer. The goal isn&amp;rsquo;t to ensure that information passes from person A to person B in a meeting. The goal is to ensure that the information exists in a form that person B can find when they need it, regardless of whether person A is online.&lt;/p>
&lt;p>GitLab, operating as a fully remote company since its founding, formalised this as handbook-first communication: any decision, policy, or significant piece of information gets documented before it gets communicated or implemented. Not as bureaucracy, but because documentation creates the retrievable artifact. A verbal decision made in a meeting exists only in the attendees&amp;rsquo; memories. A written decision linked from a PR exists for the lifetime of the repository.&lt;/p>
&lt;p>For data teams, this translates practically into a few habits. Meeting decisions get a one-paragraph summary posted to a Teams thread, linked from the relevant GitHub issue. A verbal conversation about model design gets a brief written follow-up: &amp;ldquo;Following on from our chat — going with incremental strategy for &lt;code>fct_orders&lt;/code> because full refresh at scale was timing out in staging. ADR drafted at &lt;code>docs/adr/0012-incremental-fct-orders.md&lt;/code>.&amp;rdquo; The conversation happened. Now it has an artifact that travels.&lt;/p>
&lt;p>Long-running projects — migrations, platform changes, major refactors — benefit from a written brief at the start that defines the scope, the success criteria, and the milestone cadence. Not a Jira epic with fifty tasks. A two-page document that answers: what are we doing, why are we doing it, how will we know when we&amp;rsquo;re done, and what are the milestones where we&amp;rsquo;ll regroup. The brief is the thing you hand to a new stakeholder who asks what&amp;rsquo;s happening with the platform migration. The brief is the thing you update when the scope changes. The brief is what prevents the three-week silence.&lt;/p>
&lt;p>Locally Optimistic describes a useful scripting pattern for circular projects — the kind where the outcomes depend heavily on what you discover along the way. Rather than promising a deliverable, you promise a process: &amp;ldquo;We will investigate X, Y, and Z over the next two weeks. Those might not give us a conclusive answer, but we&amp;rsquo;ll regroup to discuss the findings and determine next steps.&amp;rdquo; That framing is honest. It sets appropriate expectations. And critically, it includes a defined moment of resurfacing — the regroup — which prevents the silence from becoming open-ended.&lt;/p>
&lt;p>Julia Evans has written persuasively about the value of maintaining a running document of your own accomplishments — not as a vanity project but as a memory aid. The problem she identifies is real: at review time, you&amp;rsquo;ve forgotten half of what you did in the last six months, and your manager has forgotten more than that. A brag document is a log of things that mattered: bugs found before they became incidents, migrations completed, models refactored, stakeholders unblocked, junior engineers mentored.&lt;/p>
&lt;p>This isn&amp;rsquo;t about inflating your achievements. It&amp;rsquo;s about making them visible in a context where the work itself produces no natural evidence trail.&lt;/p>
&lt;p>Equally useful: weekly notes. A short message to your team lead, every Friday, that says &amp;ldquo;this week I did X, I&amp;rsquo;m picking up Y on Monday, and I want to flag Z as a potential blocker.&amp;rdquo; It doesn&amp;rsquo;t have to be long. It doesn&amp;rsquo;t have to be polished. It just has to exist. Will Larson calls this the drip — communication on cadence, regardless of whether there&amp;rsquo;s exciting news. The drip is what makes silence informative. When you&amp;rsquo;ve been sending a Friday note every week for three months and then you stop, the absence is a signal. When you&amp;rsquo;ve never sent one, silence tells nobody anything.&lt;/p>
&lt;p>&lt;strong>Tanya Reilly&amp;rsquo;s glue trap — for senior ICs especially.&lt;/strong>&lt;/p>
&lt;p>If you&amp;rsquo;re a staff or principal data engineer, you&amp;rsquo;re probably doing a lot of work that doesn&amp;rsquo;t look like work from a promotion perspective: reviewing designs, unblocking others, noticing what&amp;rsquo;s slipping, maintaining relationships with stakeholders. Reilly calls this glue work — the coordination and communication and knowledge-transfer that holds projects together. It&amp;rsquo;s real work. It&amp;rsquo;s often the most valuable work on the team. And it&amp;rsquo;s completely invisible without a deliberate effort to create artifacts.&lt;/p>
&lt;p>Design proposals, written-up meeting decisions, documented onboarding pathways, group emails that summarise a decision made in a verbal conversation — these are the artifacts that make glue work visible. Without them, you&amp;rsquo;re building organisational infrastructure that has no evidence it exists.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-fear-dimension-whats-really-happening-when-engineers-go-dark">The Fear Dimension: What&amp;rsquo;s Really Happening When Engineers Go Dark&lt;/h3>
&lt;p>Here&amp;rsquo;s something I think we don&amp;rsquo;t talk about honestly enough — and I&amp;rsquo;m including myself in this.&lt;/p>
&lt;p>I&amp;rsquo;ve been on both sides of it. I&amp;rsquo;ve gone dark on work I wasn&amp;rsquo;t confident about, telling myself I&amp;rsquo;d surface when I had something worth showing. I&amp;rsquo;ve also been the lead who didn&amp;rsquo;t notice, or didn&amp;rsquo;t ask, until a stakeholder beat me to it. Neither version felt like failure at the time. Both were.&lt;/p>
&lt;p>Most going-dark behaviour isn&amp;rsquo;t strategic. It isn&amp;rsquo;t laziness. It&amp;rsquo;s fear — specifically, the fear that surfacing in-progress work will reveal that you don&amp;rsquo;t have it as together as you implied you did.&lt;/p>
&lt;p>Collins-Sussman put this precisely in the piece Atwood quoted: developers don&amp;rsquo;t want their peers to see mistakes or failures. They want to present themselves as infallible. They&amp;rsquo;re comfortable sharing code once it&amp;rsquo;s polished. The messy middle, the exploratory branches, the models with big question marks in the comments — that stays private.&lt;/p>
&lt;p>This fear is rational in the short term. Showing a half-finished thing invites feedback that can derail you. Showing a broken thing invites questions you can&amp;rsquo;t answer yet. Staying silent avoids both. The problem is that in the medium term, silence accumulates into something much worse: an engineer who surfaces three weeks into a problem with nothing to show, facing a stakeholder conversation where the first question is &amp;ldquo;why didn&amp;rsquo;t you just tell us?&amp;rdquo;&lt;/p>
&lt;p>Paloma Medina&amp;rsquo;s BICEPS framework identifies six core needs that, when threatened, produce withdrawal: Belonging, Improvement, Choice, Equality, Predictability, Significance. (Yes, it spells BICEPS. Yes, it will appear on a workshop slide near you. The underlying insight is still useful.) Lara Hogan, who has done more than anyone I know of to translate this framework into engineering management practice, makes the observation that going dark is often a flight response to one of these needs being threatened. A project pivot that makes the last three months of work feel wasted threatens Significance — and the engineer who feels that their work no longer matters is likely to disengage before they articulate why. An abrupt change in priorities threatens Predictability and Choice simultaneously. A team reorganisation threatens Belonging.&lt;/p>
&lt;p>The leader&amp;rsquo;s job, when they notice going-dark behaviour, isn&amp;rsquo;t to increase accountability pressure. It&amp;rsquo;s to identify which need has been threatened and address it directly. Fournier and others recommend this as the actual content of the 1:1 when something feels off: not &amp;ldquo;how&amp;rsquo;s the project going?&amp;rdquo; but &amp;ldquo;how are you doing?&amp;rdquo; The project status comes second. The human state comes first.&lt;/p>
&lt;p>Hogan&amp;rsquo;s Red/Yellow/Green check-in is a lightweight mechanism for this. At the start of a 1:1, an engineer can say &amp;ldquo;red&amp;rdquo; without having to explain why — just to signal that something is wrong. The permission to say &amp;ldquo;red&amp;rdquo; without an explanation is itself the thing that makes the saying possible. When explanation is required before disclosure, disclosure gets deferred until the explanation is formed — and by then, the problem has grown. Hogan&amp;rsquo;s observation: &lt;em>only having to say red, and not having to explain the why, is huge.&lt;/em> The absence of the explanation requirement is the mechanism. It lowers the cost of disclosure below the threshold where people start editing themselves.&lt;/p>
&lt;p>Amy Edmondson&amp;rsquo;s research on psychological safety is directly relevant here. Her finding, counterintuitively, was that better hospital units didn&amp;rsquo;t make fewer medication errors — they reported more of them. Her reframe: the better units weren&amp;rsquo;t more error-prone, they were more willing to surface problems early. The error rate looked higher because the environment was safe enough that people actually said what was happening.&lt;/p>
&lt;p>Google&amp;rsquo;s Project Aristotle, which studied 180 teams over several years, found the same pattern at scale: the single strongest predictor of team performance was whether team members felt safe to take interpersonal risks. Safe to say &amp;ldquo;I&amp;rsquo;m stuck.&amp;rdquo; Safe to say &amp;ldquo;I found a bug.&amp;rdquo; Safe to say &amp;ldquo;I don&amp;rsquo;t know how long this will take.&amp;rdquo;&lt;/p>
&lt;p>The data engineering implication is uncomfortable but important: if your team never has visible incidents, never surfaces broken assumptions, never flags a model that&amp;rsquo;s been silently wrong — that&amp;rsquo;s probably not a sign of excellence. It&amp;rsquo;s probably a sign that the environment doesn&amp;rsquo;t feel safe enough to tell the truth. The team that surfaces the most broken pipelines, the most failed tests, the most &amp;ldquo;we got this wrong&amp;rdquo; disclosures might be your best team. They&amp;rsquo;re the ones telling you what&amp;rsquo;s actually happening.&lt;/p>
&lt;p>The &amp;ldquo;I&amp;rsquo;ll tell them when it&amp;rsquo;s fixed&amp;rdquo; trap is the going-dark failure mode in its purest form. An engineer discovers a problem — a data quality issue, a budget overrun, a model that produces subtly wrong results. Rather than surface it, they decide to fix it first. Each day that passes raises the psychological cost of disclosure. The problem grows. The fix gets harder. Eventually, either the engineer surfaces in a state of exhaustion with a fix and hopes nobody asks hard questions, or a stakeholder finds the problem externally — at which point the disclosure isn&amp;rsquo;t just about the bug. It&amp;rsquo;s about the three weeks of silence.&lt;/p>
&lt;p>Fixing the number is not the same as fixing what the number did to people who trusted it.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-lead-chapter-building-conditions-where-your-team-doesnt-go-dark">The Lead Chapter: Building Conditions Where Your Team Doesn&amp;rsquo;t Go Dark&lt;/h3>
&lt;p>If you lead a data team, the going-dark problem is partly yours to own — not because your engineers are your responsibility in a parental sense, but because the &lt;em>conditions&lt;/em> that make going dark feel safe are conditions you set.&lt;/p>
&lt;p>Camille Fournier&amp;rsquo;s framing, in &lt;em>The Manager&amp;rsquo;s Path&lt;/em>, is that 1:1s are like oil changes: if you skip them, plan to get stranded. The 1:1 is the primary diagnostic for whether someone is going dark, and it works only if it creates genuine psychological safety — which means it has to be something other than a status update. If every 1:1 is &amp;ldquo;what are you working on this week?&amp;rdquo;, you&amp;rsquo;ll get accurate project status and zero signal about whether someone is actually stuck, scared, or burning out.&lt;/p>
&lt;p>Lara Hogan recommends opening the relationship with simple questions: what makes you grumpy? How will I know when you&amp;rsquo;re grumpy? What do you need from me when you&amp;rsquo;re struggling? These aren&amp;rsquo;t soft questions — they&amp;rsquo;re diagnostic instruments. They establish that the 1:1 is a place where the real state of things can be named.&lt;/p>
&lt;p>Her Red/Yellow/Green check-in is a lightweight version of the same tool. At the start of a 1:1, an engineer can say &amp;ldquo;red&amp;rdquo; without having to explain why — just to signal that something is wrong. The permission to say &amp;ldquo;red&amp;rdquo; without an explanation is itself the thing that makes the saying possible. When explanation is required before disclosure, disclosure gets deferred until the explanation is formed — and by then, the problem has grown.&lt;/p>
&lt;p>Will Larson frames visibility for leads as a mechanical practice rather than an art form. The goal isn&amp;rsquo;t to inspire your team to communicate — it&amp;rsquo;s to build systems that make communication the path of least resistance. A weekly update template, committed to as a team norm. An async standup channel where the format is so lightweight that posting takes two minutes. A retrospective cadence where surfacing what went wrong is the expected behaviour, not the exception.&lt;/p>
&lt;p>&lt;strong>Warning signs to watch for.&lt;/strong>&lt;/p>
&lt;p>Going dark rarely announces itself. The early signals are easy to miss.&lt;/p>
&lt;p>Standups that shift from specific to vague — &amp;ldquo;still on X&amp;rdquo; across multiple days with no detail — are an indicator. PRs that open in draft and stay there, accumulating commits but never progressing to review. Estimates that consistently slip by a day or two, week after week, without the engineer naming the slippage explicitly. A previously engaged person who starts attending meetings on camera-off, contributing less, leaving threads unresponded to.&lt;/p>
&lt;p>The most useful signal is often external. When a stakeholder or adjacent team member asks you what&amp;rsquo;s happening with a project before the engineer working on it has raised anything — that&amp;rsquo;s the canary. The information is flowing around the engineer rather than through them. That&amp;rsquo;s a conversation to have in the next 1:1, and it&amp;rsquo;s worth naming directly: &amp;ldquo;I heard from finance about the revenue models. Tell me where things actually are.&amp;rdquo;&lt;/p>
&lt;p>&lt;strong>Deep work is not going dark — and the distinction matters.&lt;/strong>&lt;/p>
&lt;p>There&amp;rsquo;s a real risk of over-correcting. Data engineering requires extended, cognitively-demanding focus. Paul Graham&amp;rsquo;s framing of the maker&amp;rsquo;s schedule versus the manager&amp;rsquo;s schedule is directly applicable: a data engineer tracing lineage through a DAG, rebuilding a model from first principles, or debugging an intermittent pipeline failure needs hours of uninterrupted context, not a constant stream of check-in messages.&lt;/p>
&lt;p>Deep work and going dark are not the same thing. The difference is legibility.&lt;/p>
&lt;p>Deep work has a pre-declared output and timebox. Going dark does not. Deep work is legible at a coarse grain — &amp;ldquo;I&amp;rsquo;m heads-down on X until Thursday, update Friday&amp;rdquo; — even when it&amp;rsquo;s illegible at a fine grain. Going dark is illegible at every grain.&lt;/p>
&lt;p>The mechanism that makes the distinction: the cadence contract. Before going heads-down, the engineer declares: here&amp;rsquo;s what I&amp;rsquo;m working on, here&amp;rsquo;s when you&amp;rsquo;ll hear from me, here&amp;rsquo;s what would make me surface early. &amp;ldquo;If I&amp;rsquo;m stuck for more than a day, I&amp;rsquo;ll tell you. If I&amp;rsquo;m on track, Friday update. Ping me in #urgent if there&amp;rsquo;s a fire.&amp;rdquo; With that contract in place, silence during the week means &amp;ldquo;on track.&amp;rdquo; Without it, silence means nothing — or worse, triggers anxiety that prompts the exact interruptions that fragment deep work.&lt;/p>
&lt;p>The goal is to make silence meaningful. That requires prior investment in communication, not more communication during the work itself.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-visible-data-team-playbook">The Visible Data Team Playbook&lt;/h3>
&lt;p>There are a handful of practices that the best data teams I&amp;rsquo;ve seen use consistently. None of them are revolutionary. Most of them are embarrassingly simple. They work because they&amp;rsquo;re habitual, not because they&amp;rsquo;re clever.&lt;/p>
&lt;p>&lt;strong>Weekly insight posts.&lt;/strong>&lt;/p>
&lt;p>A practice described from data teams at companies like Netlify involves a dedicated Teams channel where each analyst posts one insight per week — not a status update, not a task list, but one thing they found interesting in the data. The format is designed to answer five questions: what am I looking at, why should you care, what caught my eye, where can you learn more, and who do you ask if you have questions.&lt;/p>
&lt;p>The explicit benefit isn&amp;rsquo;t just visibility — it&amp;rsquo;s demonstrating the range of what the data team does. Stakeholders who only see the data team when something breaks, or when they submit a request, have no model of what the team actually knows. A weekly insight post gives them one. Over time, it becomes the evidence base for why the team should be resourced, why data work is complex, why the business should care about data quality. It&amp;rsquo;s not reporting. It&amp;rsquo;s positioning.&lt;/p>
&lt;p>&lt;strong>Parity dashboards during migrations.&lt;/strong>&lt;/p>
&lt;p>One of the most effective anti-going-dark tools for long migrations is a parity dashboard: a side-by-side view of critical metrics computed in the old and new systems simultaneously. Before you&amp;rsquo;ve finished the migration, you build the parallel calculation and make it visible. Divergence between the two numbers becomes the artifact of progress — something demoable, something reviewable, something that gives a stakeholder a concrete thing to look at even when the underlying models are still being rebuilt.&lt;/p>
&lt;p>The parity dashboard also builds confidence in a way that purely technical communication can&amp;rsquo;t. When a finance stakeholder sees that their ARR number matches between the legacy and new system for three weeks running, they start to trust the migration in a way that no amount of &amp;ldquo;we&amp;rsquo;ve tested this thoroughly&amp;rdquo; can produce. Evidence beats assertion.&lt;/p>
&lt;p>&lt;strong>Incident communication as a team norm.&lt;/strong>&lt;/p>
&lt;p>The instinct to fix a data quality issue quietly and move on is almost universal among data engineers. The impulse is understandable — nobody wants to be the person who caused the problem, and surfacing it draws attention to the failure. But the silent fix has a consistent failure mode: it gets discovered later, retrospectively, in a context where the silence becomes the story rather than the fix.&lt;/p>
&lt;p>Monte Carlo&amp;rsquo;s research showed that the data teams with the highest stakeholder trust tend to over-communicate around incidents — sending post-mortems after significant issues, flagging discoveries before fixes are complete, maintaining a clear status trail from &amp;ldquo;investigating&amp;rdquo; to &amp;ldquo;resolved.&amp;rdquo; What one team described as transformative wasn&amp;rsquo;t their mean time to resolution. It was the shift from reactive firefighting to proactive communication that demonstrated awareness and control. The communication is the product, as much as the fix.&lt;/p>
&lt;p>A simple practice: any time you discover a data quality issue, the first act is naming it in writing — a Teams message, a GitHub issue, an incident channel post — before you start fixing it. Not to perform distress, but to start the paper trail. &amp;ldquo;Found an issue with the &lt;code>fct_orders&lt;/code> null handling — investigating now, will update by EOD.&amp;rdquo; That message, and the chain of messages that follow it, is what turns an incident into evidence of a well-run team rather than evidence of a careless one.&lt;/p>
&lt;p>&lt;strong>ADRs for the decisions people will question.&lt;/strong>&lt;/p>
&lt;p>The architecture decisions that feel obvious when you make them feel arbitrary when someone encounters them six months later without context. Why incremental and not table? Why Snowflake and not BigQuery? Why is revenue calculated at invoice date rather than payment receipt date? These questions have answers. The answers live in someone&amp;rsquo;s head, or in a Confluence page that nobody linked to the model.&lt;/p>
&lt;p>ADRs — short markdown files in &lt;code>docs/adr/&lt;/code>, committed to the repository — are the durable artifact of those decisions. They don&amp;rsquo;t need to be long. They need a title, a status, a description of the context, the decision, and the consequences. Once written and committed, they&amp;rsquo;re part of the repository&amp;rsquo;s history. They review in PRs. They can be linked from model descriptions. They don&amp;rsquo;t disappear when someone leaves the team.&lt;/p>
&lt;p>&lt;strong>dbt Exposures for impact legibility.&lt;/strong>&lt;/p>
&lt;p>If you&amp;rsquo;re using dbt and you haven&amp;rsquo;t adopted Exposures, this is worth doing this week. Exposures are YAML-defined references to downstream artifacts — dashboards, ML models, reverse-ETL syncs — that extend the dbt DAG past the warehouse. A few lines in &lt;code>schema.yml&lt;/code> that say &amp;ldquo;this model feeds the Weekly Revenue dashboard&amp;rdquo; means that when someone runs &lt;code>dbt ls --select +exposure:weekly_revenue&lt;/code>, they get the complete list of upstream models that would need to change to affect that dashboard.&lt;/p>
&lt;p>For a migration, this means you can answer &amp;ldquo;will this change break the CEO&amp;rsquo;s dashboard?&amp;rdquo; without archaeology. For a refactor, it means PR reviewers have a concrete list of stakeholders to notify. For a data quality incident, it means you know who to tell before they find out themselves.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="going-back-to-mia">Going Back to Mia&lt;/h3>
&lt;p>After my conversation with Mia — the one where she walked me through three weeks of extraordinary, invisible work — we had a more direct conversation about why none of it had been visible.&lt;/p>
&lt;p>Her answer was the one Collins-Sussman identified in 2008: she didn&amp;rsquo;t want to show something until it was right. She was worried about raising a false alarm. She was concerned that surfacing work in progress would invite questions she couldn&amp;rsquo;t answer yet. There was also something she said that stuck with me: &amp;ldquo;I thought if I kept my head down long enough, I&amp;rsquo;d come up with something worth showing.&amp;rdquo;&lt;/p>
&lt;p>All of those concerns were reasonable. None of them were wrong, exactly. What she&amp;rsquo;d underestimated was the cost of the silence — not to herself, but to the people around her. The finance stakeholder who&amp;rsquo;d gone to me with questions would have gone to Mia first if they&amp;rsquo;d had any signal that the work was in progress and legible. My not knowing, in a week where I&amp;rsquo;d been asked directly by leadership about the revenue model timeline, was a problem I hadn&amp;rsquo;t been equipped to solve.&lt;/p>
&lt;p>What I should have done differently, and what I&amp;rsquo;ve done differently since: established the cadence contract at the start of the project, not after the silence had already built. Agreed, before the first commit, what &amp;ldquo;on track&amp;rdquo; communication would look like — what the update cadence would be, what would trigger an early surface, and what silence during a declared heads-down period meant. With that contract in place, three weeks of quiet is three weeks of &amp;ldquo;on track, as agreed.&amp;rdquo; Without it, three weeks of quiet is three weeks of &amp;ldquo;nobody knows what&amp;rsquo;s happening.&amp;rdquo;&lt;/p>
&lt;p>The artifact we built together after that conversation was simple. For the remaining six weeks of the migration, Mia opened a draft PR on Monday morning with a brief plan for the week. She posted a comment each Friday with what had been done, what was pending, and any surprises. She added a parity dashboard that the finance lead could check at any time. She wrote three ADRs for the decisions that had taken the most deliberation.&lt;/p>
&lt;p>The work didn&amp;rsquo;t change. The visibility changed. And the experience of the migration — for Mia, for the stakeholders, for me — changed entirely.&lt;/p>
&lt;p>At the end of the project, the finance lead sent a message that I still think about. She said: &amp;ldquo;I&amp;rsquo;ve never felt so informed during a data migration. Usually I find out things are done when the thing I was waiting for just starts working.&amp;rdquo;&lt;/p>
&lt;p>Mia had gone dark for three weeks. For the remaining six, she hadn&amp;rsquo;t. The difference wasn&amp;rsquo;t effort. It was artifacts.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-closing-thought">The Closing Thought&lt;/h3>
&lt;p>Jeff Atwood closed his post with something direct: don&amp;rsquo;t go dark. Don&amp;rsquo;t be the developer in the room who hides their code until it&amp;rsquo;s done. Sharing work in progress feels riskier than it is. The feedback and connection it produces are worth more than the protection the silence provides.&lt;/p>
&lt;p>He was right in 2008. The principle has only compounded since, as data teams have become more distributed, more remote, and more responsible for the kind of long-horizon work that naturally produces invisibility.&lt;/p>
&lt;p>The tradeoff is real, though. Better visibility takes time that could go to the work itself. An ADR written at the end of a hard week is time you&amp;rsquo;re not spending fixing the problem. A parity dashboard adds engineering overhead before the migration is even done. Writing commit messages properly requires discipline you don&amp;rsquo;t always have at 5pm on a Friday when you just want to push and go home.&lt;/p>
&lt;p>The case I&amp;rsquo;m making isn&amp;rsquo;t that artifacts are free. It&amp;rsquo;s that the alternative — silence — costs more than you can see while you&amp;rsquo;re inside it. The cost lands on the stakeholder who went to your manager instead of you. On the on-call engineer at 2am who has nothing to go on. On the project that looked fine until it suddenly didn&amp;rsquo;t.&lt;/p>
&lt;p>Amy Edmondson&amp;rsquo;s hospital study keeps coming back to me. She&amp;rsquo;d expected to find that better units made fewer errors. Instead she found that they reported more. Her reframe: maybe the better units don&amp;rsquo;t have fewer problems. Maybe they&amp;rsquo;re just more willing to say so.&lt;/p>
&lt;p>The best data teams are the same. Not the ones with the fewest broken pipelines — the ones where a broken pipeline doesn&amp;rsquo;t stay quiet for long.&lt;/p>
&lt;p>Three weeks is going dark. One good artifact is enough. The rest follows from that.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Career Development</category><category>Leadership</category><category>Communication</category><category>Data Quality</category><category>GitHub</category><category>dbt</category><category>Remote Work</category><category>Career Development</category><category>Engineering Culture</category><category>Psychological Safety</category></item><item><title>The Broken Window in Your Data Pipeline</title><link>https://ghostinthedata.info/posts/2026/2026-05-09-broken-window-theory/</link><pubDate>Sat, 09 May 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-05-09-broken-window-theory/</guid><author>Chris Hillman</author><description>A single ignored data quality issue doesn't stay local. In pipelines, broken windows travel — and by the time anyone notices, the damage is already downstream.</description><content:encoded>&lt;p>There&amp;rsquo;s a particular kind of data problem that doesn&amp;rsquo;t announce itself. It accumulates.&lt;/p>
&lt;p>We were receiving Salesforce data through delta extraction — sensible in theory, because full snapshots can run to hundreds of terabytes and less than 1% of records change on any given day. The problem is that deltas require someone to know what &amp;ldquo;changed&amp;rdquo; means. In Salesforce, that&amp;rsquo;s less obvious than it sounds. Watch a &lt;code>last_modified&lt;/code> column and you&amp;rsquo;ll miss objects that get updated when a related object changes, without their own timestamp reflecting it.&lt;/p>
&lt;p>Over time: drift. Orphan records. Data that &lt;em>looks&lt;/em> current because the record exists, but isn&amp;rsquo;t. The fix was documented — run a full snapshot periodically to correct the accumulated drift, painful as that was — and everyone with any context on the system knew it.&lt;/p>
&lt;p>What happened was roughly this: the drift accumulated silently until something downstream looked wrong. An investigation traced it back to the delta logic. The full snapshot was run. The problem was resolved. Notes were written. The workaround was filed away.&lt;/p>
&lt;p>And then, about twelve months later, the same conversation happened again.&lt;/p>
&lt;hr>
&lt;p>What made that moment stick wasn&amp;rsquo;t the technical failure. It was the recognition that the window had been broken for a while. Everybody who walked past it knew it was broken. We&amp;rsquo;d even put a note next to it.&lt;/p>
&lt;p>I&amp;rsquo;ve been that person — the one who wrote the documentation, filed the Jira ticket, and moved on. Which is probably why I remember the twelve-month cycle so clearly. And why I started paying attention to the pattern of it, across teams and companies, long after that particular incident.&lt;/p>
&lt;p>We just hadn&amp;rsquo;t fixed it.&lt;/p>
&lt;hr>
&lt;p>Bear with me here, because what I&amp;rsquo;m about to describe starts with an abandoned car in the Bronx in 1969 — and I promise it ends somewhere relevant to your dbt models.&lt;/p>
&lt;/br>
&lt;/br>
&lt;h3 id="a-broken-window-in-the-bronx-a-broken-window-in-your-warehouse">A broken window in the Bronx, a broken window in your warehouse&lt;/h3>
&lt;/br>
&lt;p>A Stanford psychologist named Philip Zimbardo ran a strange experiment. He abandoned two cars — one in the Bronx, one in Palo Alto — and watched what happened.&lt;/p>
&lt;p>The Bronx car was stripped within twenty-four hours. Within three days it was completely gutted.
The Palo Alto car sat untouched for a week — until Zimbardo himself walked up with a sledgehammer and broke a window. Within hours, it had been stripped too.&lt;/p>
&lt;p>Same car. Different signal.&lt;/p>
&lt;p>Thirteen years later, criminologists James Q. Wilson and George Kelling built a theory on that experiment. If a window in a building is broken and left unrepaired, they argued, all the rest will soon follow. Not because there&amp;rsquo;s a particular breed of window-breaker lurking around, but because an unrepaired window sends a message: &lt;em>nobody here cares&lt;/em>. And once that signal is broadcast, breaking more windows costs nothing.&lt;/p>
&lt;p>The insight was semiotic, not structural. It wasn&amp;rsquo;t about windows. It was about what an unrepaired window communicates about the norms of a place.&lt;/p>
&lt;p>The theory eventually found its way into software. Andrew Hunt and David Thomas put it into &lt;em>The Pragmatic Programmer&lt;/em> almost verbatim: don&amp;rsquo;t leave broken windows — bad designs, wrong decisions, poor code — unrepaired. They&amp;rsquo;d watched clean systems deteriorate quickly once windows started breaking. The mechanism was the same. A developer looking at a messy codebase thinks: &lt;em>if someone else got away with being careless, maybe I can too.&lt;/em> The norm shifts. The entropy accelerates.&lt;/p>
&lt;p>Researchers at Empirical Software Engineering ran a controlled experiment — twenty-nine developers, codebases seeded with either high or low technical-debt density — and found exactly what Hunt and Thomas had intuited. Pre-existing debt measurably caused developers to introduce &lt;em>more&lt;/em> debt. Non-descriptive variable names. Duplicated logic instead of reuse. Additional code smells. The broken window was statistically contagious.&lt;/p>
&lt;p>It&amp;rsquo;s a compelling idea, and it maps well to software. But it doesn&amp;rsquo;t map perfectly to data engineering. And the gap between &amp;ldquo;maps well&amp;rdquo; and &amp;ldquo;maps perfectly&amp;rdquo; is where teams get into serious trouble.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-thing-thats-different-about-data">The thing that&amp;rsquo;s different about data&lt;/h3>
&lt;/br>
&lt;p>When a broken window appears in a codebase, it&amp;rsquo;s local. A messy function, an undocumented module, a class that&amp;rsquo;s doing five things at once — these are ugly, they invite imitation, and they slow future development. But they stay where they are. They don&amp;rsquo;t go anywhere.&lt;/p>
&lt;p>A broken window in a data pipeline doesn&amp;rsquo;t stay where it is.&lt;/p>
&lt;p>It travels.&lt;/p>
&lt;p>That&amp;rsquo;s the thing nobody really talks about when they apply broken-windows thinking to data work. In a pipeline, everything is connected. A schema drift in a source table doesn&amp;rsquo;t just make that table annoying to work with — it silently corrupts every model downstream that touches that field. Which means it corrupts every dashboard that uses those models. Which means it corrupts the metrics those dashboards expose. Which means it corrupts the business decisions made from those metrics.&lt;/p>
&lt;p>The broken window is in row one. By the time someone notices, the damage is in the boardroom.&lt;/p>
&lt;p>And unlike software, where the damage is visible — a stacktrace, a failing build, a crash — data damage is often invisible. The pipeline still runs. The dashboard still renders. The report still reconciles. The numbers just happen to be wrong, quietly, for reasons nobody can easily trace.&lt;/p>
&lt;p>Software bugs produce noise. Bad data produces silence.&lt;/p>
&lt;p>In software, the broken window signals disorder and invites imitation. In data pipelines, it does all of that &lt;em>and&lt;/em> it propagates at machine speed through every downstream system that trusts the upstream to be clean. By the time the propagation is discovered, it&amp;rsquo;s usually been underway for a while.&lt;/p>
&lt;p>This is the propagation problem. It&amp;rsquo;s not just that bad data begets more bad data. It&amp;rsquo;s that one bad window can contaminate an entire watershed.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="what-a-contaminated-watershed-actually-looks-like">What a contaminated watershed actually looks like&lt;/h3>
&lt;/br>
&lt;p>In May 2022, Unity Software — the game engine company — disclosed something remarkable in their SEC earnings filing. Their Audience Pinpointer ML model, the system that powered their ad targeting business, had ingested bad data from a large customer. That data had corrupted their training set.&lt;/p>
&lt;p>CEO John Riccitiello on the earnings call: &amp;ldquo;we lost the value of a portion of our training data due in part to us ingesting bad data from a large customer.&amp;rdquo;&lt;/p>
&lt;p>The bad data didn&amp;rsquo;t just produce one wrong prediction. It poisoned the model weights. Every subsequent retrain was built on a contaminated foundation. The model had to be taken offline, the bad data removed, and training restarted from scratch. The estimated impact was $110 million in revenue for 2022. The stock dropped 37% in a single day. Market cap losses in the billions.&lt;/p>
&lt;p>The broken window wasn&amp;rsquo;t in Unity&amp;rsquo;s systems — it was in data they were &lt;em>ingesting&lt;/em>. Once it crossed the boundary into their training pipeline, it propagated in the only direction data knows: forward and downstream, embedding itself into every layer of the system that touched it.&lt;/p>
&lt;p>Riccitiello&amp;rsquo;s pledge after the fact was almost poignant: &amp;ldquo;We are deploying monitoring, alerting and recovery systems and processes to promptly mitigate future events.&amp;rdquo; The observability came after the disaster. The window had already broken every other window in the building.&lt;/p>
&lt;hr>
&lt;p>These are the normal failure mode of connected data systems. The propagation isn&amp;rsquo;t a bug in the design — it&amp;rsquo;s inherent to the architecture. Data flows in one direction. Trust flows with it. The moment you have a pipeline, you have propagation risk.&lt;/p>
&lt;p>The question isn&amp;rsquo;t whether your broken windows will propagate. It&amp;rsquo;s how far they&amp;rsquo;ll travel before anyone notices.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-tools-we-built-to-tolerate-this">The tools we built to tolerate this&lt;/h3>
&lt;/br>
&lt;p>Modern data tooling has given us genuinely sophisticated ways to formally accept broken windows.&lt;/p>
&lt;p>dbt — the transformation tool most data engineering teams live inside — has a &lt;code>severity&lt;/code> configuration on data tests. You can set a test to &lt;code>warn&lt;/code> instead of &lt;code>error&lt;/code>. The test runs, detects a problem, and&amp;hellip; doesn&amp;rsquo;t fail the pipeline. It records the warning and moves on.&lt;/p>
&lt;p>The documentation is cheerful about it: &amp;ldquo;Maybe 1 duplicate record can count as a warning, but 10 duplicate records should count as an error.&amp;rdquo;&lt;/p>
&lt;p>In principle, this is sensible. In practice, the &lt;code>warn&lt;/code> threshold becomes a Schelling point. Teams configure tests to warn to avoid breaking CI. The warnings accumulate. And then — almost invariably — the warnings become wallpaper. &lt;em>(Go check your own test results right now. Count how many have been in warn state for more than two weeks. I&amp;rsquo;ll wait.)&lt;/em>&lt;/p>
&lt;p>dbt also has a feature called &lt;code>store_failures&lt;/code> that saves failed records to an audit schema table. The test is doing its job. The failures are being recorded. They overwrite the previous run&amp;rsquo;s failures on the next run. If nobody is actively querying that audit table — and almost nobody is — the failures exist only to make the test feel like it&amp;rsquo;s being taken seriously.&lt;/p>
&lt;p>It&amp;rsquo;s a passive graveyard. The window is monitored. Nobody fixes the glass.&lt;/p>
&lt;p>Airflow has an equivalent pattern. The &lt;code>soft_fail&lt;/code> parameter on sensors means that if an exception is raised — the source system is down, the file hasn&amp;rsquo;t arrived — the task is marked as &lt;em>skipped&lt;/em> rather than &lt;em>failed&lt;/em>. Downstream tasks, depending on their trigger rules, may also skip. An entire branch of your DAG quietly collapses to a skipped state, which most pipelines treat as benign, and your stakeholders get a dashboard with no data in it instead of an error message.&lt;/p>
&lt;p>Retries do something similar. A flaky source that fails 30% of the time gets &lt;code>retries=3&lt;/code> configured. The task eventually succeeds on the third attempt. The 30% failure rate never surfaces as an anomaly in any meaningful way. Until the source dies entirely, at which point the symptom everyone responds to is not &amp;ldquo;this has been flaky for six months&amp;rdquo; but &amp;ldquo;this suddenly started failing today.&amp;rdquo;&lt;/p>
&lt;p>None of this is the fault of the tools. dbt and Airflow are doing what they&amp;rsquo;re designed to do. The issue is that the default ergonomics of both make &lt;em>tolerating&lt;/em> failure significantly easier than &lt;em>stopping propagation&lt;/em>. &amp;ldquo;Don&amp;rsquo;t break the build&amp;rdquo; is a more convenient goal than &amp;ldquo;don&amp;rsquo;t ship broken data downstream.&amp;rdquo; The tools give you excellent knobs for the former and adequate knobs for the latter.&lt;/p>
&lt;p>Chad Sanderson — formerly at Convoy, now building in the data contracts space — has a name for what this produces over time. He calls it the POSIWID principle: the Purpose Of a System Is What It Does. If your data pipelines are consistently producing low-quality data, and your team consistently tolerates that, then whatever you think the purpose of your data platform is, its actual purpose is to enable teams to move fast, ship without accountability, and tolerate breakages.&lt;/p>
&lt;p>The broken windows aren&amp;rsquo;t exceptions. They&amp;rsquo;re the product.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="weve-built-monitoring-why-isnt-it-working">We&amp;rsquo;ve built monitoring. Why isn&amp;rsquo;t it working?&lt;/h3>
&lt;/br>
&lt;p>Data observability emerged as a formal discipline. The concept — largely popularised by Barr Moses at Monte Carlo — drew directly from the software world&amp;rsquo;s SRE and distributed systems monitoring practices: freshness, distribution, volume, schema, lineage. Five pillars. Automated anomaly detection. Alerts in Slack when something looks wrong.&lt;/p>
&lt;p>The framing was right. The problem is what happened next.&lt;/p>
&lt;p>Monte Carlo&amp;rsquo;s own telemetry — across millions of monitored tables — shows that alert engagement drops 15% when a Slack channel exceeds 50 alerts per week, and a further 20% past 100 per week. The data observability system has a data quality problem: the signal-to-noise ratio degrades, and so does the human response.&lt;/p>
&lt;p>This isn&amp;rsquo;t specific to data. The clinical literature on alarm fatigue is sobering. ICU environments average more than 150 alarms per bed per day. Studies consistently find that 72% to 99% of those alarms are non-actionable. The clinical community&amp;rsquo;s response to this was institutional — the Joint Commission made alarm safety a National Patient Safety Goal in 2014 — because they recognised that a monitoring system that produces more noise than signal doesn&amp;rsquo;t just fail to help. It actively worsens outcomes by training humans to stop responding.&lt;/p>
&lt;p>Cybersecurity teams face the same thing. Security operations centres receive thousands of alerts daily; most go unaddressed, not because analysts are lazy, but because the ratio of genuine signals to false positives has degraded to the point where sustained attention is cognitively impossible.&lt;/p>
&lt;p>The Google SRE book is direct on this: &amp;ldquo;Every page should be actionable. If a page merely merits a robotic response, it shouldn&amp;rsquo;t be a page.&amp;rdquo;&lt;/p>
&lt;p>In data engineering, the equivalent of a &amp;ldquo;robotic response&amp;rdquo; is the acknowledged-and-unresolved alert. The monitor that fires every Tuesday morning, gets a thumbs up in Teams, gets added to the &amp;ldquo;known issues&amp;rdquo; document, and never gets fixed. The window is now monitored. The fact that it&amp;rsquo;s monitored makes the team feel responsible. The window stays broken.&lt;/p>
&lt;p>There&amp;rsquo;s a term for this: observability theatre. The infrastructure of visibility exists. The dashboards are green. Nobody&amp;rsquo;s actually looking at the glass.&lt;/p>
&lt;p>This is where the broken windows metaphor earns its keep most fully. Wilson and Kelling weren&amp;rsquo;t saying that disorder is bad because it looks bad. They were saying that an unrepaired broken window sends a signal that no one cares — and that signal is the actual mechanism of decay. Monitoring a broken window without repairing it sends exactly the same signal. Possibly a worse one, because now everyone knows the problem is being tracked and nobody&amp;rsquo;s acting on it. The norm becomes: acknowledged problems are not necessarily fixed problems.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-stakeholder-sees-it-differently">The stakeholder sees it differently&lt;/h3>
&lt;/br>
&lt;p>Here is the asymmetry that data teams consistently underestimate.&lt;/p>
&lt;p>When an engineer looks at a broken window in a data pipeline — a column with a known null rate, a test that&amp;rsquo;s been in warn state for three weeks, a DAG that soft-fails every second run — they see technical debt. A problem to eventually be addressed. Something that exists on a spectrum of severity, probably not at the top of the list.&lt;/p>
&lt;p>When a business stakeholder encounters the downstream consequence of that broken window — a dashboard that contradicts a number they just presented to the CFO, a metric that moved in an unexplained direction, a report that doesn&amp;rsquo;t reconcile with another report — they don&amp;rsquo;t think &amp;ldquo;technical debt.&amp;rdquo; They think: &lt;em>can I trust the data team?&lt;/em>&lt;/p>
&lt;p>Benn Stancil — one of the clearest thinkers writing about this — put it well: &amp;ldquo;Trust is built, and blown up, by the outputs — and specifically, the consistency of those outputs.&amp;rdquo; He uses a vivid image that I keep coming back to: &amp;ldquo;Data and the dashboards that display it create a shared sense of reality. Looking at two dashboards that don&amp;rsquo;t match is like looking out two adjacent windows and not seeing the same thing.&amp;rdquo;&lt;/p>
&lt;p>That&amp;rsquo;s the experience of broken-window propagation from the stakeholder&amp;rsquo;s side. Two adjacent windows. Different views. A reality that doesn&amp;rsquo;t cohere.&lt;/p>
&lt;p>Monte Carlo&amp;rsquo;s 2023 State of Data Quality research put a number to the trust inversion that most data engineers already sense. In 2022, 47% of respondents reported that business stakeholders identified data issues first &amp;ldquo;all or most of the time&amp;rdquo; — more often than the data team itself. By 2023, that figure had risen to 74%. &lt;em>(Three in four. Let that land.)&lt;/em>&lt;/p>
&lt;p>Think about what that means. In three out of four data incidents, the people who rely on the data found the problem before the people who built it. The pipeline runs, the test passes, the dashboard renders, and somewhere downstream a business analyst is staring at a number that doesn&amp;rsquo;t look right and is about to send a message that starts with: &amp;ldquo;Quick question about this figure&amp;hellip;&amp;rdquo;&lt;/p>
&lt;p>The damage isn&amp;rsquo;t technical. It&amp;rsquo;s relational. And it compounds in a way that&amp;rsquo;s harder to reverse than any schema migration.&lt;/p>
&lt;p>Thomas Redman — who has spent decades studying data quality — made this point in Harvard Business Review: &amp;ldquo;When data are unreliable, managers quickly lose faith in them and fall back on their intuition to make decisions.&amp;rdquo; Once that happens, you haven&amp;rsquo;t just produced bad data. You&amp;rsquo;ve trained decision-makers to ignore good data too, because they can no longer distinguish between the two.&lt;/p>
&lt;p>The broken window didn&amp;rsquo;t just propagate through the pipeline. It propagated into the culture.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="nobody-owns-the-broken-window">Nobody owns the broken window&lt;/h3>
&lt;/br>
&lt;p>There&amp;rsquo;s a specific pattern worth naming, because it&amp;rsquo;s where most data teams I&amp;rsquo;ve worked with have the most unacknowledged broken windows: the orphaned asset.&lt;/p>
&lt;p>The pipeline that was built by an engineer who left eighteen months ago. The table in the warehouse that exists in the data catalogue but whose owner field says &amp;ldquo;Unknown&amp;rdquo; or, worse, the name of someone who&amp;rsquo;s no longer at the company. The dbt model that runs in production, that feeds two dashboards, that nobody on the current team can fully explain.&lt;/p>
&lt;p>These aren&amp;rsquo;t broken in any obvious sense. They run. They load. The tests pass, or they&amp;rsquo;re set to warn. But they&amp;rsquo;re broken windows in the original sense: they signal that nobody cares. And because nobody owns them, nobody&amp;rsquo;s in a position to repair them even when something goes wrong.&lt;/p>
&lt;p>What typically happens is this: a new engineer joins the team. They need to understand how the data flows. They look at the catalogue. They look at the undocumented table. They look at the model that references it. They look at the Jira ticket from fourteen months ago that says &amp;ldquo;Known issue — downstream teams aware.&amp;rdquo; And they make a rational decision: don&amp;rsquo;t touch it, build around it, replicate it if necessary.&lt;/p>
&lt;p>The broken window has now inspired a new window. Same mechanism as the Bronx car. Different materials.&lt;/p>
&lt;p>The organisational research on this is unambiguous. Ron Westrum&amp;rsquo;s typology of organisational cultures — validated empirically by the DORA research programme across thousands of software teams — found that information flow predicts safety and performance more reliably than almost any structural variable. In pathological cultures, information is hoarded or withheld for political reasons. In generative cultures, information flows freely, failures are shared, and problems get fixed because surfacing problems is rewarded rather than penalised.&lt;/p>
&lt;p>Amy Edmondson&amp;rsquo;s research on psychological safety adds the other half: 85% of employees have withheld important information from their manager due to fear of speaking up. In data teams, this looks like: the junior engineer who noticed the null rate had been climbing for two weeks and didn&amp;rsquo;t raise it because they weren&amp;rsquo;t sure it was their call to make. The analyst who suspected the metric definition had drifted but didn&amp;rsquo;t want to slow down the dashboard delivery. The data engineer who knew the pipeline was flaky but figured someone more senior would have noticed if it really mattered.&lt;/p>
&lt;p>The broken window gets left unrepaired not because nobody sees it, but because the culture hasn&amp;rsquo;t made repair feel safe or worthwhile.&lt;/p>
&lt;p>Edmondson on this: &amp;ldquo;If there&amp;rsquo;s no bad news, remind yourself: It&amp;rsquo;s not that it&amp;rsquo;s not there. It&amp;rsquo;s that you&amp;rsquo;re not hearing about it.&amp;rdquo;&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-original-theorys-mistake--and-ours">The original theory&amp;rsquo;s mistake — and ours&lt;/h3>
&lt;/br>
&lt;p>Before we get to what to actually do about this, it&amp;rsquo;s worth pausing on where the original theory went wrong. Because data engineering is at risk of making exactly the same mistake.&lt;/p>
&lt;p>When Wilson and Kelling published their 1982 Atlantic article, the theory was nuanced. It was about signals and norms. It was explicitly a theory about community — about how residents and police could work together to maintain shared standards in public spaces. It was, in Kelling&amp;rsquo;s own framing, a theory of collective efficacy.&lt;/p>
&lt;p>What cities did with it was zero tolerance. Mass arrests for minor offences. Stop-and-frisk. 685,000 stops in New York City in a single year, more than 85% of them finding nothing at all.&lt;/p>
&lt;p>Kelling&amp;rsquo;s reaction, when he learned how comprehensively his theory had been misapplied: &amp;ldquo;Oh, shit.&amp;rdquo;&lt;/p>
&lt;p>The software world made a version of the same mistake when it imported broken windows thinking. &amp;ldquo;Don&amp;rsquo;t live with broken windows&amp;rdquo; became, for some teams, a linting rule for everything, a zero-tolerance policy for code smells, a culture of perfectionism that burned people out and produced beautiful codebases that never shipped.&lt;/p>
&lt;p>Adam Tornhill&amp;rsquo;s research at CodeScene offers the corrective: not all broken windows matter equally. In a 400,000-line codebase with 89 developers, his analysis found that 4% of the code was responsible for 72% of the defects. The broken windows that needed fixing weren&amp;rsquo;t distributed evenly. They clustered in hotspots — files that were both frequently changed and highly complex. Fix the hotspots. Let the cold, ugly, stable corner of the codebase sit.&lt;/p>
&lt;p>The principle for data engineering is the same. Zero-tolerance data quality enforcement — failing every pipeline on every test at severity error — produces fragile systems and tired teams. It&amp;rsquo;s the data equivalent of arresting everyone for jaywalking. The signal gets lost in the noise.&lt;/p>
&lt;p>What the original theory was actually pointing at was something more like: maintain the norm. Make it visible that people care. Ensure that the broken windows that matter — the ones that propagate, the ones in the high-traffic, high-trust parts of the system — get fixed promptly and publicly. Not every window. The ones that signal whether anyone&amp;rsquo;s paying attention.&lt;/p>
&lt;hr>
&lt;h3 id="what-collective-efficacy-looks-like-in-a-data-team">What collective efficacy looks like in a data team&lt;/h3>
&lt;/br>
&lt;p>Sampson, Raudenbush, and Earls — the Chicago sociologists who ran the most rigorous empirical test of broken windows theory — found that the variable that actually predicted neighbourhood safety wasn&amp;rsquo;t the presence or absence of disorder. It was &lt;em>collective efficacy&lt;/em>: the combination of social cohesion and shared willingness to intervene. Neighbours who knew each other, trusted each other, and were willing to act on behalf of each other&amp;rsquo;s wellbeing.&lt;/p>
&lt;p>The policy prescription that emerges from their work is very different from zero tolerance. It&amp;rsquo;s community investment. Shared ownership. Making the costs of non-intervention visible.&lt;/p>
&lt;p>The equivalent in data engineering starts with one thing, and if you do nothing else in this list, do this one:&lt;/p>
&lt;p>&lt;strong>Make propagation visible before anything else.&lt;/strong>&lt;/p>
&lt;p>Column-level data lineage — now available in most modern data observability platforms and increasingly in the open-source OpenLineage standard — lets you answer the question: if this column is wrong, what does it break? That question should be answerable in seconds, not hours. Teams that can visualise propagation chains respond to broken windows faster because they can see the radius of the damage before they decide whether to act.&lt;/p>
&lt;p>This matters beyond incident response. When you can show an engineer that the null column they&amp;rsquo;re tolerating in a staging model feeds seven downstream gold-layer tables, three dashboards, and a Snowflake share that two other teams consume — the calculus on whether to fix it changes. The broken window stops being an abstract code quality concern and becomes a blast radius. That&amp;rsquo;s a much more compelling argument for repair than &amp;ldquo;we should improve our data quality culture.&amp;rdquo;&lt;/p>
&lt;p>&lt;strong>Treat ownership as load-bearing infrastructure, not housekeeping.&lt;/strong>&lt;/p>
&lt;p>Every pipeline, every table, every model should have a named owner. Not a team. A person. The Jira ticket that says &amp;ldquo;Known issue — downstream teams aware&amp;rdquo; with no assigned owner is the data equivalent of a broken window with an orange cone next to it. The cone acknowledges the hazard. Nobody&amp;rsquo;s fixing the glass.&lt;/p>
&lt;p>The counterargument is always resourcing — &amp;ldquo;we don&amp;rsquo;t have time to own everything properly.&amp;rdquo; That&amp;rsquo;s true, and worth taking seriously. But the right response isn&amp;rsquo;t to pretend you own things you don&amp;rsquo;t. It&amp;rsquo;s to make orphaned assets visible and have an honest conversation about whether the organisation can afford to run production pipelines with no accountable maintainer. Most of the time, when the question is asked that directly, the answer is no.&lt;/p>
&lt;p>&lt;strong>Calibrate your tolerance patterns deliberately — and actually revisit them.&lt;/strong>&lt;/p>
&lt;p>The dbt &lt;code>severity: warn&lt;/code> setting and Airflow&amp;rsquo;s &lt;code>soft_fail&lt;/code> are legitimate tools when they&amp;rsquo;re conscious decisions: &amp;ldquo;this condition is a signal worth tracking but not a pipeline-stopper, and here&amp;rsquo;s the threshold at which that changes.&amp;rdquo; The problem is that almost nobody uses them that way. They&amp;rsquo;re the path of least resistance to avoid a broken CI build — and six months later you audit your test results and discover forty tests set to warn that haven&amp;rsquo;t been at zero failures since the day they were written.&lt;/p>
&lt;p>Treat warn-severity tests the way you treat Jira tickets that never get triaged. Set a review cadence. If a test has been consistently warning for more than two weeks without an associated investigation, it&amp;rsquo;s either a bug that needs fixing or a threshold that needs changing. &amp;ldquo;Known issue&amp;rdquo; is not a status. It&amp;rsquo;s an admission.&lt;/p>
&lt;p>&lt;strong>Fix alert fatigue before it hollows out your monitoring culture.&lt;/strong>&lt;/p>
&lt;p>The Teams channel where every data quality alert lands is the digital equivalent of a neighbourhood where every broken window gets photographed and logged and nobody ever fixes one. The log is evidence that someone noticed. It&amp;rsquo;s not evidence that anyone will act.&lt;/p>
&lt;p>Set alert thresholds that produce actionable signals. The Google SRE principle applies directly: if an alert merits a robotic response, it shouldn&amp;rsquo;t be an alert. If your first instinct on seeing a particular monitor fire is to click acknowledge and move on, that monitor is producing noise, not signal. Change it or delete it. Grouping alerts by lineage — &amp;ldquo;these five monitors fired because of one upstream schema change&amp;rdquo; rather than five separate pings — reduces volume while making propagation visible at the moment of failure, which is exactly when you want it.&lt;/p>
&lt;p>&lt;strong>Make broken windows visible to leadership, because right now the cost is invisible.&lt;/strong>&lt;/p>
&lt;p>Data quality work is famously hard to demonstrate. Pipelines that run cleanly leave no artefact. A well-maintained model with a 0% null rate looks identical to a freshly built one. There&amp;rsquo;s no natural demo format for &amp;ldquo;nothing bad happened this week.&amp;rdquo;&lt;/p>
&lt;p>This invisibility is part of why broken windows accumulate. The cost of prevention is hidden in maintenance time that doesn&amp;rsquo;t get counted. The cost of failure gets absorbed by analysts who spend their Tuesdays reconciling numbers instead of answering strategic questions, by data engineers who spend their Fridays investigating stakeholder tickets, by business decisions made on incorrect information that can&amp;rsquo;t be traced back to a specific incident.&lt;/p>
&lt;p>Quantify the propagation radius when incidents do occur. How many downstream models were affected? How many stakeholders were exposed to incorrect data, and for how long? What was the resolution effort in hours? Those numbers, tracked consistently, build a case for investment that abstract arguments about data quality never will.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-window-was-broken-before-i-noticed">The window was broken before I noticed&lt;/h3>
&lt;/br>
&lt;p>What we ended up doing was building the fix into the operating rhythm. A scheduled data drift catch-up: a full snapshot weekly, or monthly, depending on the volatility of the Salesforce object — not to respond to an incident, but to systematically correct the accumulated drift before it became visible to anyone downstream. We stopped waiting for the escalation. We made the repair part of the architecture.&lt;/p>
&lt;p>The broken window was still there, technically. The delta logic still had its blind spots around hidden relationships. But we stopped waiting for it to propagate before we addressed it. Some broken windows aren&amp;rsquo;t things you permanently seal — they&amp;rsquo;re things you build a maintenance routine around. Knowing that is the fix.&lt;/p>
&lt;p>What stayed with me wasn&amp;rsquo;t the technical solution, which was straightforward enough once we committed to it. What stayed with me was the calendar predictability of the escalation cycle before the fix. The fact that we had a known issue, a known workaround, documentation of both — and still managed to have the same downstream discovery, the same investigation, the same remediation conversation, almost exactly twelve months apart.&lt;/p>
&lt;p>The window wasn&amp;rsquo;t hidden. It was visible in the delta logs if you knew where to look. We just hadn&amp;rsquo;t made it someone&amp;rsquo;s job to look, and we hadn&amp;rsquo;t made repair part of the routine. We&amp;rsquo;d fixed the glass, filed the paperwork, and assumed the problem was solved. Until the same signal appeared in the same downstream reports, a year later.&lt;/p>
&lt;hr>
&lt;p>What I keep coming back to, though, is something that was true of that experience and is true of most of the data quality failures I&amp;rsquo;ve seen since: the window wasn&amp;rsquo;t broken secretly. It wasn&amp;rsquo;t hidden. It was visible, it was acknowledged, and it was left.&lt;/p>
&lt;p>We had the monitoring. We had the test. We had the Jira ticket.&lt;/p>
&lt;p>We had, in other words, all the infrastructure of concern. What we didn&amp;rsquo;t have was the collective agreement that repair mattered — that the propagation radius of that one broken window was large enough, and trust-eroding enough, that it justified stopping what we were doing and fixing the glass.&lt;/p>
&lt;p>That&amp;rsquo;s the thing broken windows theory is actually about, underneath all the criminology and the code smells and the schema drift. It&amp;rsquo;s about the signal that unrepaired damage sends. Not to the criminals or the developers or the data consumers. To the people who are supposed to care about the system.&lt;/p>
&lt;p>When a broken window sits long enough in a data pipeline, it stops being a problem and starts being a norm. The new engineer doesn&amp;rsquo;t flag it — they work around it. The analyst doesn&amp;rsquo;t escalate it — they add a caveat to their report. The data engineer doesn&amp;rsquo;t fix it — they document it.&lt;/p>
&lt;p>And somewhere downstream, a business decision gets made on numbers that were broken before anyone thought to check.&lt;/p>
&lt;p>The window isn&amp;rsquo;t just in the pipeline. The window is in the standard you&amp;rsquo;re willing to keep.&lt;/p></content:encoded><category>Data Engineering</category><category>Data Quality</category><category>Data Quality</category><category>Data Pipelines</category><category>Technical Debt</category><category>Data Observability</category><category>dbt</category><category>Apache Airflow</category><category>Data Culture</category><category>Pipeline Architecture</category></item><item><title>Five Worlds of Data Engineering</title><link>https://ghostinthedata.info/posts/2026/2026-05-02-five-worlds-data-engineering/</link><pubDate>Sat, 02 May 2026 09:00:00 +1000</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-05-02-five-worlds-data-engineering/</guid><author>Chris Hillman</author><description>Not all data engineering is the same. The modern analytics shop, the enterprise legacy estate, the product engine, the regulated pipeline, and the internal platform each play by different rules — and most advice only applies to one of them.</description><content:encoded>&lt;p>You watch a conference talk about implementing data contracts, and nobody mentions that the advice assumes you have multiple teams producing data — which you don&amp;rsquo;t. You read a post declaring &amp;ldquo;if you&amp;rsquo;re still using stored procedures in 2026, you&amp;rsquo;re doing it wrong,&amp;rdquo; and the comments erupt. Half the people are nodding along. Half are furious. Both sides are right. They&amp;rsquo;re just living in different worlds and don&amp;rsquo;t realise it.&lt;/p>
&lt;p>That mismatch — smart people giving each other advice that doesn&amp;rsquo;t apply — is the thing that almost never gets named. And the reason it doesn&amp;rsquo;t get named is that most public data engineering discourse is produced by and for one particular world, while pretending to speak for all of them.&lt;/p>
&lt;p>I&amp;rsquo;ve worked across a few of these worlds myself — SQL Server migrations to Teradata, then Teradata to S3, regulatory reporting under APRA and ASIC, a data mesh initiative at scale, and now a university environment that runs closer to a startup than anything I expected. The advice that kept me out of trouble in one would have gotten me fired in the other. That experience is what&amp;rsquo;s behind this taxonomy.&lt;/p>
&lt;p>I think there are five worlds here. Sometimes they intersect. Often they don&amp;rsquo;t.&lt;/p>
&lt;br>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="world-1-the-modern-analytics-shop">World 1: The Modern Analytics Shop&lt;/h3>
&lt;br>
&lt;p>This is the world most people picture when they hear &amp;ldquo;data engineering.&amp;rdquo; A team of two to ten engineers at a startup, scale-up, or mid-size company. Cloud-native from day one, or recently migrated. The stack reads like a vendor sponsor list: Snowflake or BigQuery or Databricks, dbt for transformations, Fivetran or Airbyte for ingestion, Looker or Metabase for dashboards. Everything managed. Everything SaaS. GitHub Actions wiring it together.&lt;/p>
&lt;p>What matters here is speed. Getting value to stakeholders fast. Shipping a working dashboard before the quarterly review. Iterating on models without a change advisory board. The team is small enough that everyone knows the codebase, and requirements come from a Teams message, not a 40-page specification document.&lt;/p>
&lt;p>I find myself in this world more than I expected right now. The current role is an institution that by any measure is large and complex — but the data team is small, the autonomy is real, and the energy genuinely feels like a startup. New tools, fast decisions, the ability to actually make change. It&amp;rsquo;s a reminder that World 1 isn&amp;rsquo;t only about company size. It&amp;rsquo;s about how the team operates.&lt;/p>
&lt;p>This is also a great world to work in, especially early in your career. The tooling is mature, the feedback loops are tight, and the problems are tractable.&lt;/p>
&lt;p>Here&amp;rsquo;s the thing, though. This is also the world that roughly 80% of public data engineering content is written for and about. Not because it&amp;rsquo;s the most common world — it isn&amp;rsquo;t — but because it&amp;rsquo;s the most fundable. The vendor-funded conference talks, the sponsored podcasts, the developer advocate blog posts, the LinkedIn hot takes — they overwhelmingly reflect this world. Developer advocates write about the stack their employer sells (this is not a dig — it&amp;rsquo;s the job). Conference sponsors want talks that showcase their tools in the most flattering light. The technical depth suffers because the incentive is awareness, not education.&lt;/p>
&lt;p>None of that is wrong. But it creates a gravitational pull that distorts the whole discourse. When someone writes &amp;ldquo;the right way to do data engineering,&amp;rdquo; they almost always mean &lt;em>this&lt;/em> world. And if you&amp;rsquo;re in a different one and don&amp;rsquo;t recognise the mismatch, you end up feeling like you&amp;rsquo;re doing it wrong when you&amp;rsquo;re actually just solving a different problem.&lt;/p>
&lt;p>&lt;strong>Advice that works here but rarely travels:&lt;/strong> &amp;ldquo;Just use dbt.&amp;rdquo; &amp;ldquo;Schema-on-read is fine for now.&amp;rdquo; &amp;ldquo;You don&amp;rsquo;t need a data catalog yet.&amp;rdquo; &amp;ldquo;Start with a &lt;a href="https://ghostinthedata.info/posts/2025/2025-11-07-effective-data-modelling/" target="_blank" rel="noopener">star schema&lt;/a> and iterate.&amp;rdquo;&lt;/p>
&lt;br>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="world-2-the-enterprise-legacy-estate">World 2: The Enterprise Legacy Estate&lt;/h3>
&lt;br>
&lt;p>This is the world of large organisations — banks, insurers, manufacturers, utilities, healthcare systems, government agencies — with data infrastructure that predates the cloud. Often predates the people currently maintaining it. Teams of twenty to two hundred data professionals spread across business units that don&amp;rsquo;t always talk to each other.&lt;/p>
&lt;p>The stack tells a different story: Informatica, SSIS, Ab Initio, Teradata, Oracle, maybe an on-prem Hadoop cluster that someone championed in 2014 and nobody&amp;rsquo;s had the political capital to decommission. Perhaps some Snowflake or Databricks grafted on top, creating a hybrid that&amp;rsquo;s more complex than either system alone.&lt;/p>
&lt;p>I lived this migration arc twice — once from SQL Server to Teradata, and again from Teradata to S3 and Starburst. Each time, the temptation was to treat the old system as a problem to escape rather than a body of knowledge to decode. Each time, the engineers who did the most damage were the ones who arrived with the new stack and immediately started designing around the old one rather than understanding it first.&lt;/p>
&lt;p>What matters here is stability. Not breaking things. Migration plans that span years, not sprints. Governance and lineage aren&amp;rsquo;t aspirational — they&amp;rsquo;re audit requirements. And politics, because the data warehouse your predecessor built in 2011 is somebody&amp;rsquo;s empire (you know the one), and you can&amp;rsquo;t &amp;ldquo;just replace it&amp;rdquo; without navigating a web of organisational power dynamics that no architecture diagram captures.&lt;/p>
&lt;p>This is the world where &lt;a href="https://ghostinthedata.info/posts/2026/2026-03-09-data-model-not-broken-part-1/" target="_blank" rel="noopener">refactoring beats rebuilding&lt;/a> every time. That fact table with 200 columns? Those bridge tables nobody understands? The slowly-changing-dimension-within-a-slowly-changing-dimension? They&amp;rsquo;re not bugs — they&amp;rsquo;re reality encoded. Every weird modelling choice represents a business rule someone fought to understand. The &amp;ldquo;clean&amp;rdquo; data vault remodel will eventually end up with the same complexity, just distributed across more tables with more confusing names.&lt;/p>
&lt;p>Advice from World 1 can be actively dangerous here. &amp;ldquo;Just rewrite it&amp;rdquo; destroys institutional memory. &amp;ldquo;Adopt a lakehouse architecture&amp;rdquo; sounds great until you realise you have 4,000 stored procedures that encode fifteen years of business rules, and nobody documented them. The engineer in this world isn&amp;rsquo;t slow because they&amp;rsquo;re behind the curve. They&amp;rsquo;re careful because the cost of breaking something is measured in regulatory findings and executive phone calls, not a failed CI check.&lt;/p>
&lt;p>&lt;strong>Advice that works here but rarely travels:&lt;/strong> &amp;ldquo;Document everything before you touch it.&amp;rdquo; &amp;ldquo;Strangler fig, never big bang.&amp;rdquo; &amp;ldquo;Spend more time understanding &lt;em>why&lt;/em> it was built this way than planning what to replace it with.&amp;rdquo; &amp;ldquo;The weird WHERE clause isn&amp;rsquo;t a bug — it&amp;rsquo;s institutional memory.&amp;rdquo;&lt;/p>
&lt;br>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="world-3-the-product-engine">World 3: The Product Engine&lt;/h3>
&lt;br>
&lt;p>This is the world where data engineering directly powers an outcome the business delivers. That might be real-time personalisation, recommendation engines, or pricing algorithms. But it also includes building data products that drive campaigns, marketing attribution, and business decisions — cases where the engineer&amp;rsquo;s output flows into something customers or commercial teams act on directly, rather than landing in a dashboard that an analyst reviews on Monday morning.&lt;/p>
&lt;p>The real-time variant of this world has its own stack: Kafka or Kinesis for streaming, Spark or Flink for processing, feature stores for serving machine learning models. Often custom infrastructure because off-the-shelf tools can&amp;rsquo;t meet the latency requirements.&lt;/p>
&lt;p>But the defining characteristic isn&amp;rsquo;t milliseconds — it&amp;rsquo;s consequence. If a pipeline breaks in this world, someone notices immediately. A campaign fires with the wrong audience. A pricing decision gets made on stale data. A recommendation engine serves the same product to everyone because the features stopped updating overnight. The engineer is accountable to an outcome, not just a pipeline.&lt;/p>
&lt;p>I&amp;rsquo;ve worked in this space building data products that powered business campaigns and marketing decisions. Not real-time serving in the p99 latency sense, but consequential enough that a broken pipeline meant a broken business process. That accountability changes how you think about testing, monitoring, and what &amp;ldquo;done&amp;rdquo; actually means.&lt;/p>
&lt;p>The skills that matter here — operational thinking, system design, understanding how your data is consumed downstream — overlap more with software engineering and product thinking than with SQL and dbt. When someone says &amp;ldquo;data engineering is just SQL and orchestration,&amp;rdquo; someone in this world quietly closes the tab.&lt;/p>
&lt;p>&lt;strong>Advice that works here but rarely travels:&lt;/strong> &amp;ldquo;Treat pipelines like production services.&amp;rdquo; &amp;ldquo;Your tests need to run in CI, not in a notebook.&amp;rdquo; &amp;ldquo;Schema evolution is a deployment problem, not a modelling problem.&amp;rdquo; &amp;ldquo;The consumer of your data is your customer — know what breaks their day.&amp;rdquo;&lt;/p>
&lt;br>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="world-4-the-regulated-pipeline">World 4: The Regulated Pipeline&lt;/h3>
&lt;br>
&lt;p>This is the world defined by compliance. In Australia: APRA for prudential standards, ASIC for financial services conduct, AFCA for dispute resolution obligations, and the Privacy Act for anything touching personal information. Internationally: HIPAA, SOX, MiFID, GDPR. The defining characteristic isn&amp;rsquo;t the technology — it&amp;rsquo;s that regulatory requirements shape every architectural decision before the first line of code is written.&lt;/p>
&lt;p>The stack is whatever passed the security review. Often years behind the cutting edge because new tools need months of compliance evaluation before they&amp;rsquo;re approved. Data lineage tools aren&amp;rsquo;t nice-to-haves — they&amp;rsquo;re audit requirements. Encryption isn&amp;rsquo;t a best practice — it&amp;rsquo;s a legal mandate. Access controls aren&amp;rsquo;t &amp;ldquo;we should get around to that&amp;rdquo; — they&amp;rsquo;re the first thing you build.&lt;/p>
&lt;p>What matters is auditability. Can you answer &amp;ldquo;who accessed what, when, and why&amp;rdquo; for any record in the system? Can you prove data lineage from source to report? Can you demonstrate that a deletion request was honoured within the legally mandated timeframe? Your data deletion pipeline is as important as your ingestion pipeline, and it needs the same rigour.&lt;/p>
&lt;p>I spent significant time in this world — financial services, where regulatory reporting wasn&amp;rsquo;t a background concern but a core deliverable. APRA and ASIC don&amp;rsquo;t ask nicely. The audit wasn&amp;rsquo;t a hypothetical and the regulator&amp;rsquo;s question list arrived without warning. The engineers I worked alongside weren&amp;rsquo;t slow because they lacked ambition. They were deliberate because the cost of getting it wrong wasn&amp;rsquo;t a postmortem — it was a regulatory finding, a remediation programme, and occasionally a front-page story.&lt;/p>
&lt;p>The trap for engineers entering this world from World 1 is assuming that governance is bureaucracy. It isn&amp;rsquo;t. Governance &lt;em>is&lt;/em> architecture. The compliance requirements aren&amp;rsquo;t obstacles to good engineering — they&amp;rsquo;re constraints that shape what good engineering looks like. The best engineers I&amp;rsquo;ve worked with in regulated environments don&amp;rsquo;t fight the constraints. They design systems where compliance is a property of the architecture itself, not a layer bolted on top.&lt;/p>
&lt;p>&lt;strong>Advice that works here but rarely travels:&lt;/strong> &amp;ldquo;Governance is architecture, not bureaucracy.&amp;rdquo; &amp;ldquo;If you can&amp;rsquo;t prove lineage, you can&amp;rsquo;t ship it.&amp;rdquo; &amp;ldquo;The security review IS the sprint.&amp;rdquo; &amp;ldquo;Your data deletion pipeline is as important as your data ingestion pipeline.&amp;rdquo;&lt;/p>
&lt;br>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="world-5-the-internal-platform">World 5: The Internal Platform&lt;/h3>
&lt;br>
&lt;p>This is the world of organisations large enough that the data team has split into &amp;ldquo;platform&amp;rdquo; and &amp;ldquo;domain&amp;rdquo; functions. The platform team builds the infrastructure, tooling, and self-service capabilities that other data teams consume. Data mesh adopters. Companies with fifty or more data practitioners who realised that a centralised team can&amp;rsquo;t scale to serve every domain&amp;rsquo;s needs.&lt;/p>
&lt;p>The stack is about abstraction and enablement: Kubernetes, self-hosted Airflow or Dagster or Prefect, internal developer portals, data contracts, schema registries, internal PyPI packages. Often custom abstractions layered on top of cloud services, designed to give domain teams guardrails without bottlenecks.&lt;/p>
&lt;p>I experienced this world through ANZx — ANZ&amp;rsquo;s digital platform initiative — where data mesh concepts were applied at scale. Not as a theoretical framework on a conference slide, but as a real attempt to distribute data ownership across domains while maintaining coherence at the platform level. The challenge wasn&amp;rsquo;t the technology. It was getting domain teams to think like data product owners rather than data consumers. That shift is harder than any infrastructure problem.&lt;/p>
&lt;p>What matters here is adoption. Not how many pipelines &lt;em>you&lt;/em> build, but how many pipelines your consumers build without needing your help. Your success metric isn&amp;rsquo;t &amp;ldquo;pipelines shipped&amp;rdquo; — it&amp;rsquo;s &amp;ldquo;time to first pipeline for a new domain team.&amp;rdquo; Developer experience for internal consumers is your product, and if the experience is poor, your consumers will route around you. They&amp;rsquo;ll spin up their own Snowflake account, write their own ingestion scripts, and create exactly the kind of ungoverned sprawl your platform was supposed to prevent.&lt;/p>
&lt;p>If adopting your platform requires a Jira ticket and a two-week wait, you&amp;rsquo;ve already lost. The shadow pipelines will multiply, and nobody will tell you until the audit.&lt;/p>
&lt;p>&lt;strong>Advice that works here but rarely travels:&lt;/strong> &amp;ldquo;Treat internal teams as customers with choices.&amp;rdquo; &amp;ldquo;Self-service is the goal, not centralised delivery.&amp;rdquo; &amp;ldquo;Your documentation IS the product.&amp;rdquo; &amp;ldquo;Measure adoption, not output.&amp;rdquo;&lt;/p>
&lt;br>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="so-what">So What?&lt;/h3>
&lt;br>
&lt;p>Most things in data engineering are the same no matter which world you&amp;rsquo;re in. Data quality matters everywhere. Testing matters everywhere. Documentation matters everywhere. Understanding the business context behind the data — that&amp;rsquo;s universal, and I&amp;rsquo;d argue it&amp;rsquo;s the &lt;a href="https://ghostinthedata.info/posts/2025/2025-02-08-business-context-guide/" target="_blank" rel="noopener">most undervalued skill&lt;/a> in the profession.&lt;/p>
&lt;p>But not everything transfers. And when somebody tells you about methodology — at a conference, on LinkedIn, in a blog post (including mine) — it&amp;rsquo;s worth thinking about which world they&amp;rsquo;re coming from.&lt;/p>
&lt;p>Here&amp;rsquo;s something I&amp;rsquo;ve noticed when hiring vetran data engineers: the ones who stand out aren&amp;rsquo;t necessarily the deepest specialists in any one world. They&amp;rsquo;re the ones who&amp;rsquo;ve visited enough worlds to know the laws. They don&amp;rsquo;t have to have lived in each one for years — but they&amp;rsquo;ve spent enough time in the regulated environment to understand why governance isn&amp;rsquo;t bureaucracy. Enough time in the enterprise to know that &amp;ldquo;just rewrite it&amp;rdquo; destroys things. Enough time in the product engine to feel what accountability to an outcome actually means. That layered experience is what hardens an engineer. It&amp;rsquo;s what lets them walk into an unfamiliar environment and read it quickly — recognising the constraints, the trade-offs, the things that will matter before anyone tells them.&lt;/p>
&lt;p>There&amp;rsquo;s also something worth saying about what these worlds are not: interchangeable, or combinable. I&amp;rsquo;ve never seen an organisation that genuinely encompasses all five. You occasionally see attempts — usually during a data mesh initiative — and they tend to produce something more complex than any individual world, with the clarity of none of them. The worlds don&amp;rsquo;t collapse into a super-universe. They coexist, imperfectly, inside large organisations. A regulated pipeline team and a modern analytics shop can operate within the same company and have almost nothing in common in terms of how they work, what they value, and what good looks like.&lt;/p>
&lt;p>The tradeoff is real and worth naming plainly: the content that&amp;rsquo;s most abundant — the conference talks, the vendor posts, the LinkedIn hot takes — is calibrated for a world that isn&amp;rsquo;t yours if you&amp;rsquo;re in World 2, 3, 4, or 5. There&amp;rsquo;s more of it than ever, and it&amp;rsquo;s easier than ever to access. What you give up is the ability to consume it passively. Filtering for relevance is now part of the job, and nobody puts that in the job description.&lt;/p>
&lt;p>When you read advice — any advice, including everything on &lt;a href="https://ghostinthedata.info/posts/" target="_blank" rel="noopener">this blog&lt;/a> — ask yourself which world it&amp;rsquo;s coming from. If it doesn&amp;rsquo;t apply to yours, that&amp;rsquo;s not a failing on your part or theirs. It just means they&amp;rsquo;re in a different world.&lt;/p>
&lt;p>Now you know to notice.&lt;/p>
&lt;br></content:encoded><category>Data Engineering</category><category>Career Development</category><category>Data Engineering</category><category>Modern Data Stack</category><category>Enterprise</category><category>Data Architecture</category><category>Career Development</category><category>Leadership</category></item><item><title>Your Data Platform Costs More Than It Should</title><link>https://ghostinthedata.info/posts/2026/2026-04-25-cost-management/</link><pubDate>Sat, 25 Apr 2026 09:00:00 +1100</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-04-25-cost-management/</guid><author>Chris Hillman</author><description>A practical guide to understanding, measuring, and reducing your Snowflake and AWS data platform costs — starting with the habits that actually move the needle.</description><content:encoded>&lt;p>Let me tell you about the moment I stopped treating cloud costs as someone else&amp;rsquo;s problem.&lt;/p>
&lt;p>We were three months into a Snowflake migration. Everything was humming. Pipelines were green, dashboards were fast, the analytics team was happier than I&amp;rsquo;d seen them before. I felt good about the work we&amp;rsquo;d done.&lt;/p>
&lt;p>Then finance forwarded me the invoice.&lt;/p>
&lt;p>The number wasn&amp;rsquo;t catastrophic. But it was significantly higher than what we&amp;rsquo;d budgeted, and when I started digging, I couldn&amp;rsquo;t explain where most of it was going. I knew we had warehouses running. I knew we had pipelines executing. But I couldn&amp;rsquo;t tell you which warehouse was responsible for what cost, which pipelines were the expensive ones, or whether the money was well spent. I had built a platform I was proud of — and I had no idea what it actually cost to operate.&lt;/p>
&lt;p>That&amp;rsquo;s the moment that changed how I think about data engineering. Not because of the dollar amount, but because of the realisation underneath it: &lt;strong>I had built something I couldn&amp;rsquo;t explain to the people paying for it.&lt;/strong> And if I couldn&amp;rsquo;t explain it, I couldn&amp;rsquo;t defend it. And if I couldn&amp;rsquo;t defend it, someone else would make the decisions for me — someone who didn&amp;rsquo;t understand why the platform mattered.&lt;/p>
&lt;br>
&lt;p>I&amp;rsquo;m telling you this because cost management is one of those things that sounds like a finance problem until you experience the consequences firsthand. It&amp;rsquo;s not about being cheap. It&amp;rsquo;s about being intentional. It&amp;rsquo;s about knowing that every credit you spend is buying something valuable — and being able to prove it when someone asks.&lt;/p>
&lt;p>The data engineers who understand their costs don&amp;rsquo;t just save money. They earn trust. They get budget for the projects that matter. They sleep better because they&amp;rsquo;ve eliminated the waste that eventually becomes someone else&amp;rsquo;s excuse to cut headcount or freeze hiring.&lt;/p>
&lt;p>This article is about building that understanding. Not with a vendor&amp;rsquo;s optimisation tool or a consultant&amp;rsquo;s audit — but with the habits, queries, and mental models that let you own your platform&amp;rsquo;s economics from the inside. Everything here is grounded in Snowflake and AWS, with specific code you can run today.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="you-cant-optimise-what-you-cant-see">You can&amp;rsquo;t optimise what you can&amp;rsquo;t see&lt;/h3>
&lt;br>
&lt;p>Before you touch a single warehouse configuration, you need to answer one question: &lt;strong>where is the money going?&lt;/strong>&lt;/p>
&lt;p>Most teams skip this step. They read a blog post about auto-suspend settings, change a few defaults, and call it optimisation. That&amp;rsquo;s like going on a diet by switching to diet soda while eating three pizzas a day. The soda wasn&amp;rsquo;t the problem.&lt;/p>
&lt;p>Here&amp;rsquo;s the query I run first on every Snowflake environment I touch. It tells you which warehouses are consuming the most credits over the last 30 days:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used) &lt;span style="color:#66d9ef">AS&lt;/span> total_credits,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used) &lt;span style="color:#f92672">*&lt;/span> &lt;span style="color:#ae81ff">3&lt;/span>.&lt;span style="color:#ae81ff">00&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> estimated_cost_usd, &lt;span style="color:#75715e">-- adjust your credit price
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(&lt;span style="color:#66d9ef">DISTINCT&lt;/span> DATE_TRUNC(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, start_time)) &lt;span style="color:#66d9ef">AS&lt;/span> active_days,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ROUND(&lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(&lt;span style="color:#66d9ef">DISTINCT&lt;/span> DATE_TRUNC(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, start_time)), &lt;span style="color:#ae81ff">2&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> credits_per_day
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span> snowflake.account_usage.warehouse_metering_history
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span> start_time &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">30&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span> warehouse_name
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span> total_credits &lt;span style="color:#66d9ef">DESC&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Run that. Right now. I&amp;rsquo;ll wait.&lt;/p>
&lt;p>If you&amp;rsquo;re like most teams, 70–80% of your credits come from two or three warehouses. That&amp;rsquo;s your starting point. Not everything — just the expensive stuff.&lt;/p>
&lt;br>
&lt;p>Now do the same thing for queries. This one finds your top 20 most expensive queries by bytes scanned:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> query_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> query_text,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> user_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> total_elapsed_time &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#ae81ff">1000&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> elapsed_seconds,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> bytes_scanned &lt;span style="color:#f92672">/&lt;/span> POWER(&lt;span style="color:#ae81ff">1024&lt;/span>, &lt;span style="color:#ae81ff">3&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> gb_scanned,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> partitions_scanned,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> partitions_total,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ROUND(partitions_scanned &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#66d9ef">NULLIF&lt;/span>(partitions_total, &lt;span style="color:#ae81ff">0&lt;/span>) &lt;span style="color:#f92672">*&lt;/span> &lt;span style="color:#ae81ff">100&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> pct_partitions_scanned
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.query_history
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> start_time &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">7&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> bytes_scanned &lt;span style="color:#f92672">&amp;gt;&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> bytes_scanned &lt;span style="color:#66d9ef">DESC&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">LIMIT&lt;/span> &lt;span style="color:#ae81ff">20&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Pay attention to that &lt;code>pct_partitions_scanned&lt;/code> column. If you see queries scanning 90–100% of a table&amp;rsquo;s partitions, those queries aren&amp;rsquo;t benefiting from clustering or partition pruning. That&amp;rsquo;s where the big wins hide.&lt;/p>
&lt;p>This 15-minute exercise — top warehouses, top queries — tells you more about your cost profile than any dashboard. It&amp;rsquo;s the equivalent of checking your bank statement before creating a budget. Obvious in hindsight. Almost nobody does it.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-idle-tax-is-your-single-biggest-waste">The idle tax is your single biggest waste&lt;/h3>
&lt;br>
&lt;p>Here&amp;rsquo;s what nobody tells you about Snowflake billing: you pay for compute in 60-second minimums. Every time a warehouse resumes from suspension, you&amp;rsquo;re billed for at least one minute — regardless of whether the query takes 2 seconds or 58 seconds.&lt;/p>
&lt;p>This matters more than it sounds. Picture a BI tool like Metabase or Tableau hitting your warehouse with 20 small metadata queries over 15 minutes. If your warehouse auto-suspends after 5 minutes (Snowflake&amp;rsquo;s default for many setups), it might suspend and resume multiple times during that window. Each resume triggers another 60-second charge.&lt;/p>
&lt;p>Twenty queries that take 3 seconds each? That&amp;rsquo;s 60 seconds of actual compute. But if the warehouse suspends and resumes 4 times, you&amp;rsquo;re billed for 240 seconds. A 4x overhead.&lt;/p>
&lt;p>The fix is straightforward but requires thought:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- For transformation warehouses (predictable, bursty workloads)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">ALTER&lt;/span> WAREHOUSE transform_wh
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SET&lt;/span> AUTO_SUSPEND &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">60&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> AUTO_RESUME &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">TRUE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> WAREHOUSE_SIZE &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;MEDIUM&amp;#39;&lt;/span>;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- For BI/dashboard warehouses (frequent small queries)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">ALTER&lt;/span> WAREHOUSE bi_wh
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SET&lt;/span> AUTO_SUSPEND &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">300&lt;/span> &lt;span style="color:#75715e">-- 5 min keeps it warm between dashboard interactions
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> AUTO_RESUME &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">TRUE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> WAREHOUSE_SIZE &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;SMALL&amp;#39;&lt;/span>;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- For ad-hoc/analyst warehouses (unpredictable, intermittent)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">ALTER&lt;/span> WAREHOUSE adhoc_wh
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SET&lt;/span> AUTO_SUSPEND &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">60&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> AUTO_RESUME &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">TRUE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> WAREHOUSE_SIZE &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;XSMALL&amp;#39;&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The principle: &lt;strong>transformation warehouses should suspend aggressively&lt;/strong> because they run in defined windows with gaps between runs. BI warehouses should stay warm a bit longer because dashboard users generate clusters of queries with short pauses in between. Ad-hoc warehouses should be small and aggressive — analysts can tolerate a 1–2 second resume delay.&lt;/p>
&lt;p>Here&amp;rsquo;s how to find warehouses that are running but idle — the silent cost killer:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used) &lt;span style="color:#66d9ef">AS&lt;/span> total_credits,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used_compute) &lt;span style="color:#66d9ef">AS&lt;/span> compute_credits,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used_cloud_services) &lt;span style="color:#66d9ef">AS&lt;/span> cloud_credits,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ROUND(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> (&lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used) &lt;span style="color:#f92672">-&lt;/span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used_compute))
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#66d9ef">NULLIF&lt;/span>(&lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used), &lt;span style="color:#ae81ff">0&lt;/span>) &lt;span style="color:#f92672">*&lt;/span> &lt;span style="color:#ae81ff">100&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ) &lt;span style="color:#66d9ef">AS&lt;/span> pct_idle_cost
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.warehouse_metering_history
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> start_time &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">30&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse_name
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">HAVING&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> pct_idle_cost &lt;span style="color:#f92672">&amp;gt;&lt;/span> &lt;span style="color:#ae81ff">20&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> total_credits &lt;span style="color:#66d9ef">DESC&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>If &lt;code>pct_idle_cost&lt;/code> is above 20% on any warehouse, you&amp;rsquo;re burning money on idle time. Tighten the auto-suspend, or investigate what&amp;rsquo;s keeping the warehouse awake between queries.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-most-expensive-query-is-the-one-that-scans-everything">The most expensive query is the one that scans everything&lt;/h3>
&lt;br>
&lt;p>Snowflake charges you for the compute time your queries consume, and nothing drives compute time like full table scans. The two most common culprits are queries that wrap filter columns in functions, and queries that select more columns than they need.&lt;/p>
&lt;p>Here&amp;rsquo;s what I mean. This query looks innocent:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Expensive: function on the filter column disables pruning
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_id, customer_id, total_amount
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> analytics.fct_orders
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> DATE(order_timestamp) &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;2026-03-01&amp;#39;&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>But it&amp;rsquo;s quietly terrible. Wrapping &lt;code>order_timestamp&lt;/code> in &lt;code>DATE()&lt;/code> forces Snowflake to evaluate every row before filtering. The query planner can&amp;rsquo;t use micro-partition metadata to skip irrelevant partitions.&lt;/p>
&lt;p>This version does the same thing, but lets Snowflake prune:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Cheap: range filter on raw column enables pruning
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_id, customer_id, total_amount
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> analytics.fct_orders
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_timestamp &lt;span style="color:#f92672">&amp;gt;=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;2026-03-01&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> order_timestamp &lt;span style="color:#f92672">&amp;lt;&lt;/span> &lt;span style="color:#e6db74">&amp;#39;2026-03-02&amp;#39;&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The difference can be enormous on large tables — I&amp;rsquo;ve seen this single change reduce bytes scanned by 95% on billion-row fact tables.&lt;/p>
&lt;p>The other silent killer is &lt;code>SELECT *&lt;/code> in intermediate transformations. In a dbt project, I often find staging models that look like this:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- stg_orders.sql (before)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span> &lt;span style="color:#f92672">*&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">source&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;raw&amp;#39;&lt;/span>, &lt;span style="color:#e6db74">&amp;#39;orders&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> _loaded_at &lt;span style="color:#f92672">&amp;gt;&lt;/span> (&lt;span style="color:#66d9ef">SELECT&lt;/span> &lt;span style="color:#66d9ef">MAX&lt;/span>(_loaded_at) &lt;span style="color:#66d9ef">FROM&lt;/span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> this &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>That &lt;code>SELECT *&lt;/code> pulls every column from the source — including columns nobody downstream ever uses. In a columnar store like Snowflake, you only pay to scan the columns you reference. Trimming to the columns you actually need is free performance:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- stg_orders.sql (after)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_timestamp,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> total_amount,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> status,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> _loaded_at
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">source&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;raw&amp;#39;&lt;/span>, &lt;span style="color:#e6db74">&amp;#39;orders&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> _loaded_at &lt;span style="color:#f92672">&amp;gt;&lt;/span> (&lt;span style="color:#66d9ef">SELECT&lt;/span> &lt;span style="color:#66d9ef">MAX&lt;/span>(_loaded_at) &lt;span style="color:#66d9ef">FROM&lt;/span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> this &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>This isn&amp;rsquo;t premature optimisation. It&amp;rsquo;s hygiene. Every &lt;code>SELECT *&lt;/code> in your transformation layer is a small tax you pay on every single run.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="full-refreshes-are-the-most-expensive-default-in-data-engineering">Full refreshes are the most expensive default in data engineering&lt;/h3>
&lt;br>
&lt;p>If you&amp;rsquo;re using dbt — and most teams on Snowflake are — the default materialisation is either &lt;code>view&lt;/code> or &lt;code>table&lt;/code>. Both have the same problem at scale: they rebuild everything, every time.&lt;/p>
&lt;p>A table materialisation on a 500-million-row fact table means Snowflake reads, transforms, and writes 500 million rows every run. Even if only 50,000 rows changed since yesterday. That&amp;rsquo;s a 10,000x overhead.&lt;/p>
&lt;p>Switching to incremental materialisation is the single highest-impact cost change most teams can make:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- fct_orders.sql
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> config(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> materialized&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;incremental&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> unique_key&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;order_id&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> incremental_strategy&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;merge&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> on_schema_change&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;append_new_columns&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> )
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> order_timestamp,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> total_amount,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> status,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> updated_at
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">ref&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;stg_orders&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#960050;background-color:#1e0010">{&lt;/span>&lt;span style="color:#f92672">%&lt;/span> &lt;span style="color:#66d9ef">if&lt;/span> is_incremental() &lt;span style="color:#f92672">%&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span> updated_at &lt;span style="color:#f92672">&amp;gt;&lt;/span> (
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SELECT&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;hour&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">3&lt;/span>, &lt;span style="color:#66d9ef">MAX&lt;/span>(updated_at))
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">FROM&lt;/span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> this &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#960050;background-color:#1e0010">{&lt;/span>&lt;span style="color:#f92672">%&lt;/span> endif &lt;span style="color:#f92672">%&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>A few things worth noting here. The &lt;code>DATEADD('hour', -3, ...)&lt;/code> creates a 3-hour lookback window. This catches late-arriving data and handles clock skew between source systems. Without it, you&amp;rsquo;ll miss rows that arrive slightly out of order — and your numbers will silently drift.&lt;/p>
&lt;p>The &lt;code>unique_key&lt;/code> with &lt;code>incremental_strategy='merge'&lt;/code> means Snowflake will update existing rows and insert new ones. This is essential for tables where source records get modified after initial load (order status changes, for example).&lt;/p>
&lt;p>The rule of thumb: &lt;strong>if a table has more than a million rows and less than 20% changes per run, make it incremental.&lt;/strong> The cost savings are usually 80–95% on that model&amp;rsquo;s compute.&lt;/p>
&lt;p>But — and this is important — schedule a periodic full refresh to correct any drift. I typically set up a weekly &lt;code>dbt build --full-refresh --select fct_orders&lt;/code> via a separate Airflow DAG or GitHub Actions workflow. Belt and suspenders.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-hidden-costs-your-snowflake-dashboard-wont-show-you">The hidden costs your Snowflake dashboard won&amp;rsquo;t show you&lt;/h3>
&lt;br>
&lt;p>Most teams monitor warehouse credits because that&amp;rsquo;s what the Snowflake UI makes visible. But there are cost vectors that don&amp;rsquo;t show up in the obvious places.&lt;/p>
&lt;p>&lt;strong>Cloud services credits&lt;/strong> accrue when queries use Snowflake&amp;rsquo;s coordination layer — query compilation, metadata operations, result set caching. Normally this is covered by a 10% &amp;ldquo;free&amp;rdquo; adjustment against your compute credits. But if you have BI tools making thousands of small metadata queries, cloud services can exceed that adjustment and start costing real money.&lt;/p>
&lt;p>Find them:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> DATE_TRUNC(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, usage_date) &lt;span style="color:#66d9ef">AS&lt;/span> &lt;span style="color:#66d9ef">day&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used) &lt;span style="color:#66d9ef">AS&lt;/span> total_credits,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_adjustment_cloud_services) &lt;span style="color:#66d9ef">AS&lt;/span> cloud_services_adjustment,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">CASE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHEN&lt;/span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_adjustment_cloud_services) &lt;span style="color:#f92672">&amp;lt;&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">THEN&lt;/span> &lt;span style="color:#66d9ef">ABS&lt;/span>(&lt;span style="color:#66d9ef">SUM&lt;/span>(credits_adjustment_cloud_services))
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">ELSE&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">END&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> excess_cloud_services_cost
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.metering_daily_history
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> usage_date &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">30&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">day&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">day&lt;/span> &lt;span style="color:#66d9ef">DESC&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>&lt;strong>Serverless feature credits&lt;/strong> — Snowpipe, automatic clustering, materialized view maintenance, search optimisation — all consume credits outside your warehouse billing. They don&amp;rsquo;t show up in &lt;code>warehouse_metering_history&lt;/code>. You need to check separately:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Snowpipe costs
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> pipe_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used) &lt;span style="color:#66d9ef">AS&lt;/span> total_credits
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.pipe_usage_history
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> start_time &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">30&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> pipe_name
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> total_credits &lt;span style="color:#66d9ef">DESC&lt;/span>;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Automatic clustering costs
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">table_name&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used) &lt;span style="color:#66d9ef">AS&lt;/span> total_credits
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.automatic_clustering_history
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> start_time &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">30&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">table_name&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> total_credits &lt;span style="color:#66d9ef">DESC&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>On the AWS side, the sneaky costs tend to be data transfer. Moving data between S3 regions, or from S3 out to the internet, adds up quietly. If your Snowflake account is in &lt;code>us-east-1&lt;/code> but your S3 landing zone is in &lt;code>ap-southeast-2&lt;/code>, every byte of ingestion carries a cross-region transfer charge. Check your AWS Cost Explorer with the &amp;ldquo;Data Transfer&amp;rdquo; service filter — the numbers are often surprising.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="tag-everything-or-youll-optimise-blind">Tag everything or you&amp;rsquo;ll optimise blind&lt;/h3>
&lt;br>
&lt;p>You cannot reduce costs you can&amp;rsquo;t attribute. The simplest and most underused tool in Snowflake is the query tag — a piece of metadata you attach to every query that tells you &lt;em>what&lt;/em> generated it.&lt;/p>
&lt;p>If you&amp;rsquo;re using dbt, this takes about five minutes to set up. Create a macro:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- macros/set_query_tag.sql
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">{&lt;/span>&lt;span style="color:#f92672">%&lt;/span> macro set_query_tag() &lt;span style="color:#f92672">%&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">{&lt;/span>&lt;span style="color:#f92672">%&lt;/span> &lt;span style="color:#66d9ef">set&lt;/span> query_tag &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#960050;background-color:#1e0010">{&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#34;dbt_model&amp;#34;&lt;/span>: model.name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#34;dbt_schema&amp;#34;&lt;/span>: model.&lt;span style="color:#66d9ef">schema&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#34;dbt_materialized&amp;#34;&lt;/span>: model.config.materialized,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#34;dbt_invocation_id&amp;#34;&lt;/span>: invocation_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#34;environment&amp;#34;&lt;/span>: target.name
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">}&lt;/span> &lt;span style="color:#f92672">%&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">{&lt;/span>&lt;span style="color:#f92672">%&lt;/span> &lt;span style="color:#66d9ef">do&lt;/span> run_query(&lt;span style="color:#e6db74">&amp;#34;ALTER SESSION SET QUERY_TAG = &amp;#39;{}&amp;#39;&amp;#34;&lt;/span>.format(query_tag &lt;span style="color:#f92672">|&lt;/span> tojson)) &lt;span style="color:#f92672">%&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#960050;background-color:#1e0010">{&lt;/span>&lt;span style="color:#f92672">%&lt;/span> endmacro &lt;span style="color:#f92672">%&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Then add it as a pre-hook in your &lt;code>dbt_project.yml&lt;/code>:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># dbt_project.yml&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">models&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">your_project&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">+pre-hook&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;{{ set_query_tag() }}&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Now every query dbt runs is tagged with the model name, materialisation type, and environment. You can query cost by model:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> PARSE_JSON(query_tag):dbt_model::STRING &lt;span style="color:#66d9ef">AS&lt;/span> model_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> PARSE_JSON(query_tag):dbt_materialized::STRING &lt;span style="color:#66d9ef">AS&lt;/span> materialization,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(&lt;span style="color:#f92672">*&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> query_count,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(total_elapsed_time) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#ae81ff">1000&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> total_seconds,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(bytes_scanned) &lt;span style="color:#f92672">/&lt;/span> POWER(&lt;span style="color:#ae81ff">1024&lt;/span>, &lt;span style="color:#ae81ff">4&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> tb_scanned
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.query_history
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> query_tag &lt;span style="color:#66d9ef">IS&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">NULL&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> TRY_PARSE_JSON(query_tag) &lt;span style="color:#66d9ef">IS&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">NULL&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> start_time &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">7&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> model_name, materialization
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> tb_scanned &lt;span style="color:#66d9ef">DESC&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">LIMIT&lt;/span> &lt;span style="color:#ae81ff">20&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>This is how you find the one dbt model that&amp;rsquo;s responsible for 40% of your bill. Without tags, it&amp;rsquo;s a guessing game.&lt;/p>
&lt;p>For non-dbt workloads — Airflow tasks, Lambda functions, BI tools — set query tags at the session level in your connection configuration. In Python:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> snowflake.connector
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> json
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>conn &lt;span style="color:#f92672">=&lt;/span> snowflake&lt;span style="color:#f92672">.&lt;/span>connector&lt;span style="color:#f92672">.&lt;/span>connect(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> account&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;your_account&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> user&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;your_user&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> password&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;your_password&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;transform_wh&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> session_parameters&lt;span style="color:#f92672">=&lt;/span>{
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;QUERY_TAG&amp;#39;&lt;/span>: json&lt;span style="color:#f92672">.&lt;/span>dumps({
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;pipeline&amp;#39;&lt;/span>: &lt;span style="color:#e6db74">&amp;#39;customer_ingestion&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;task&amp;#39;&lt;/span>: &lt;span style="color:#e6db74">&amp;#39;load_raw_customers&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;environment&amp;#39;&lt;/span>: &lt;span style="color:#e6db74">&amp;#39;production&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> })
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> }
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The goal is simple: &lt;strong>every query that runs on your platform should be attributable to a team, a pipeline, or a tool.&lt;/strong> Start with the big consumers and work outward.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="kill-the-zombies">Kill the zombies&lt;/h3>
&lt;br>
&lt;p>Every data platform accumulates dead weight. Tables nobody queries. Pipelines that run faithfully every morning, transforming data that no dashboard, no analyst, and no model has touched in months.&lt;/p>
&lt;p>These zombies cost you twice: once in compute (the pipeline that refreshes them) and once in storage (the data that sits there). Finding them is straightforward:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Tables with zero reads in the last 90 days
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> t.table_schema,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> t.&lt;span style="color:#66d9ef">table_name&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> t.&lt;span style="color:#66d9ef">row_count&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> t.bytes &lt;span style="color:#f92672">/&lt;/span> POWER(&lt;span style="color:#ae81ff">1024&lt;/span>, &lt;span style="color:#ae81ff">3&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> size_gb,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> t.last_altered,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">MAX&lt;/span>(ah.query_start_time) &lt;span style="color:#66d9ef">AS&lt;/span> last_queried
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.tables t
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">LEFT&lt;/span> &lt;span style="color:#66d9ef">JOIN&lt;/span> snowflake.account_usage.access_history ah
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">ON&lt;/span> ah.base_objects_accessed &lt;span style="color:#66d9ef">LIKE&lt;/span> &lt;span style="color:#e6db74">&amp;#39;%&amp;#39;&lt;/span> &lt;span style="color:#f92672">||&lt;/span> t.&lt;span style="color:#66d9ef">table_name&lt;/span> &lt;span style="color:#f92672">||&lt;/span> &lt;span style="color:#e6db74">&amp;#39;%&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> ah.query_start_time &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">90&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> t.table_schema &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">IN&lt;/span> (&lt;span style="color:#e6db74">&amp;#39;INFORMATION_SCHEMA&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> t.deleted &lt;span style="color:#66d9ef">IS&lt;/span> &lt;span style="color:#66d9ef">NULL&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> t.&lt;span style="color:#66d9ef">row_count&lt;/span> &lt;span style="color:#f92672">&amp;gt;&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> t.table_schema, t.&lt;span style="color:#66d9ef">table_name&lt;/span>, t.&lt;span style="color:#66d9ef">row_count&lt;/span>, t.bytes, t.last_altered
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">HAVING&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> last_queried &lt;span style="color:#66d9ef">IS&lt;/span> &lt;span style="color:#66d9ef">NULL&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> size_gb &lt;span style="color:#66d9ef">DESC&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>On the AWS side, check your S3 storage for orphaned data. Landing zones accumulate raw files that were loaded months ago and never cleaned up. A simple lifecycle policy handles this:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-json" data-lang="json">&lt;span style="display:flex;">&lt;span>{
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;Rules&amp;#34;&lt;/span>: [
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;ID&amp;#34;&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;Archive raw data after 90 days&amp;#34;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;Status&amp;#34;&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;Enabled&amp;#34;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;Filter&amp;#34;&lt;/span>: {
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;Prefix&amp;#34;&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;raw-landing/&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> },
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;Transitions&amp;#34;&lt;/span>: [
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;Days&amp;#34;&lt;/span>: &lt;span style="color:#ae81ff">90&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;StorageClass&amp;#34;&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;GLACIER_INSTANT_RETRIEVAL&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> }
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ],
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;Expiration&amp;#34;&lt;/span>: {
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">&amp;#34;Days&amp;#34;&lt;/span>: &lt;span style="color:#ae81ff">365&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> }
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> }
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>}
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The cultural shift matters as much as the tooling. Make it a quarterly habit: pull the zombie table list, review it with the team, and deprecate what nobody uses. If someone screams, you can always restore from Time Travel. But in my experience, nobody screams. The data was already dead — you&amp;rsquo;re just acknowledging it.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="cost-anomalies-will-find-you--or-you-can-find-them-first">Cost anomalies will find you — or you can find them first&lt;/h3>
&lt;br>
&lt;p>The scariest cost event isn&amp;rsquo;t the gradual creep. It&amp;rsquo;s the single bad query or misconfigured pipeline that doubles your weekly bill overnight. I&amp;rsquo;ve seen a single Snowflake query with a missing WHERE clause scan an entire 2TB table repeatedly inside a loop — burning through hundreds of credits in an hour.&lt;/p>
&lt;p>On Snowflake, set up resource monitors as a basic guardrail:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Create a resource monitor with alerts and hard stop
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">CREATE&lt;/span> &lt;span style="color:#66d9ef">OR&lt;/span> &lt;span style="color:#66d9ef">REPLACE&lt;/span> RESOURCE MONITOR monthly_budget
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WITH&lt;/span> CREDIT_QUOTA &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">5000&lt;/span> &lt;span style="color:#75715e">-- adjust to your monthly budget
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> FREQUENCY &lt;span style="color:#f92672">=&lt;/span> MONTHLY
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> START_TIMESTAMP &lt;span style="color:#f92672">=&lt;/span> IMMEDIATELY
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> TRIGGERS
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">ON&lt;/span> &lt;span style="color:#ae81ff">75&lt;/span> PERCENT &lt;span style="color:#66d9ef">DO&lt;/span> &lt;span style="color:#66d9ef">NOTIFY&lt;/span> &lt;span style="color:#75715e">-- email alert at 75%
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> &lt;span style="color:#66d9ef">ON&lt;/span> &lt;span style="color:#ae81ff">90&lt;/span> PERCENT &lt;span style="color:#66d9ef">DO&lt;/span> &lt;span style="color:#66d9ef">NOTIFY&lt;/span> &lt;span style="color:#75715e">-- email alert at 90%
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> &lt;span style="color:#66d9ef">ON&lt;/span> &lt;span style="color:#ae81ff">100&lt;/span> PERCENT &lt;span style="color:#66d9ef">DO&lt;/span> SUSPEND; &lt;span style="color:#75715e">-- hard stop at 100%
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Apply it to a warehouse
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">ALTER&lt;/span> WAREHOUSE transform_wh
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SET&lt;/span> RESOURCE_MONITOR &lt;span style="color:#f92672">=&lt;/span> monthly_budget;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>On AWS, enable Cost Anomaly Detection in the AWS Cost Management console. It uses ML to detect unusual spending patterns and sends alerts via SNS. For Snowflake-specific monitoring, a lightweight approach is a scheduled task that checks daily credit consumption against a rolling average:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Create a simple anomaly detection view
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">CREATE&lt;/span> &lt;span style="color:#66d9ef">OR&lt;/span> &lt;span style="color:#66d9ef">REPLACE&lt;/span> &lt;span style="color:#66d9ef">VIEW&lt;/span> monitoring.daily_cost_anomalies &lt;span style="color:#66d9ef">AS&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WITH&lt;/span> daily_usage &lt;span style="color:#66d9ef">AS&lt;/span> (
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> DATE_TRUNC(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, start_time) &lt;span style="color:#66d9ef">AS&lt;/span> usage_date,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(credits_used) &lt;span style="color:#66d9ef">AS&lt;/span> daily_credits
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.warehouse_metering_history
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> start_time &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">60&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> usage_date, warehouse_name
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>),
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>averages &lt;span style="color:#66d9ef">AS&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AVG&lt;/span>(daily_credits) &lt;span style="color:#66d9ef">AS&lt;/span> avg_daily_credits,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> STDDEV(daily_credits) &lt;span style="color:#66d9ef">AS&lt;/span> stddev_daily_credits
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> daily_usage
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> usage_date &lt;span style="color:#f92672">&amp;lt;&lt;/span> &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span>()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse_name
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> du.usage_date,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> du.warehouse_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> du.daily_credits,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> a.avg_daily_credits,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ROUND((du.daily_credits &lt;span style="color:#f92672">-&lt;/span> a.avg_daily_credits) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#66d9ef">NULLIF&lt;/span>(a.stddev_daily_credits, &lt;span style="color:#ae81ff">0&lt;/span>), &lt;span style="color:#ae81ff">2&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> z_score
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> daily_usage du
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">JOIN&lt;/span> averages a
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">ON&lt;/span> du.warehouse_name &lt;span style="color:#f92672">=&lt;/span> a.warehouse_name
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> du.usage_date &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span>()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> (du.daily_credits &lt;span style="color:#f92672">-&lt;/span> a.avg_daily_credits) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#66d9ef">NULLIF&lt;/span>(a.stddev_daily_credits, &lt;span style="color:#ae81ff">0&lt;/span>) &lt;span style="color:#f92672">&amp;gt;&lt;/span> &lt;span style="color:#ae81ff">2&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> z_score &lt;span style="color:#66d9ef">DESC&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>A z-score above 2 means today&amp;rsquo;s spend is more than two standard deviations above the 60-day average. That&amp;rsquo;s worth investigating. Pipe this into a Teams or Slack alert via an AWS Lambda function and you&amp;rsquo;ve got same-day cost anomaly detection for the price of a few lines of SQL.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="what-your-data-model-costs-you">What your data model costs you&lt;/h3>
&lt;br>
&lt;p>This one&amp;rsquo;s subtle but powerful. The way you model your data directly affects how much compute you burn on every query.&lt;/p>
&lt;p>In columnar warehouses like Snowflake, wide denormalised tables query faster than star schemas with multiple JOINs — because each JOIN has overhead, and Snowflake is optimised for scanning columns from flat structures. But wide tables cost more to &lt;em>maintain&lt;/em> because updating a single dimension attribute means rewriting many rows in the fact table.&lt;/p>
&lt;p>The practical approach is a layered architecture:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Staging models&lt;/strong> (views or ephemeral): Minimal transformation, column selection, type casting. Cheap to run, no storage cost.&lt;/li>
&lt;li>&lt;strong>Intermediate models&lt;/strong> (tables or incremental): Business logic, deduplication, SCD handling. Materialised because they&amp;rsquo;re referenced by multiple downstream models.&lt;/li>
&lt;li>&lt;strong>Mart models&lt;/strong> (incremental or table): Wide, denormalised, optimised for consumer queries. These are what analysts and BI tools actually hit.&lt;/li>
&lt;/ul>
&lt;p>The cost trap is materialising too much too early. Every table materialisation means Snowflake stores and maintains that data. Every &lt;code>dbt run&lt;/code> rebuilds it. If an intermediate model is only referenced by one downstream model, make it ephemeral or a view — let Snowflake inline it at query time.&lt;/p>
&lt;p>Check your dbt DAG for this pattern: a staging model materialised as a table, referenced by a single intermediate model, which is also materialised as a table, referenced by a single mart. That&amp;rsquo;s three materialisations where one (the mart) would suffice. The staging and intermediate models can be views or ephemeral — the compute happens once when the mart builds, not three separate times.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="coming-back-to-why-this-matters">Coming back to why this matters&lt;/h3>
&lt;br>
&lt;p>I started this article with a story about not being able to explain my own platform&amp;rsquo;s costs. That wasn&amp;rsquo;t a technical failure. It was a leadership failure. I&amp;rsquo;d built something valuable and then neglected the part that kept it funded.&lt;/p>
&lt;p>The data engineers I respect most aren&amp;rsquo;t the ones who build the most elegant pipelines. They&amp;rsquo;re the ones who can walk into a budget conversation and say: &amp;ldquo;Here&amp;rsquo;s what we spend. Here&amp;rsquo;s what we get for it. Here&amp;rsquo;s what I&amp;rsquo;d cut, and here&amp;rsquo;s what I&amp;rsquo;d invest more in.&amp;rdquo; That&amp;rsquo;s the kind of clarity that earns trust — and trust is what keeps data teams alive when the inevitable cost-cutting conversations happen.&lt;/p>
&lt;p>You don&amp;rsquo;t need a FinOps certification or an expensive monitoring tool to get there. You need the queries in this article, a quarterly habit of reviewing them, and the willingness to kill the zombies nobody wants to admit are dead.&lt;/p>
&lt;p>Start with the 15-minute audit. Tag your queries. Set up a resource monitor. Do those three things this week and you&amp;rsquo;ll know more about your platform&amp;rsquo;s economics than most data teams learn in a year.&lt;/p>
&lt;p>The cheapest query is the one you never need to run. But the most valuable skill is knowing which queries are worth paying for.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Cloud Architecture</category><category>Snowflake</category><category>AWS</category><category>Cost Optimization</category><category>FinOps</category><category>dbt</category><category>Data Platform</category></item><item><title>Why Your Pipeline Finishes Later Every Month</title><link>https://ghostinthedata.info/posts/2026/2026-04-18-pipeline-optimization/</link><pubDate>Sat, 18 Apr 2026 09:00:00 +1100</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-04-18-pipeline-optimization/</guid><author>Chris Hillman</author><description>A practical guide to diagnosing pipeline bottlenecks, fixing unnecessary dependencies, and getting data to consumers faster — with Snowflake and AWS patterns you can apply today.</description><content:encoded>&lt;p>Let me tell you about a graph that changed how I think about data engineering.&lt;/p>
&lt;p>A junior engineer on my team — let&amp;rsquo;s call her Priya — had been tracking something nobody asked her to track. Every morning for two months, she&amp;rsquo;d noted the timestamp when our main analytics pipeline completed. She wasn&amp;rsquo;t trying to make a point. She was just curious, because the finance team kept mentioning their dashboards weren&amp;rsquo;t ready when they arrived at 8 AM anymore.&lt;/p>
&lt;p>One afternoon she pulled me aside and showed me a scatter plot on her laptop. Pipeline completion time, plotted daily over several months. The trend was unmistakable: a slow, steady drift to the right. What used to finish at 5:47 AM was now finishing at 7:23 AM. And the slope wasn&amp;rsquo;t flattening.&lt;/p>
&lt;p>&amp;ldquo;If this keeps going,&amp;rdquo; she said, &amp;ldquo;we&amp;rsquo;ll miss the 9 AM SLA in about six weeks.&amp;rdquo;&lt;/p>
&lt;p>She was right. And nobody else on the team — including me — had noticed. We were watching for failures. Green DAGs, clean logs, no alerts. But the pipeline wasn&amp;rsquo;t failing. It was &lt;em>slowing&lt;/em>. And slow is harder to see than broken, because slow doesn&amp;rsquo;t trigger an alert. Slow just quietly erodes trust until one day someone in finance builds their own spreadsheet and stops asking you for anything.&lt;/p>
&lt;br>
&lt;p>That&amp;rsquo;s the moment I understood that pipeline health isn&amp;rsquo;t about pass/fail. It&amp;rsquo;s about &lt;em>trajectory&lt;/em>. A pipeline that runs successfully but takes 5% longer every month is a ticking clock. And the people who notice first aren&amp;rsquo;t the engineers watching the DAG — they&amp;rsquo;re the consumers waiting for their data.&lt;/p>
&lt;p>I&amp;rsquo;m telling you this because pipeline optimisation sounds like a performance engineering problem, and it is. But underneath the technical work, it&amp;rsquo;s really about a commitment: the commitment to deliver data when you said you would, every single day. That&amp;rsquo;s what builds trust between data teams and the rest of the organisation. Not fancy architectures. Not real-time everything. Just showing up on time, reliably.&lt;/p>
&lt;p>This article is about diagnosing why your pipelines get slower, identifying the bottlenecks that actually matter, and fixing the patterns that cause data to arrive later than it should. Everything is grounded in Snowflake, AWS, Airflow, and dbt — with specific patterns you can apply immediately.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="measure-the-pipeline-in-stages-not-as-a-single-number">Measure the pipeline in stages, not as a single number&lt;/h3>
&lt;br>
&lt;p>When a pipeline is slow, the instinct is to look at the longest-running task. That&amp;rsquo;s often the wrong place to start.&lt;/p>
&lt;p>A pipeline is a chain of stages: extract, transform, load, test. The total runtime is the sum of all stages on the critical path — the longest chain of dependent tasks. But the &lt;em>bottleneck&lt;/em> might not be the longest task. It might be a 30-second task that blocks five parallel branches from starting.&lt;/p>
&lt;p>Before you optimise anything, instrument your pipeline to measure each stage independently. In Airflow, the task instance metadata already captures this.&lt;/p>
&lt;p>Two numbers matter here: &lt;code>duration_seconds&lt;/code> (how long the task actually ran) and &lt;code>queue_wait_seconds&lt;/code> (how long it waited before it could run). If queue wait is high, your problem isn&amp;rsquo;t the task — it&amp;rsquo;s resource contention. Too many tasks competing for too few Airflow worker slots, or too many queries competing for the same Snowflake warehouse.&lt;/p>
&lt;p>For dbt runs specifically, the &lt;code>run_results.json&lt;/code> that dbt generates after every invocation is a goldmine:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> json
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">with&lt;/span> open(&lt;span style="color:#e6db74">&amp;#39;target/run_results.json&amp;#39;&lt;/span>) &lt;span style="color:#66d9ef">as&lt;/span> f:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> results &lt;span style="color:#f92672">=&lt;/span> json&lt;span style="color:#f92672">.&lt;/span>load(f)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Find your slowest models&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>models &lt;span style="color:#f92672">=&lt;/span> sorted(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> [r &lt;span style="color:#66d9ef">for&lt;/span> r &lt;span style="color:#f92672">in&lt;/span> results[&lt;span style="color:#e6db74">&amp;#39;results&amp;#39;&lt;/span>] &lt;span style="color:#66d9ef">if&lt;/span> r[&lt;span style="color:#e6db74">&amp;#39;status&amp;#39;&lt;/span>] &lt;span style="color:#f92672">==&lt;/span> &lt;span style="color:#e6db74">&amp;#39;success&amp;#39;&lt;/span>],
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> key&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#66d9ef">lambda&lt;/span> x: x[&lt;span style="color:#e6db74">&amp;#39;execution_time&amp;#39;&lt;/span>],
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> reverse&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#66d9ef">True&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">for&lt;/span> m &lt;span style="color:#f92672">in&lt;/span> models[:&lt;span style="color:#ae81ff">10&lt;/span>]:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> print(&lt;span style="color:#e6db74">f&lt;/span>&lt;span style="color:#e6db74">&amp;#34;&lt;/span>&lt;span style="color:#e6db74">{&lt;/span>m[&lt;span style="color:#e6db74">&amp;#39;unique_id&amp;#39;&lt;/span>]&lt;span style="color:#e6db74">:&lt;/span>&lt;span style="color:#e6db74">60s&lt;/span>&lt;span style="color:#e6db74">}&lt;/span>&lt;span style="color:#e6db74"> &lt;/span>&lt;span style="color:#e6db74">{&lt;/span>m[&lt;span style="color:#e6db74">&amp;#39;execution_time&amp;#39;&lt;/span>]&lt;span style="color:#e6db74">:&lt;/span>&lt;span style="color:#e6db74">8.1f&lt;/span>&lt;span style="color:#e6db74">}&lt;/span>&lt;span style="color:#e6db74">s&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Run that after your next &lt;code>dbt build&lt;/code>. You&amp;rsquo;ll immediately see which models dominate your pipeline runtime. In my experience, 3–5 models account for 60–80% of total execution time. Those are the only models worth optimising.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="your-dag-probably-has-dependencies-it-doesnt-need">Your DAG probably has dependencies it doesn&amp;rsquo;t need&lt;/h3>
&lt;br>
&lt;p>This is the most common structural problem in data pipelines, and it comes in two forms. One is obvious once you know to look for it. The other is sneaky, because it starts as a completely reasonable engineering decision.&lt;/p>
&lt;p>&lt;strong>The first form: phantom dependencies.&lt;/strong>&lt;/p>
&lt;p>An engineer builds Model C and declares it depends on Model B — because Model B produces a table that Model C reads. Fair enough. But six months later, someone refactors Model C and it no longer reads from Model B. The &lt;code>ref()&lt;/code> gets removed from the SQL. But the dependency in the Airflow DAG? That stays. Or in dbt, someone adds a &lt;code>ref()&lt;/code> to a model they don&amp;rsquo;t actually need, just to &amp;ldquo;make sure it runs first.&amp;rdquo;&lt;/p>
&lt;p>The result is a DAG with phantom dependencies — tasks that wait for other tasks to complete even though they don&amp;rsquo;t use those tasks&amp;rsquo; outputs. Every phantom dependency adds serial wait time to your pipeline.&lt;/p>
&lt;p>&lt;strong>The second form: dependency monsters.&lt;/strong>&lt;/p>
&lt;p>This one is trickier, because it starts with good intentions.&lt;/p>
&lt;p>The customer team needs three new enrichment attributes on &lt;code>dim_customers&lt;/code> — regional segment codes sourced from a CRM export, tenure tier derived from a subscription history table, and a propensity score from a data science model. All reasonable requests. Each one approved and added without much ceremony.&lt;/p>
&lt;p>But &lt;code>dim_customers&lt;/code> is upstream of 40+ models across your DAG. Adding those three attributes means pulling in the CRM extract, joining the subscription history table — which has its own upstream dependencies — and waiting on the propensity score model to complete before any of those 40 downstream models can start. Your pipeline used to finish at 9 AM. Eighteen months and a dozen enrichment requests later, it finishes at 3 PM.&lt;/p>
&lt;p>Each addition was individually justified. Nobody modelled what they&amp;rsquo;d cost collectively. That&amp;rsquo;s how you build a dependency monster.&lt;/p>
&lt;p>The fix isn&amp;rsquo;t to refuse enrichment requests — it&amp;rsquo;s to stop embedding enrichment into your core spine model. Keep &lt;code>dim_customers&lt;/code> lean: identifiers, names, status, the attributes that nearly every consumer genuinely needs. Build a separate &lt;code>dim_customers_extended&lt;/code> model that joins in the expensive enrichment for the consumers who actually need it. Most of your downstream models will never touch the propensity score. There&amp;rsquo;s no reason to make them wait for it.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- dim_customers.sql: lean spine — runs fast, unblocks the DAG
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> email,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> status,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> created_at
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">ref&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;stg_customers&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- dim_customers_extended.sql: enrichment for consumers who need it
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- runs independently, doesn&amp;#39;t block the main pipeline
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">c&lt;/span>.customer_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">c&lt;/span>.customer_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">c&lt;/span>.status,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> crm.regional_segment,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> sub.tenure_tier,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ps.propensity_score
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">ref&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;dim_customers&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span> &lt;span style="color:#66d9ef">c&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">LEFT&lt;/span> &lt;span style="color:#66d9ef">JOIN&lt;/span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">ref&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;stg_crm_segments&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span> crm
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">on&lt;/span> crm.customer_id &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">c&lt;/span>.customer_id
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">LEFT&lt;/span> &lt;span style="color:#66d9ef">JOIN&lt;/span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">ref&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;int_subscription_tenure&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span> sub
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">on&lt;/span> sub.customer_id &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">c&lt;/span>.customer_id
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">LEFT&lt;/span> &lt;span style="color:#66d9ef">JOIN&lt;/span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">ref&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;ml_propensity_scores&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span> ps
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">on&lt;/span> ps.customer_id &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">c&lt;/span>.customer_id
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Now &lt;code>dim_customers_extended&lt;/code> and its expensive upstream dependencies sit on their own branch of the DAG. Consumers who need the enrichment reference the extended model. The rest of the pipeline is unblocked.&lt;/p>
&lt;p>A useful governance rule of thumb before adding a field to a core model: if fewer than half your consumers will ever query that attribute, it probably doesn&amp;rsquo;t belong there. The team who needs it should own the enrichment themselves, as a downstream model they maintain.&lt;/p>
&lt;br>
&lt;p>&lt;strong>Auditing for both problems&lt;/strong>&lt;/p>
&lt;p>In dbt, the &lt;code>dbt_project_evaluator&lt;/code> package surfaces structural issues systematically. Install it:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># packages.yml&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">packages&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">package&lt;/span>: &lt;span style="color:#ae81ff">dbt-labs/dbt_project_evaluator&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">version&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;&amp;gt;=0.8.0 &amp;lt;1.0.0&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Then run:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>dbt build --select package:dbt_project_evaluator
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>It will flag models with fan-out issues (one model referenced by too many downstream models), unnecessary dependencies, and models that could be parallelised but aren&amp;rsquo;t. It won&amp;rsquo;t catch dependency monsters directly — those require a human asking &amp;ldquo;does every consumer of this model actually use every attribute in it?&amp;rdquo; — but it&amp;rsquo;s a good starting point for the structural audit.&lt;/p>
&lt;p>For Airflow DAGs, the visual DAG view tells you a lot at a glance. A healthy DAG looks like a tree with wide parallel branches. An unhealthy DAG looks like a chain — or worse, a funnel where everything converges through a single overloaded task before it can branch out again. That funnel shape is the visual signature of a dependency monster: many expensive upstreams flowing into one hub model, with dozens of downstream models queued behind it.&lt;/p>
&lt;p>If your DAG looks like a chain, ask this question for every dependency edge: &lt;strong>does downstream task B actually read data produced by upstream task A?&lt;/strong> If the answer is no, remove the dependency. Let them run in parallel.&lt;/p>
&lt;p>I once audited a DAG with 47 tasks running in a strict serial chain. After removing phantom dependencies and restructuring, 31 of those tasks could run in parallel. The pipeline went from 2 hours 15 minutes to 38 minutes. Same tasks, same compute, same data. Just fewer unnecessary wait states.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-critical-path-is-the-only-thing-worth-optimising">The critical path is the only thing worth optimising&lt;/h3>
&lt;br>
&lt;p>Here&amp;rsquo;s a concept from project management that applies directly to pipeline engineering: the &lt;strong>critical path&lt;/strong>.&lt;/p>
&lt;p>The critical path is the longest chain of dependent tasks through your DAG. It determines your pipeline&amp;rsquo;s minimum possible runtime. Every other path through the DAG has &amp;ldquo;float&amp;rdquo; — slack time where tasks can be delayed without affecting the overall completion time.&lt;/p>
&lt;p>This means: &lt;strong>optimising a task that isn&amp;rsquo;t on the critical path has zero impact on your pipeline&amp;rsquo;s end-to-end runtime.&lt;/strong> You could make a non-critical task 10x faster and your pipeline would finish at exactly the same time.&lt;/p>
&lt;p>Finding the critical path requires knowing each task&amp;rsquo;s duration and dependency structure. For a dbt project, you can approximate it:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Using dbt&amp;#39;s run results stored in Snowflake (if you log them)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Or query your orchestrator&amp;#39;s task instance table
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">WITH&lt;/span> task_durations &lt;span style="color:#66d9ef">AS&lt;/span> (
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> task_id &lt;span style="color:#66d9ef">AS&lt;/span> model_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AVG&lt;/span>(&lt;span style="color:#66d9ef">EXTRACT&lt;/span>(EPOCH &lt;span style="color:#66d9ef">FROM&lt;/span> (end_date &lt;span style="color:#f92672">-&lt;/span> start_date))) &lt;span style="color:#66d9ef">AS&lt;/span> avg_duration
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">FROM&lt;/span> task_instance
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHERE&lt;/span> dag_id &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;dbt_daily&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> &lt;span style="color:#66d9ef">state&lt;/span> &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;success&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> execution_date &lt;span style="color:#f92672">&amp;gt;=&lt;/span> &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span> &lt;span style="color:#f92672">-&lt;/span> INTERVAL &lt;span style="color:#e6db74">&amp;#39;7 days&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span> task_id
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> model_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ROUND(avg_duration, &lt;span style="color:#ae81ff">1&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> avg_seconds,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ROUND(avg_duration &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#ae81ff">60&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> avg_minutes
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span> task_durations
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span> avg_duration &lt;span style="color:#66d9ef">DESC&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">LIMIT&lt;/span> &lt;span style="color:#ae81ff">20&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The longest tasks are &lt;em>candidates&lt;/em> for the critical path, but only if they&amp;rsquo;re on the dependency chain that determines total runtime. A 20-minute model that runs in parallel with a 45-minute model isn&amp;rsquo;t the bottleneck — the 45-minute model is.&lt;/p>
&lt;p>Once you&amp;rsquo;ve identified the critical path, you have three options for shortening it: make the slow tasks faster (query optimisation, incremental models), reduce the number of tasks on the path (remove unnecessary dependencies), or parallelise sequential tasks (split a monolithic model into independent pieces).&lt;/p>
&lt;hr>
&lt;h3 id="shifting-right-diagnose-it-before-it-breaks-your-sla">Shifting right: diagnose it before it breaks your SLA&lt;/h3>
&lt;br>
&lt;p>Remember Priya&amp;rsquo;s scatter plot? That pattern — pipeline completion drifting later and later — has a name: &lt;strong>shifting right&lt;/strong>. And it has predictable causes.&lt;/p>
&lt;p>Track it with a simple query against your orchestrator&amp;rsquo;s metadata:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Airflow: track pipeline completion time drift
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> execution_date::DATE &lt;span style="color:#66d9ef">AS&lt;/span> run_date,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">MAX&lt;/span>(end_date)::TIME &lt;span style="color:#66d9ef">AS&lt;/span> completion_time,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">EXTRACT&lt;/span>(EPOCH &lt;span style="color:#66d9ef">FROM&lt;/span> (&lt;span style="color:#66d9ef">MAX&lt;/span>(end_date) &lt;span style="color:#f92672">-&lt;/span> &lt;span style="color:#66d9ef">MIN&lt;/span>(start_date))) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#ae81ff">60&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> total_runtime_minutes
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> task_instance
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> dag_id &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;daily_analytics&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> &lt;span style="color:#66d9ef">state&lt;/span> &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;success&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> execution_date &lt;span style="color:#f92672">&amp;gt;=&lt;/span> &lt;span style="color:#66d9ef">CURRENT_DATE&lt;/span> &lt;span style="color:#f92672">-&lt;/span> INTERVAL &lt;span style="color:#e6db74">&amp;#39;60 days&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> run_date
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> run_date;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Plot &lt;code>completion_time&lt;/code> over &lt;code>run_date&lt;/code>. If it trends upward, you&amp;rsquo;re shifting right. The four root causes, in order of how often I see them:&lt;/p>
&lt;p>&lt;strong>1. Data volume growth.&lt;/strong> Your table had 10 million rows when you built the model. Now it has 200 million. That JOIN that took 8 seconds now takes 3 minutes. This is the most common cause and the easiest to fix — switch to incremental materialisation and the growth stops mattering.&lt;/p>
&lt;p>&lt;strong>2. Resource contention.&lt;/strong> You&amp;rsquo;ve added more pipelines and they all run in the same window, competing for the same Snowflake warehouse. In Snowflake, you can see this directly:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Find warehouse queueing (queries waiting for compute)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> DATE_TRUNC(&lt;span style="color:#e6db74">&amp;#39;hour&amp;#39;&lt;/span>, start_time) &lt;span style="color:#66d9ef">AS&lt;/span> hour,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(&lt;span style="color:#f92672">*&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> total_queries,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SUM&lt;/span>(&lt;span style="color:#66d9ef">CASE&lt;/span> &lt;span style="color:#66d9ef">WHEN&lt;/span> queued_overload_time &lt;span style="color:#f92672">&amp;gt;&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span> &lt;span style="color:#66d9ef">THEN&lt;/span> &lt;span style="color:#ae81ff">1&lt;/span> &lt;span style="color:#66d9ef">ELSE&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span> &lt;span style="color:#66d9ef">END&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> queued_queries,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AVG&lt;/span>(queued_overload_time) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#ae81ff">1000&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> avg_queue_seconds
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.query_history
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> start_time &lt;span style="color:#f92672">&amp;gt;=&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">7&lt;/span>, &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>())
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> warehouse_name &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;TRANSFORM_WH&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GROUP&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> warehouse_name, hour
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">HAVING&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> queued_queries &lt;span style="color:#f92672">&amp;gt;&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> hour;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>If &lt;code>queued_queries&lt;/code> is non-zero during your pipeline window, queries are waiting for warehouse compute. Either scale up the warehouse during that window, split workloads across dedicated warehouses, or stagger pipeline start times to reduce concurrency.&lt;/p>
&lt;p>&lt;strong>3. Dependency chain lengthening.&lt;/strong> Someone added a new intermediate model between your staging and mart layers. That model takes 4 minutes. But because it&amp;rsquo;s on the critical path, the entire pipeline now finishes 4 minutes later. This is death by a thousand cuts — each addition is small, but they accumulate.&lt;/p>
&lt;p>&lt;strong>4. Upstream source delays.&lt;/strong> Your pipeline starts at 4 AM because the source system&amp;rsquo;s extract used to land in S3 by 3:45 AM. But the source system has grown too, and now the extract doesn&amp;rsquo;t land until 4:30 AM. Your pipeline sensors wait, and everything shifts right.&lt;/p>
&lt;p>For upstream delays in an AWS environment, replace time-based scheduling with event-driven triggers. Instead of scheduling your Airflow DAG at 4 AM and hoping the data is there, trigger it when the data actually arrives:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Airflow DAG: trigger on S3 file landing using AWS sensor&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> airflow.providers.amazon.aws.sensors.s3 &lt;span style="color:#f92672">import&lt;/span> S3KeySensor
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> airflow &lt;span style="color:#f92672">import&lt;/span> DAG
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> datetime &lt;span style="color:#f92672">import&lt;/span> datetime
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">with&lt;/span> DAG(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;daily_analytics&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> start_date&lt;span style="color:#f92672">=&lt;/span>datetime(&lt;span style="color:#ae81ff">2026&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>),
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> schedule_interval&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#66d9ef">None&lt;/span>, &lt;span style="color:#75715e"># triggered externally, not on a schedule&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> catchup&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#66d9ef">False&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>) &lt;span style="color:#66d9ef">as&lt;/span> dag:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> wait_for_source &lt;span style="color:#f92672">=&lt;/span> S3KeySensor(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> task_id&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;wait_for_orders_extract&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> bucket_name&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;your-data-lake&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> bucket_key&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;raw/orders/dt={{ ds }}/orders.parquet&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> aws_conn_id&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;aws_default&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> mode&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;reschedule&amp;#39;&lt;/span>, &lt;span style="color:#75715e"># frees the worker slot while waiting&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> poke_interval&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">300&lt;/span>, &lt;span style="color:#75715e"># check every 5 minutes&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> timeout&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">7200&lt;/span>, &lt;span style="color:#75715e"># fail after 2 hours if file never arrives&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> )
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The &lt;code>mode='reschedule'&lt;/code> is critical. Without it, the sensor occupies a worker slot for the entire time it&amp;rsquo;s waiting. With &lt;code>reschedule&lt;/code>, it checks, releases the slot, and checks again later. This prevents the classic deadlock where all your worker slots are consumed by sensors and no actual work can run.&lt;/p>
&lt;p>Even better: use S3 event notifications to trigger a Lambda function that kicks off the DAG via Airflow&amp;rsquo;s REST API. Zero polling, zero wasted slots:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Lambda function triggered by S3 PutObject event&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> boto3
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> requests
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> os
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">handler&lt;/span>(event, context):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> bucket &lt;span style="color:#f92672">=&lt;/span> event[&lt;span style="color:#e6db74">&amp;#39;Records&amp;#39;&lt;/span>][&lt;span style="color:#ae81ff">0&lt;/span>][&lt;span style="color:#e6db74">&amp;#39;s3&amp;#39;&lt;/span>][&lt;span style="color:#e6db74">&amp;#39;bucket&amp;#39;&lt;/span>][&lt;span style="color:#e6db74">&amp;#39;name&amp;#39;&lt;/span>]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> key &lt;span style="color:#f92672">=&lt;/span> event[&lt;span style="color:#e6db74">&amp;#39;Records&amp;#39;&lt;/span>][&lt;span style="color:#ae81ff">0&lt;/span>][&lt;span style="color:#e6db74">&amp;#39;s3&amp;#39;&lt;/span>][&lt;span style="color:#e6db74">&amp;#39;object&amp;#39;&lt;/span>][&lt;span style="color:#e6db74">&amp;#39;key&amp;#39;&lt;/span>]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#75715e"># Trigger Airflow DAG via REST API&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> airflow_url &lt;span style="color:#f92672">=&lt;/span> os&lt;span style="color:#f92672">.&lt;/span>environ[&lt;span style="color:#e6db74">&amp;#39;AIRFLOW_API_URL&amp;#39;&lt;/span>]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> response &lt;span style="color:#f92672">=&lt;/span> requests&lt;span style="color:#f92672">.&lt;/span>post(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">f&lt;/span>&lt;span style="color:#e6db74">&amp;#34;&lt;/span>&lt;span style="color:#e6db74">{&lt;/span>airflow_url&lt;span style="color:#e6db74">}&lt;/span>&lt;span style="color:#e6db74">/api/v1/dags/daily_analytics/dagRuns&amp;#34;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> json&lt;span style="color:#f92672">=&lt;/span>{
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#34;conf&amp;#34;&lt;/span>: {&lt;span style="color:#e6db74">&amp;#34;source_bucket&amp;#34;&lt;/span>: bucket, &lt;span style="color:#e6db74">&amp;#34;source_key&amp;#34;&lt;/span>: key}
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> },
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> auth&lt;span style="color:#f92672">=&lt;/span>(os&lt;span style="color:#f92672">.&lt;/span>environ[&lt;span style="color:#e6db74">&amp;#39;AIRFLOW_USER&amp;#39;&lt;/span>], os&lt;span style="color:#f92672">.&lt;/span>environ[&lt;span style="color:#e6db74">&amp;#39;AIRFLOW_PASSWORD&amp;#39;&lt;/span>]),
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> headers&lt;span style="color:#f92672">=&lt;/span>{&lt;span style="color:#e6db74">&amp;#34;Content-Type&amp;#34;&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;application/json&amp;#34;&lt;/span>}
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> )
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">return&lt;/span> {&lt;span style="color:#e6db74">&amp;#34;statusCode&amp;#34;&lt;/span>: response&lt;span style="color:#f92672">.&lt;/span>status_code}
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>This eliminates the &amp;ldquo;safety margin&amp;rdquo; scheduling pattern entirely. Your pipeline runs as soon as the data is available — not 30 minutes after you &lt;em>hope&lt;/em> it will be available.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="partition-and-cluster-to-eliminate-full-table-scans">Partition and cluster to eliminate full table scans&lt;/h3>
&lt;br>
&lt;p>The single most effective performance optimisation in Snowflake is making sure your queries only read the data they need. Snowflake&amp;rsquo;s micro-partition pruning does this automatically — but only if your data is physically organised in a way that aligns with your query patterns.&lt;/p>
&lt;p>Clustering keys tell Snowflake how to organise data within micro-partitions. If you consistently filter on &lt;code>order_date&lt;/code>, clustering on that column means queries with a &lt;code>WHERE order_date = '2026-03-01'&lt;/code> clause scan a tiny fraction of the table instead of the whole thing.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Add clustering to a large fact table
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">ALTER&lt;/span> &lt;span style="color:#66d9ef">TABLE&lt;/span> analytics.fct_orders
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">CLUSTER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span> (order_date, customer_segment);
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Choose clustering keys based on how the table is actually queried, not how it&amp;rsquo;s loaded. The best candidates are columns that appear in WHERE clauses, JOIN conditions, and range filters. Limit yourself to 2–3 keys — more than that and Snowflake can&amp;rsquo;t maintain effective clustering.&lt;/p>
&lt;p>Check whether your existing tables benefit from clustering:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Check clustering depth (lower is better, 0 is perfect)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">table_name&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> clustering_key,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> total_constant_partition_count,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> total_partition_count,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> average_overlaps,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> average_depth
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.table_storage_metrics
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> table_catalog &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;YOUR_DATABASE&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> clustering_key &lt;span style="color:#66d9ef">IS&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">NULL&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> active_bytes &lt;span style="color:#f92672">&amp;gt;&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> average_depth &lt;span style="color:#66d9ef">DESC&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>An &lt;code>average_depth&lt;/code> above 5 means your clustering isn&amp;rsquo;t effective — queries are still scanning more partitions than they should. Either the clustering key doesn&amp;rsquo;t match query patterns, or the table has had so many small writes that the clustering has degraded. Running &lt;code>ALTER TABLE ... RECLUSTER&lt;/code> manually or relying on automatic clustering (which costs credits) addresses the latter.&lt;/p>
&lt;p>On the AWS side, if you&amp;rsquo;re using S3 as a data lake with Parquet or Iceberg, the equivalent is &lt;strong>partition layout&lt;/strong>. Partition your S3 data by the columns you filter most:&lt;/p>



&lt;div class="goat svg-container ">
 
 &lt;svg
 xmlns="http://www.w3.org/2000/svg"
 font-family="Menlo,Lucida Console,monospace"
 
 viewBox="0 0 312 121"
 >
 &lt;g transform='translate(8,16)'>
&lt;polygon points='176.000000,80.000000 164.000000,74.400002 164.000000,85.599998' fill='currentColor' transform='rotate(90.000000, 168.000000, 80.000000)'>&lt;/polygon>
&lt;text text-anchor='middle' x='0' y='4' fill='currentColor' style='font-size:1em'>s&lt;/text>
&lt;text text-anchor='middle' x='8' y='4' fill='currentColor' style='font-size:1em'>3&lt;/text>
&lt;text text-anchor='middle' x='16' y='4' fill='currentColor' style='font-size:1em'>:&lt;/text>
&lt;text text-anchor='middle' x='24' y='4' fill='currentColor' style='font-size:1em'>/&lt;/text>
&lt;text text-anchor='middle' x='32' y='4' fill='currentColor' style='font-size:1em'>/&lt;/text>
&lt;text text-anchor='middle' x='32' y='20' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='40' y='4' fill='currentColor' style='font-size:1em'>y&lt;/text>
&lt;text text-anchor='middle' x='40' y='20' fill='currentColor' style='font-size:1em'>v&lt;/text>
&lt;text text-anchor='middle' x='48' y='4' fill='currentColor' style='font-size:1em'>o&lt;/text>
&lt;text text-anchor='middle' x='48' y='20' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='56' y='4' fill='currentColor' style='font-size:1em'>u&lt;/text>
&lt;text text-anchor='middle' x='56' y='20' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='64' y='4' fill='currentColor' style='font-size:1em'>r&lt;/text>
&lt;text text-anchor='middle' x='64' y='20' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='64' y='36' fill='currentColor' style='font-size:1em'>y&lt;/text>
&lt;text text-anchor='middle' x='72' y='4' fill='currentColor' style='font-size:1em'>-&lt;/text>
&lt;text text-anchor='middle' x='72' y='20' fill='currentColor' style='font-size:1em'>s&lt;/text>
&lt;text text-anchor='middle' x='72' y='36' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='80' y='4' fill='currentColor' style='font-size:1em'>d&lt;/text>
&lt;text text-anchor='middle' x='80' y='20' fill='currentColor' style='font-size:1em'>/&lt;/text>
&lt;text text-anchor='middle' x='80' y='36' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='88' y='4' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='88' y='36' fill='currentColor' style='font-size:1em'>r&lt;/text>
&lt;text text-anchor='middle' x='96' y='4' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='96' y='36' fill='currentColor' style='font-size:1em'>=&lt;/text>
&lt;text text-anchor='middle' x='96' y='52' fill='currentColor' style='font-size:1em'>m&lt;/text>
&lt;text text-anchor='middle' x='104' y='4' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='104' y='36' fill='currentColor' style='font-size:1em'>2&lt;/text>
&lt;text text-anchor='middle' x='104' y='52' fill='currentColor' style='font-size:1em'>o&lt;/text>
&lt;text text-anchor='middle' x='112' y='4' fill='currentColor' style='font-size:1em'>-&lt;/text>
&lt;text text-anchor='middle' x='112' y='36' fill='currentColor' style='font-size:1em'>0&lt;/text>
&lt;text text-anchor='middle' x='112' y='52' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='120' y='4' fill='currentColor' style='font-size:1em'>l&lt;/text>
&lt;text text-anchor='middle' x='120' y='36' fill='currentColor' style='font-size:1em'>2&lt;/text>
&lt;text text-anchor='middle' x='120' y='52' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='128' y='4' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='128' y='36' fill='currentColor' style='font-size:1em'>6&lt;/text>
&lt;text text-anchor='middle' x='128' y='52' fill='currentColor' style='font-size:1em'>h&lt;/text>
&lt;text text-anchor='middle' x='128' y='68' fill='currentColor' style='font-size:1em'>d&lt;/text>
&lt;text text-anchor='middle' x='136' y='4' fill='currentColor' style='font-size:1em'>k&lt;/text>
&lt;text text-anchor='middle' x='136' y='36' fill='currentColor' style='font-size:1em'>/&lt;/text>
&lt;text text-anchor='middle' x='136' y='52' fill='currentColor' style='font-size:1em'>=&lt;/text>
&lt;text text-anchor='middle' x='136' y='68' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='144' y='4' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='144' y='52' fill='currentColor' style='font-size:1em'>0&lt;/text>
&lt;text text-anchor='middle' x='144' y='68' fill='currentColor' style='font-size:1em'>y&lt;/text>
&lt;text text-anchor='middle' x='152' y='4' fill='currentColor' style='font-size:1em'>/&lt;/text>
&lt;text text-anchor='middle' x='152' y='52' fill='currentColor' style='font-size:1em'>3&lt;/text>
&lt;text text-anchor='middle' x='152' y='68' fill='currentColor' style='font-size:1em'>=&lt;/text>
&lt;text text-anchor='middle' x='160' y='52' fill='currentColor' style='font-size:1em'>/&lt;/text>
&lt;text text-anchor='middle' x='160' y='68' fill='currentColor' style='font-size:1em'>1&lt;/text>
&lt;text text-anchor='middle' x='160' y='84' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='160' y='100' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='168' y='68' fill='currentColor' style='font-size:1em'>5&lt;/text>
&lt;text text-anchor='middle' x='168' y='100' fill='currentColor' style='font-size:1em'>v&lt;/text>
&lt;text text-anchor='middle' x='176' y='68' fill='currentColor' style='font-size:1em'>/&lt;/text>
&lt;text text-anchor='middle' x='176' y='84' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='176' y='100' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='184' y='84' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='184' y='100' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='192' y='84' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='192' y='100' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='200' y='84' fill='currentColor' style='font-size:1em'>s&lt;/text>
&lt;text text-anchor='middle' x='200' y='100' fill='currentColor' style='font-size:1em'>s&lt;/text>
&lt;text text-anchor='middle' x='208' y='84' fill='currentColor' style='font-size:1em'>_&lt;/text>
&lt;text text-anchor='middle' x='208' y='100' fill='currentColor' style='font-size:1em'>_&lt;/text>
&lt;text text-anchor='middle' x='216' y='84' fill='currentColor' style='font-size:1em'>0&lt;/text>
&lt;text text-anchor='middle' x='216' y='100' fill='currentColor' style='font-size:1em'>0&lt;/text>
&lt;text text-anchor='middle' x='224' y='84' fill='currentColor' style='font-size:1em'>0&lt;/text>
&lt;text text-anchor='middle' x='224' y='100' fill='currentColor' style='font-size:1em'>0&lt;/text>
&lt;text text-anchor='middle' x='232' y='84' fill='currentColor' style='font-size:1em'>1&lt;/text>
&lt;text text-anchor='middle' x='232' y='100' fill='currentColor' style='font-size:1em'>2&lt;/text>
&lt;text text-anchor='middle' x='240' y='84' fill='currentColor' style='font-size:1em'>.&lt;/text>
&lt;text text-anchor='middle' x='240' y='100' fill='currentColor' style='font-size:1em'>.&lt;/text>
&lt;text text-anchor='middle' x='248' y='84' fill='currentColor' style='font-size:1em'>p&lt;/text>
&lt;text text-anchor='middle' x='248' y='100' fill='currentColor' style='font-size:1em'>p&lt;/text>
&lt;text text-anchor='middle' x='256' y='84' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='256' y='100' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='264' y='84' fill='currentColor' style='font-size:1em'>r&lt;/text>
&lt;text text-anchor='middle' x='264' y='100' fill='currentColor' style='font-size:1em'>r&lt;/text>
&lt;text text-anchor='middle' x='272' y='84' fill='currentColor' style='font-size:1em'>q&lt;/text>
&lt;text text-anchor='middle' x='272' y='100' fill='currentColor' style='font-size:1em'>q&lt;/text>
&lt;text text-anchor='middle' x='280' y='84' fill='currentColor' style='font-size:1em'>u&lt;/text>
&lt;text text-anchor='middle' x='280' y='100' fill='currentColor' style='font-size:1em'>u&lt;/text>
&lt;text text-anchor='middle' x='288' y='84' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='288' y='100' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='296' y='84' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='296' y='100' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;/g>

 &lt;/svg>
 
&lt;/div>
&lt;p>When Snowflake external tables or AWS Athena query this structure with a date filter, they skip entire directories. A query for March 15th reads two files instead of scanning the entire &lt;code>events/&lt;/code> prefix. The savings compound with data volume — at a billion rows, proper partitioning can reduce query times from minutes to seconds.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="make-every-task-idempotent-or-debugging-becomes-a-nightmare">Make every task idempotent or debugging becomes a nightmare&lt;/h3>
&lt;br>
&lt;p>Here&amp;rsquo;s a pattern I see constantly: a pipeline fails midway through, the engineer reruns it, and now there are duplicate rows in the target table. Because the first run inserted half the data before failing, and the rerun inserted &lt;em>all&lt;/em> the data — including the half that already existed.&lt;/p>
&lt;p>Idempotency means a task produces the same result whether it runs once or ten times. This isn&amp;rsquo;t just a nice-to-have — it&amp;rsquo;s the foundation that makes everything else in pipeline engineering possible. Without it, you can&amp;rsquo;t safely retry. You can&amp;rsquo;t backfill. You can&amp;rsquo;t debug.&lt;/p>
&lt;p>In dbt, incremental models with a &lt;code>unique_key&lt;/code> are idempotent by default — the MERGE statement handles duplicates:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> config(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> materialized&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;incremental&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> unique_key&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;event_id&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> incremental_strategy&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;merge&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> )
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> event_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> user_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> event_type,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> event_timestamp,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> properties
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">ref&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;stg_events&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#960050;background-color:#1e0010">{&lt;/span>&lt;span style="color:#f92672">%&lt;/span> &lt;span style="color:#66d9ef">if&lt;/span> is_incremental() &lt;span style="color:#f92672">%&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> event_timestamp &lt;span style="color:#f92672">&amp;gt;=&lt;/span> (
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SELECT&lt;/span> DATEADD(&lt;span style="color:#e6db74">&amp;#39;day&amp;#39;&lt;/span>, &lt;span style="color:#f92672">-&lt;/span>&lt;span style="color:#ae81ff">3&lt;/span>, &lt;span style="color:#66d9ef">MAX&lt;/span>(event_timestamp))
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">FROM&lt;/span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> this &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#960050;background-color:#1e0010">{&lt;/span>&lt;span style="color:#f92672">%&lt;/span> endif &lt;span style="color:#f92672">%&lt;/span>&lt;span style="color:#960050;background-color:#1e0010">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>For non-dbt loads — say, a Lambda function loading data from an API into Snowflake — use MERGE instead of INSERT:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>MERGE &lt;span style="color:#66d9ef">INTO&lt;/span> raw.api_customers &lt;span style="color:#66d9ef">AS&lt;/span> target
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">USING&lt;/span> (
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">SELECT&lt;/span> &lt;span style="color:#960050;background-color:#1e0010">$&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>:id::STRING &lt;span style="color:#66d9ef">AS&lt;/span> customer_id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">$&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>:name::STRING &lt;span style="color:#66d9ef">AS&lt;/span> customer_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">$&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>:email::STRING &lt;span style="color:#66d9ef">AS&lt;/span> email,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">$&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>:updated_at::TIMESTAMP_NTZ &lt;span style="color:#66d9ef">AS&lt;/span> updated_at
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">FROM&lt;/span> &lt;span style="color:#f92672">@&lt;/span>raw.s3_stage&lt;span style="color:#f92672">/&lt;/span>customers&lt;span style="color:#f92672">/&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> (FILE_FORMAT &lt;span style="color:#f92672">=&amp;gt;&lt;/span> &lt;span style="color:#e6db74">&amp;#39;json_format&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>) &lt;span style="color:#66d9ef">AS&lt;/span> &lt;span style="color:#66d9ef">source&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ON&lt;/span> target.customer_id &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">source&lt;/span>.customer_id
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHEN&lt;/span> MATCHED &lt;span style="color:#66d9ef">AND&lt;/span> &lt;span style="color:#66d9ef">source&lt;/span>.updated_at &lt;span style="color:#f92672">&amp;gt;&lt;/span> target.updated_at &lt;span style="color:#66d9ef">THEN&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">UPDATE&lt;/span> &lt;span style="color:#66d9ef">SET&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> customer_name &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">source&lt;/span>.customer_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> email &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">source&lt;/span>.email,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> updated_at &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">source&lt;/span>.updated_at
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHEN&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> MATCHED &lt;span style="color:#66d9ef">THEN&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">INSERT&lt;/span> (customer_id, customer_name, email, updated_at)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">VALUES&lt;/span> (&lt;span style="color:#66d9ef">source&lt;/span>.customer_id, &lt;span style="color:#66d9ef">source&lt;/span>.customer_name, &lt;span style="color:#66d9ef">source&lt;/span>.email, &lt;span style="color:#66d9ef">source&lt;/span>.updated_at);
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The &lt;code>WHEN MATCHED AND source.updated_at &amp;gt; target.updated_at&lt;/code> condition prevents overwriting newer data with older data during reruns. This matters when you&amp;rsquo;re replaying historical loads or running overlapping backfills.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="most-we-need-real-time-requests-actually-need-we-need-faster-batch">Most &amp;ldquo;we need real-time&amp;rdquo; requests actually need &amp;ldquo;we need faster batch&amp;rdquo;&lt;/h3>
&lt;br>
&lt;p>Before you reach for Kafka, Kinesis, or any streaming infrastructure, have this conversation with your stakeholders: &amp;ldquo;When you say real-time, what do you actually mean?&amp;rdquo;&lt;/p>
&lt;p>In my experience, the answers fall into three categories:&lt;/p>
&lt;p>&lt;strong>&amp;ldquo;I want data from today, not yesterday.&amp;rdquo;&lt;/strong> That&amp;rsquo;s daily batch with a morning refresh. You already have this. Maybe you need to move the refresh earlier.&lt;/p>
&lt;p>&lt;strong>&amp;ldquo;I want data that&amp;rsquo;s at most an hour old.&amp;rdquo;&lt;/strong> That&amp;rsquo;s micro-batch — running your existing pipeline every 15–60 minutes instead of once a day. No new infrastructure required. Your existing SQL, dbt, and Airflow tools work unchanged at shorter intervals.&lt;/p>
&lt;p>&lt;strong>&amp;ldquo;I need to see events within seconds of them happening.&amp;rdquo;&lt;/strong> &lt;em>This&lt;/em> is actual real-time. And it&amp;rsquo;s genuinely rare. Fraud detection, safety monitoring, real-time bidding — these need streaming. Your internal sales dashboard almost certainly does not.&lt;/p>
&lt;p>For the micro-batch pattern in Airflow, it&amp;rsquo;s a scheduling change:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">with&lt;/span> DAG(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;micro_batch_analytics&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> schedule_interval&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;*/15 * * * *&amp;#39;&lt;/span>, &lt;span style="color:#75715e"># every 15 minutes&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> catchup&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#66d9ef">False&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> max_active_runs&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>, &lt;span style="color:#75715e"># prevent overlap&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>) &lt;span style="color:#66d9ef">as&lt;/span> dag:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#75715e"># ...tasks here&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The &lt;code>max_active_runs=1&lt;/code> is important. Without it, if a run takes longer than 15 minutes, the next run starts before the previous one finishes. They compete for the same warehouse, both run slower, and you&amp;rsquo;ve created the resource contention problem from the shifting-right section.&lt;/p>
&lt;p>In Snowflake, pair this with Snowpipe for continuous S3 ingestion:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Create a pipe for continuous loading from S3
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">CREATE&lt;/span> &lt;span style="color:#66d9ef">OR&lt;/span> &lt;span style="color:#66d9ef">REPLACE&lt;/span> PIPE raw.orders_pipe
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> AUTO_INGEST &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">TRUE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AS&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">COPY&lt;/span> &lt;span style="color:#66d9ef">INTO&lt;/span> raw.orders_stream
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">FROM&lt;/span> &lt;span style="color:#f92672">@&lt;/span>raw.s3_orders_stage
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> FILE_FORMAT &lt;span style="color:#f92672">=&lt;/span> (&lt;span style="color:#66d9ef">TYPE&lt;/span> &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;PARQUET&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> MATCH_BY_COLUMN_NAME &lt;span style="color:#f92672">=&lt;/span> CASE_INSENSITIVE;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Snowpipe loads files within minutes of them landing in S3. Your micro-batch dbt pipeline then transforms whatever arrived since the last run. The combination delivers data freshness measured in minutes — not seconds, but close enough for the vast majority of business use cases — without any streaming infrastructure.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="measure-time-to-consumer-not-just-pipeline-runtime">Measure time-to-consumer, not just pipeline runtime&lt;/h3>
&lt;br>
&lt;p>Pipeline runtime tells you how long your DAG takes. Time-to-consumer tells you how long a business event takes to reach a human decision-maker. These are different numbers, and the second one is what your stakeholders actually care about.&lt;/p>
&lt;p>Time-to-consumer decomposes into stages:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Source extraction delay&lt;/strong>: Time between the event occurring and the data landing in your lake&lt;/li>
&lt;li>&lt;strong>Pipeline processing time&lt;/strong>: Your DAG runtime (the part you control most directly)&lt;/li>
&lt;li>&lt;strong>Warehouse serving time&lt;/strong>: Query execution time when a dashboard or report reads the data&lt;/li>
&lt;li>&lt;strong>Cache/refresh delay&lt;/strong>: How often the BI tool refreshes its cache&lt;/li>
&lt;/ul>
&lt;p>Track it by stamping timestamps at each handoff:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Build a freshness tracking model in dbt
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- mart_data_freshness.sql
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;fct_orders&amp;#39;&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> model_name,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">MAX&lt;/span>(order_timestamp) &lt;span style="color:#66d9ef">AS&lt;/span> latest_source_event,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">MAX&lt;/span>(_loaded_at) &lt;span style="color:#66d9ef">AS&lt;/span> latest_load_time,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>() &lt;span style="color:#66d9ef">AS&lt;/span> measured_at,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> TIMESTAMPDIFF(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;minute&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">MAX&lt;/span>(order_timestamp),
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ) &lt;span style="color:#66d9ef">AS&lt;/span> minutes_since_latest_event,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> TIMESTAMPDIFF(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;minute&amp;#39;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">MAX&lt;/span>(_loaded_at),
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span>()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ) &lt;span style="color:#66d9ef">AS&lt;/span> minutes_since_latest_load
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#960050;background-color:#1e0010">{{&lt;/span> &lt;span style="color:#66d9ef">ref&lt;/span>(&lt;span style="color:#e6db74">&amp;#39;fct_orders&amp;#39;&lt;/span>) &lt;span style="color:#960050;background-color:#1e0010">}}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Schedule this model to run after every pipeline completion. Over time, you&amp;rsquo;ll see whether &lt;code>minutes_since_latest_event&lt;/code> is stable, improving, or drifting. If it&amp;rsquo;s drifting, you can pinpoint which stage is responsible by comparing the event timestamp, load timestamp, and measurement timestamp.&lt;/p>
&lt;p>In dbt, you can also use source freshness tests to alert when upstream data is stale:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># models/staging/_sources.yml&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">sources&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">raw&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">database&lt;/span>: &lt;span style="color:#ae81ff">raw_db&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">schema&lt;/span>: &lt;span style="color:#ae81ff">public&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">loaded_at_field&lt;/span>: &lt;span style="color:#ae81ff">_loaded_at&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">freshness&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">warn_after&lt;/span>: { &lt;span style="color:#f92672">count: 2, period&lt;/span>: &lt;span style="color:#ae81ff">hour }&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">error_after&lt;/span>: { &lt;span style="color:#f92672">count: 6, period&lt;/span>: &lt;span style="color:#ae81ff">hour }&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">tables&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">orders&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">customers&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">events&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Running &lt;code>dbt source freshness&lt;/code> checks whether the upstream data is as fresh as you expect. If &lt;code>raw.orders&lt;/code> hasn&amp;rsquo;t received new rows in 6 hours, the test fails — and that&amp;rsquo;s a signal that the source system&amp;rsquo;s extract is delayed, not that your pipeline is broken.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-view-on-view-problem-the-compound-interest-of-bad-performance">The view-on-view problem: the compound interest of bad performance&lt;/h3>
&lt;br>
&lt;p>This is the architectural anti-pattern I see most often in Snowflake environments, and it&amp;rsquo;s one of those problems that starts small and compounds quietly.&lt;/p>
&lt;p>A view in Snowflake doesn&amp;rsquo;t store any data — it re-executes its SQL every time it&amp;rsquo;s queried. That&amp;rsquo;s fine for a simple view. But when View C references View B, which references View A, which scans a large table — every query against View C triggers the entire chain from scratch.&lt;/p>
&lt;p>I&amp;rsquo;ve seen environments where analysts querying a &amp;ldquo;simple&amp;rdquo; dashboard view were unknowingly triggering five layers of nested views, each with its own JOINs and aggregations. The query took 4 minutes. After materialising the two most expensive intermediate layers as tables (refreshed daily), the same query took 3 seconds.&lt;/p>
&lt;p>The fix is to audit your materialisation strategy:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Find views that reference other views (potential nesting)
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> referencing_object_name &lt;span style="color:#66d9ef">AS&lt;/span> downstream_view,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> referenced_object_name &lt;span style="color:#66d9ef">AS&lt;/span> upstream_object,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> referenced_object_type
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> snowflake.account_usage.object_dependencies
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">WHERE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> referencing_object_type &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;VIEW&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">AND&lt;/span> referenced_object_type &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;VIEW&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">ORDER&lt;/span> &lt;span style="color:#66d9ef">BY&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> downstream_view;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>If this query returns results, you have view-on-view nesting. Not all of it is bad — a simple renaming view on top of a complex view is fine. But if you find three or more layers, or if any of the intermediate views involve heavy JOINs or aggregations, materialise the expensive layers as tables.&lt;/p>
&lt;p>In dbt terms, the decision framework is:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Ephemeral&lt;/strong>: Simple CTEs, column renaming, type casting. Zero cost.&lt;/li>
&lt;li>&lt;strong>View&lt;/strong>: Light transformations queried infrequently by a few consumers.&lt;/li>
&lt;li>&lt;strong>Table&lt;/strong>: Complex transformations queried frequently. Rebuilt every run.&lt;/li>
&lt;li>&lt;strong>Incremental&lt;/strong>: Large tables where only a fraction of rows change. The right choice for most fact tables.&lt;/li>
&lt;/ul>
&lt;p>If a model takes more than 30 seconds to build and is queried more than once between builds, it should be a table or incremental — not a view.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="coming-back-to-priyas-scatter-plot">Coming back to Priya&amp;rsquo;s scatter plot&lt;/h3>
&lt;br>
&lt;p>Six weeks after Priya showed me that graph, we&amp;rsquo;d fixed the three root causes: an intermediate model that had grown from 5 million to 180 million rows without being switched to incremental, a phantom dependency chain that added 12 minutes of serial wait time, and a warehouse that was being shared between transformation and BI queries during the morning refresh window.&lt;/p>
&lt;p>Pipeline completion moved from 7:23 AM back to 5:15 AM. The finance team got their dashboards before their morning coffee. And Priya? She built a monitoring dashboard that tracked completion time drift automatically — the thing I should have built from the start.&lt;/p>
&lt;p>Here&amp;rsquo;s what I learned from that experience, and what I hope you take from this article: &lt;strong>pipeline performance isn&amp;rsquo;t a problem you solve once.&lt;/strong> It&amp;rsquo;s a trajectory you manage. Data grows. Teams add models. Source systems change. The pipeline that&amp;rsquo;s fast today will be slow in six months unless someone is watching the trend line.&lt;/p>
&lt;p>The engineers who build reliable data platforms aren&amp;rsquo;t the ones who build the fastest pipeline on day one. They&amp;rsquo;re the ones who notice when completion time drifts by 3 minutes per week — and fix it before anyone else has to ask.&lt;/p>
&lt;p>Track your completion times. Audit your dependencies quarterly. Profile your critical path after every significant change. And teach your team that a green DAG doesn&amp;rsquo;t mean a healthy pipeline. Sometimes green just means it hasn&amp;rsquo;t failed &lt;em>yet&lt;/em>.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Cloud Architecture</category><category>Snowflake</category><category>AWS</category><category>Airflow</category><category>Pipeline Optimization</category><category>dbt</category><category>Data Freshness</category></item><item><title>Stop Building Salesforce Integrations From Scratch</title><link>https://ghostinthedata.info/posts/2026/2026-04-04-snowflake-openflow/</link><pubDate>Sat, 04 Apr 2026 09:00:00 +1100</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-04-04-snowflake-openflow/</guid><author>Chris Hillman</author><description>A hands-on guide to Snowflake's OpenFlow Salesforce connector — why managed connectors beat custom code, how to set one up step by step, and the schema evolution feature that makes it all worth it.</description><content:encoded>&lt;p>Let me tell you about Marcus.&lt;/p>
&lt;/br>
&lt;p>Marcus was on a team I led a few years back. Sharp, motivated, the kind of engineer who actually read documentation before writing code. When the business asked us to get Salesforce data into our warehouse, Marcus volunteered. He&amp;rsquo;d done API work before. He figured a few weeks, tops.&lt;/p>
&lt;p>He scoped it carefully. Built a Python service that authenticated via OAuth, pulled Account, Contact, and Opportunity objects through the Bulk API, flattened the nested JSON into relational tables, handled pagination, managed rate limits. Wrote solid tests. Documented everything. The kind of work you&amp;rsquo;d point to in a code review and say &lt;em>this is how it&amp;rsquo;s done&lt;/em>.&lt;/p>
&lt;p>It worked beautifully. For about three months.&lt;/p>
&lt;/br>
&lt;p>Then our Salesforce admin added a custom field to the Opportunity object — &lt;code>Renewal_Likelihood__c&lt;/code> — and nobody told the data team. The pipeline didn&amp;rsquo;t fail. That&amp;rsquo;s the insidious part. It kept running, kept landing data. It just quietly dropped the new field on the floor.&lt;/p>
&lt;p>When Marcus tracked it down, he added the field. Then he realised there were &lt;em>eleven&lt;/em> other custom fields that had been added since go-live that the pipeline was silently ignoring. And the compound address fields — &lt;code>BillingAddress&lt;/code>, &lt;code>ShippingAddress&lt;/code> — had never worked with the Bulk API in the first place. He&amp;rsquo;d been extracting the component parts (&lt;code>BillingStreet&lt;/code>, &lt;code>BillingCity&lt;/code>) as a workaround, but the workaround had a bug that truncated postal codes for international addresses.&lt;/p>
&lt;p>Marcus spent three weeks patching all of it. Three weeks of a talented engineer doing work that a managed connector handles automatically.&lt;/p>
&lt;/br>
&lt;p>When I watched a good engineer spend the best part of a month on a problem that shouldn&amp;rsquo;t exist. Marcus wasn&amp;rsquo;t learning anything. He wasn&amp;rsquo;t growing. He was hand-coding schema detection logic for the fourth time because Salesforce ships three major releases a year and our sales ops team adds custom fields like they&amp;rsquo;re decorating a Christmas tree.&lt;/p>
&lt;p>I&amp;rsquo;m telling you this because &lt;strong>Snowflake OpenFlow&lt;/strong> is one of those things that sounds like &amp;ldquo;just another integration tool&amp;rdquo; until you&amp;rsquo;ve lived through the alternative. What I want to show you today is how to set it up with Salesforce — step by step, with enough detail that you could do it tomorrow — and more importantly, &lt;em>why&lt;/em> the schema evolution feature alone justifies the switch from custom code.&lt;/p>
&lt;p>If you&amp;rsquo;ve ever maintained a custom Salesforce integration, you already know why this matters.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-problem-nobody-warns-you-about">The Problem Nobody Warns You About&lt;/h3>
&lt;/br>
&lt;p>Here&amp;rsquo;s what the &amp;ldquo;just build it yourself&amp;rdquo; crowd doesn&amp;rsquo;t tell you: the Salesforce API isn&amp;rsquo;t one API. It&amp;rsquo;s an ecosystem of overlapping interfaces, each with its own quirks, limits, and failure modes.&lt;/p>
&lt;p>The REST API handles small, synchronous requests well — but try pulling a million Opportunity records through it and you&amp;rsquo;ll burn through your rate limits before lunch. The Bulk API 2.0 is designed for volume — it can handle up to 150 million records in a 24-hour window — but it doesn&amp;rsquo;t support compound fields like &lt;code>BillingAddress&lt;/code> or &lt;code>MailingAddress&lt;/code>. Those silently return nothing. Not an error. Nothing.&lt;/p>
&lt;p>Then there&amp;rsquo;s SOQL, Salesforce&amp;rsquo;s proprietary query language, which looks enough like SQL to trick you into thinking you understand it. Until you hit the 2,000-record offset limit, or try to resolve a polymorphic relationship where a &lt;code>WhoId&lt;/code> field on a Task could reference either a Contact or a Lead depending on the record. That&amp;rsquo;s not a data modelling problem — that&amp;rsquo;s a &amp;ldquo;your pipeline logic needs to branch based on the referenced object type&amp;rdquo; problem.&lt;/p>
&lt;p>And rate limits. Enterprise Edition gives you 100,000 API calls per 24 hours, plus 1,000 per user licence. Sounds generous until you remember that your BI tool, marketing automation platform, customer support system, and your data pipeline are all drinking from the same well.&lt;/p>
&lt;/br>
&lt;p>But schema evolution is the silent killer.&lt;/p>
&lt;p>Every Salesforce org is different because every business is different. Your Salesforce admin creates custom objects and fields to match how your company actually works. That&amp;rsquo;s the point — it&amp;rsquo;s a customisation platform. But every custom field added to Salesforce is a potential break point for any integration that doesn&amp;rsquo;t handle schema changes automatically.&lt;/p>
&lt;p>And those changes happen constantly. Research suggests schemas in enterprise SaaS tools change roughly every three days on average. In Salesforce specifically, between admin-driven customisation and Salesforce&amp;rsquo;s own three-times-a-year release cycle (Spring, Summer, Winter), the schema you built against last quarter may not be the schema you&amp;rsquo;re ingesting from today.&lt;/p>
&lt;p>This is the trap. Building the initial integration isn&amp;rsquo;t hard. Any competent engineer can get Salesforce data into a warehouse. &lt;strong>Keeping it working&lt;/strong> — across schema changes, API version deprecations, authentication token rotations, and rate limit adjustments — is where it eats your life.&lt;/p>
&lt;p>I&amp;rsquo;ve seen teams eliminate two full days of monthly engineering maintenance by migrating off custom connectors. Two days, every month, spent on work that a managed tool does automatically. That&amp;rsquo;s not engineering — it&amp;rsquo;s janitorial work wearing an engineering costume.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="what-is-snowflake-openflow">What Is Snowflake OpenFlow?&lt;/h3>
&lt;/br>
&lt;p>OpenFlow is Snowflake&amp;rsquo;s native, fully managed data integration service. It&amp;rsquo;s not a partnership or a marketplace app — it&amp;rsquo;s a first-party feature built into the platform. If you&amp;rsquo;re already running Snowflake, OpenFlow runs inside your existing environment.&lt;/p>
&lt;p>The backstory matters. In November 2024, Snowflake acquired a company called &lt;strong>Datavolo&lt;/strong> for roughly $110 million. Datavolo was founded by Joe Witt, who co-created Apache NiFi back when it was an NSA project. That NiFi heritage is the foundation of OpenFlow — curated, versioned connector definitions powered by NiFi&amp;rsquo;s flow engine, wrapped in a managed Snowflake experience.&lt;/p>
&lt;p>Architecturally, it splits into two pieces. The &lt;strong>control plane&lt;/strong> lives inside Snowflake and gives you pipeline management, scheduling, and monitoring through the Snowsight UI. The &lt;strong>data plane&lt;/strong> handles the actual work and can run in two modes: &lt;strong>Snowflake Deployments&lt;/strong> (running on Snowpark Container Services within Snowflake&amp;rsquo;s own infrastructure) or &lt;strong>BYOC&lt;/strong> (Bring Your Own Cloud, running as a Kubernetes cluster in your VPC). For most teams starting out, the SPCS option is simpler — it went generally available in November 2025 on AWS and Azure.&lt;/p>
&lt;p>Data flows into Snowflake primarily through Snowpipe Streaming for the initial load, with merge queries handling incremental updates. The result is structured tables in your Snowflake database — not staged JSON files, not semi-structured VARIANT columns. Actual, queryable tables with proper column types.&lt;/p>
&lt;/br>
&lt;p>The connector catalogue currently covers about 20 curated sources: databases (PostgreSQL, MySQL, SQL Server, Oracle), SaaS platforms (Salesforce, Workday, Jira, LinkedIn Ads, Meta Ads, Google Ads), cloud storage (Google Drive, SharePoint, Box), and streaming sources (Kafka, Kinesis). It&amp;rsquo;s not Fivetran&amp;rsquo;s 500+ connector library, and nobody&amp;rsquo;s pretending it is. But the sources it does cover are the ones that account for the vast majority of enterprise data movement.&lt;/p>
&lt;p>One important caveat up front: the &lt;strong>Salesforce Bulk API connector is still in preview&lt;/strong> as of March 2026. It works. Preview in Snowflake terms means the feature is functional but not yet covered by production support SLAs. Keep that in your mental model as we walk through the setup.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="setting-up-the-salesforce-connector-a-walkthrough">Setting Up the Salesforce Connector: A Walkthrough&lt;/h3>
&lt;/br>
&lt;p>This section is the step-by-step. I&amp;rsquo;m going to walk through the full setup — Salesforce configuration, Snowflake preparation, and connector deployment — with enough detail that you could follow along on your own org.&lt;/p>
&lt;/br>
&lt;h4 id="phase-1-salesforce-configuration">Phase 1: Salesforce Configuration&lt;/h4>
&lt;p>The connector uses JWT Bearer Flow for authentication, which means you need an RSA key pair and a Connected App configured in Salesforce.&lt;/p>
&lt;p>&lt;strong>Generate your RSA key pair.&lt;/strong> Open a terminal and run:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Generate 2048-bit RSA private key&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>openssl genrsa -out salesforce_private_key.pem &lt;span style="color:#ae81ff">2048&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Generate the corresponding public certificate (valid 1 year)&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>openssl req -new -x509 -key salesforce_private_key.pem &lt;span style="color:#ae81ff">\
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#ae81ff">&lt;/span> -out salesforce_certificate.pem -days &lt;span style="color:#ae81ff">365&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Keep that private key safe. You&amp;rsquo;ll need it when configuring the connector in Snowflake.&lt;/p>
&lt;p>&lt;strong>Create the External Client App in Salesforce.&lt;/strong> Navigate to Setup → Apps → App Manager → New External Client App. Enable OAuth and add two scopes: &lt;code>api&lt;/code> and &lt;code>refresh_token&lt;/code> (sometimes labelled &lt;code>offline_access&lt;/code>). Enable the JWT Bearer Flow toggle and upload your &lt;code>salesforce_certificate.pem&lt;/code> public certificate. Save the app and record the &lt;strong>Consumer Key&lt;/strong> and &lt;strong>Consumer Secret&lt;/strong> that Salesforce generates.&lt;/p>
&lt;p>&lt;strong>Configure access.&lt;/strong> On the app&amp;rsquo;s Manage settings, set the Permitted Users policy to &amp;ldquo;Admin approved users are pre-authorized.&amp;rdquo; Then assign the appropriate user profiles or permission sets. The user account the connector will authenticate as needs read access to every object you intend to sync.&lt;/p>
&lt;/br>
&lt;h4 id="phase-2-snowflake-preparation">Phase 2: Snowflake Preparation&lt;/h4>
&lt;p>Before deploying the connector, you need a database, schema, role, and warehouse ready for it.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Create a dedicated database for Salesforce data
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">CREATE&lt;/span> &lt;span style="color:#66d9ef">DATABASE&lt;/span> &lt;span style="color:#66d9ef">IF&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">EXISTS&lt;/span> SALESFORCE_RAW;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">CREATE&lt;/span> &lt;span style="color:#66d9ef">SCHEMA&lt;/span> &lt;span style="color:#66d9ef">IF&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">EXISTS&lt;/span> SALESFORCE_RAW.OPENFLOW;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Create a connector-specific role
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">CREATE&lt;/span> &lt;span style="color:#66d9ef">ROLE&lt;/span> &lt;span style="color:#66d9ef">IF&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">EXISTS&lt;/span> OPENFLOW_SF_CONNECTOR;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Grant the role permissions to write to the destination
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">GRANT&lt;/span> &lt;span style="color:#66d9ef">USAGE&lt;/span> &lt;span style="color:#66d9ef">ON&lt;/span> &lt;span style="color:#66d9ef">DATABASE&lt;/span> SALESFORCE_RAW &lt;span style="color:#66d9ef">TO&lt;/span> &lt;span style="color:#66d9ef">ROLE&lt;/span> OPENFLOW_SF_CONNECTOR;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GRANT&lt;/span> &lt;span style="color:#66d9ef">USAGE&lt;/span> &lt;span style="color:#66d9ef">ON&lt;/span> &lt;span style="color:#66d9ef">SCHEMA&lt;/span> SALESFORCE_RAW.OPENFLOW &lt;span style="color:#66d9ef">TO&lt;/span> &lt;span style="color:#66d9ef">ROLE&lt;/span> OPENFLOW_SF_CONNECTOR;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GRANT&lt;/span> &lt;span style="color:#66d9ef">CREATE&lt;/span> &lt;span style="color:#66d9ef">TABLE&lt;/span> &lt;span style="color:#66d9ef">ON&lt;/span> &lt;span style="color:#66d9ef">SCHEMA&lt;/span> SALESFORCE_RAW.OPENFLOW &lt;span style="color:#66d9ef">TO&lt;/span> &lt;span style="color:#66d9ef">ROLE&lt;/span> OPENFLOW_SF_CONNECTOR;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Create a warehouse for merge operations
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">CREATE&lt;/span> WAREHOUSE &lt;span style="color:#66d9ef">IF&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">EXISTS&lt;/span> OPENFLOW_WH
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> WAREHOUSE_SIZE &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">&amp;#39;SMALL&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> AUTO_SUSPEND &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">60&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> AUTO_RESUME &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">TRUE&lt;/span>;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GRANT&lt;/span> &lt;span style="color:#66d9ef">USAGE&lt;/span> &lt;span style="color:#66d9ef">ON&lt;/span> WAREHOUSE OPENFLOW_WH &lt;span style="color:#66d9ef">TO&lt;/span> &lt;span style="color:#66d9ef">ROLE&lt;/span> OPENFLOW_SF_CONNECTOR;
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">GRANT&lt;/span> OPERATE &lt;span style="color:#66d9ef">ON&lt;/span> WAREHOUSE OPENFLOW_WH &lt;span style="color:#66d9ef">TO&lt;/span> &lt;span style="color:#66d9ef">ROLE&lt;/span> OPENFLOW_SF_CONNECTOR;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>For SPCS deployments, you also need a &lt;strong>network rule&lt;/strong> allowing the connector to reach your Salesforce instance. OpenFlow runs inside Snowflake&amp;rsquo;s container environment, and by default it can&amp;rsquo;t reach external endpoints without explicit permission:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Allow egress to your Salesforce instance
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">CREATE&lt;/span> NETWORK &lt;span style="color:#66d9ef">RULE&lt;/span> &lt;span style="color:#66d9ef">IF&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">EXISTS&lt;/span> OPENFLOW_SF_EGRESS
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">TYPE&lt;/span> &lt;span style="color:#f92672">=&lt;/span> HOST_PORT
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">MODE&lt;/span> &lt;span style="color:#f92672">=&lt;/span> EGRESS
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> VALUE_LIST &lt;span style="color:#f92672">=&lt;/span> (&lt;span style="color:#e6db74">&amp;#39;yourinstance.salesforce.com:443&amp;#39;&lt;/span>);
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Wrap it in an external access integration
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">CREATE&lt;/span> &lt;span style="color:#66d9ef">EXTERNAL&lt;/span> &lt;span style="color:#66d9ef">ACCESS&lt;/span> INTEGRATION &lt;span style="color:#66d9ef">IF&lt;/span> &lt;span style="color:#66d9ef">NOT&lt;/span> &lt;span style="color:#66d9ef">EXISTS&lt;/span> OPENFLOW_SF_ACCESS
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ALLOWED_NETWORK_RULES &lt;span style="color:#f92672">=&lt;/span> (OPENFLOW_SF_EGRESS)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ENABLED &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#66d9ef">TRUE&lt;/span>;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Replace &lt;code>yourinstance.salesforce.com&lt;/code> with your actual Salesforce instance hostname. If you&amp;rsquo;re not sure what that is, log in to Salesforce and look at the URL bar — it&amp;rsquo;s the domain before &lt;code>.salesforce.com&lt;/code>.&lt;/p>
&lt;/br>
&lt;h4 id="phase-3-deploying-the-connector">Phase 3: Deploying the Connector&lt;/h4>
&lt;p>Now the fun part. Navigate to &lt;strong>Snowsight → Data → Openflow&lt;/strong> in the left sidebar. Find the &amp;ldquo;Openflow connector for Salesforce Bulk API&amp;rdquo; tile and add it to your runtime.&lt;/p>
&lt;p>This opens the NiFi canvas — a visual flow editor where the connector&amp;rsquo;s pre-built process groups appear. Right-click the Salesforce process group and select &lt;strong>Configure Parameters&lt;/strong>. Here&amp;rsquo;s where you&amp;rsquo;ll enter:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Salesforce Instance URL&lt;/strong> — e.g., &lt;code>https://yourinstance.salesforce.com&lt;/code>&lt;/li>
&lt;li>&lt;strong>Consumer Key&lt;/strong> — from the Connected App you created&lt;/li>
&lt;li>&lt;strong>Consumer Secret&lt;/strong> — same source&lt;/li>
&lt;li>&lt;strong>Private Key&lt;/strong> — the contents of &lt;code>salesforce_private_key.pem&lt;/code> (the full PEM text, headers included)&lt;/li>
&lt;li>&lt;strong>Username&lt;/strong> — the Salesforce user the connector authenticates as&lt;/li>
&lt;li>&lt;strong>Filter&lt;/strong> — the objects to sync, by API name&lt;/li>
&lt;/ul>
&lt;p>The &lt;strong>Filter&lt;/strong> parameter is where you define what to pull. List standard objects by name (&lt;code>Account&lt;/code>, &lt;code>Contact&lt;/code>, &lt;code>Opportunity&lt;/code>, &lt;code>Lead&lt;/code>, &lt;code>Task&lt;/code>, &lt;code>Event&lt;/code>) and custom objects using their API name with the &lt;code>__c&lt;/code> suffix:&lt;/p>



&lt;div class="goat svg-container ">
 
 &lt;svg
 xmlns="http://www.w3.org/2000/svg"
 font-family="Menlo,Lucida Console,monospace"
 
 viewBox="0 0 720 25"
 >
 &lt;g transform='translate(8,16)'>
&lt;text text-anchor='middle' x='0' y='4' fill='currentColor' style='font-size:1em'>A&lt;/text>
&lt;text text-anchor='middle' x='8' y='4' fill='currentColor' style='font-size:1em'>c&lt;/text>
&lt;text text-anchor='middle' x='16' y='4' fill='currentColor' style='font-size:1em'>c&lt;/text>
&lt;text text-anchor='middle' x='24' y='4' fill='currentColor' style='font-size:1em'>o&lt;/text>
&lt;text text-anchor='middle' x='32' y='4' fill='currentColor' style='font-size:1em'>u&lt;/text>
&lt;text text-anchor='middle' x='40' y='4' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='48' y='4' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='56' y='4' fill='currentColor' style='font-size:1em'>,&lt;/text>
&lt;text text-anchor='middle' x='72' y='4' fill='currentColor' style='font-size:1em'>C&lt;/text>
&lt;text text-anchor='middle' x='80' y='4' fill='currentColor' style='font-size:1em'>o&lt;/text>
&lt;text text-anchor='middle' x='88' y='4' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='96' y='4' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='104' y='4' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='112' y='4' fill='currentColor' style='font-size:1em'>c&lt;/text>
&lt;text text-anchor='middle' x='120' y='4' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='128' y='4' fill='currentColor' style='font-size:1em'>,&lt;/text>
&lt;text text-anchor='middle' x='144' y='4' fill='currentColor' style='font-size:1em'>O&lt;/text>
&lt;text text-anchor='middle' x='152' y='4' fill='currentColor' style='font-size:1em'>p&lt;/text>
&lt;text text-anchor='middle' x='160' y='4' fill='currentColor' style='font-size:1em'>p&lt;/text>
&lt;text text-anchor='middle' x='168' y='4' fill='currentColor' style='font-size:1em'>o&lt;/text>
&lt;text text-anchor='middle' x='176' y='4' fill='currentColor' style='font-size:1em'>r&lt;/text>
&lt;text text-anchor='middle' x='184' y='4' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='192' y='4' fill='currentColor' style='font-size:1em'>u&lt;/text>
&lt;text text-anchor='middle' x='200' y='4' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='208' y='4' fill='currentColor' style='font-size:1em'>i&lt;/text>
&lt;text text-anchor='middle' x='216' y='4' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='224' y='4' fill='currentColor' style='font-size:1em'>y&lt;/text>
&lt;text text-anchor='middle' x='232' y='4' fill='currentColor' style='font-size:1em'>,&lt;/text>
&lt;text text-anchor='middle' x='248' y='4' fill='currentColor' style='font-size:1em'>L&lt;/text>
&lt;text text-anchor='middle' x='256' y='4' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='264' y='4' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='272' y='4' fill='currentColor' style='font-size:1em'>d&lt;/text>
&lt;text text-anchor='middle' x='280' y='4' fill='currentColor' style='font-size:1em'>,&lt;/text>
&lt;text text-anchor='middle' x='296' y='4' fill='currentColor' style='font-size:1em'>T&lt;/text>
&lt;text text-anchor='middle' x='304' y='4' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='312' y='4' fill='currentColor' style='font-size:1em'>s&lt;/text>
&lt;text text-anchor='middle' x='320' y='4' fill='currentColor' style='font-size:1em'>k&lt;/text>
&lt;text text-anchor='middle' x='328' y='4' fill='currentColor' style='font-size:1em'>,&lt;/text>
&lt;text text-anchor='middle' x='344' y='4' fill='currentColor' style='font-size:1em'>E&lt;/text>
&lt;text text-anchor='middle' x='352' y='4' fill='currentColor' style='font-size:1em'>v&lt;/text>
&lt;text text-anchor='middle' x='360' y='4' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='368' y='4' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='376' y='4' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='384' y='4' fill='currentColor' style='font-size:1em'>,&lt;/text>
&lt;text text-anchor='middle' x='400' y='4' fill='currentColor' style='font-size:1em'>C&lt;/text>
&lt;text text-anchor='middle' x='408' y='4' fill='currentColor' style='font-size:1em'>u&lt;/text>
&lt;text text-anchor='middle' x='416' y='4' fill='currentColor' style='font-size:1em'>s&lt;/text>
&lt;text text-anchor='middle' x='424' y='4' fill='currentColor' style='font-size:1em'>t&lt;/text>
&lt;text text-anchor='middle' x='432' y='4' fill='currentColor' style='font-size:1em'>o&lt;/text>
&lt;text text-anchor='middle' x='440' y='4' fill='currentColor' style='font-size:1em'>m&lt;/text>
&lt;text text-anchor='middle' x='448' y='4' fill='currentColor' style='font-size:1em'>_&lt;/text>
&lt;text text-anchor='middle' x='456' y='4' fill='currentColor' style='font-size:1em'>P&lt;/text>
&lt;text text-anchor='middle' x='464' y='4' fill='currentColor' style='font-size:1em'>i&lt;/text>
&lt;text text-anchor='middle' x='472' y='4' fill='currentColor' style='font-size:1em'>p&lt;/text>
&lt;text text-anchor='middle' x='480' y='4' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='488' y='4' fill='currentColor' style='font-size:1em'>l&lt;/text>
&lt;text text-anchor='middle' x='496' y='4' fill='currentColor' style='font-size:1em'>i&lt;/text>
&lt;text text-anchor='middle' x='504' y='4' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='512' y='4' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='520' y='4' fill='currentColor' style='font-size:1em'>_&lt;/text>
&lt;text text-anchor='middle' x='528' y='4' fill='currentColor' style='font-size:1em'>_&lt;/text>
&lt;text text-anchor='middle' x='536' y='4' fill='currentColor' style='font-size:1em'>c&lt;/text>
&lt;text text-anchor='middle' x='544' y='4' fill='currentColor' style='font-size:1em'>,&lt;/text>
&lt;text text-anchor='middle' x='560' y='4' fill='currentColor' style='font-size:1em'>R&lt;/text>
&lt;text text-anchor='middle' x='568' y='4' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='576' y='4' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='584' y='4' fill='currentColor' style='font-size:1em'>e&lt;/text>
&lt;text text-anchor='middle' x='592' y='4' fill='currentColor' style='font-size:1em'>w&lt;/text>
&lt;text text-anchor='middle' x='600' y='4' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='608' y='4' fill='currentColor' style='font-size:1em'>l&lt;/text>
&lt;text text-anchor='middle' x='616' y='4' fill='currentColor' style='font-size:1em'>_&lt;/text>
&lt;text text-anchor='middle' x='624' y='4' fill='currentColor' style='font-size:1em'>T&lt;/text>
&lt;text text-anchor='middle' x='632' y='4' fill='currentColor' style='font-size:1em'>r&lt;/text>
&lt;text text-anchor='middle' x='640' y='4' fill='currentColor' style='font-size:1em'>a&lt;/text>
&lt;text text-anchor='middle' x='648' y='4' fill='currentColor' style='font-size:1em'>c&lt;/text>
&lt;text text-anchor='middle' x='656' y='4' fill='currentColor' style='font-size:1em'>k&lt;/text>
&lt;text text-anchor='middle' x='664' y='4' fill='currentColor' style='font-size:1em'>i&lt;/text>
&lt;text text-anchor='middle' x='672' y='4' fill='currentColor' style='font-size:1em'>n&lt;/text>
&lt;text text-anchor='middle' x='680' y='4' fill='currentColor' style='font-size:1em'>g&lt;/text>
&lt;text text-anchor='middle' x='688' y='4' fill='currentColor' style='font-size:1em'>_&lt;/text>
&lt;text text-anchor='middle' x='696' y='4' fill='currentColor' style='font-size:1em'>_&lt;/text>
&lt;text text-anchor='middle' x='704' y='4' fill='currentColor' style='font-size:1em'>c&lt;/text>
&lt;/g>

 &lt;/svg>
 
&lt;/div>
&lt;p>There&amp;rsquo;s also a &lt;strong>Special Objects Filter&lt;/strong> for objects that require different handling (like &lt;code>User&lt;/code> or &lt;code>RecordType&lt;/code>).&lt;/p>
&lt;p>Once the parameters are configured, enable all controller services in the process group, then start the connector. The initial sync will pull a full snapshot of every object in your filter list. Subsequent runs perform incremental syncs using Salesforce&amp;rsquo;s &lt;code>SystemModstamp&lt;/code> field to identify changed records.&lt;/p>
&lt;/br>
&lt;h4 id="what-you-get">What You Get&lt;/h4>
&lt;p>Once the connector runs, you&amp;rsquo;ll find one table per Salesforce object in &lt;code>SALESFORCE_RAW.OPENFLOW&lt;/code>. The tables are &lt;strong>flattened and structured&lt;/strong> — every Salesforce field becomes a column with an appropriate Snowflake data type. No VARIANT columns to parse. No nested JSON to unpack.&lt;/p>
&lt;p>An &lt;code>Account&lt;/code> table will have columns like &lt;code>ID&lt;/code>, &lt;code>NAME&lt;/code>, &lt;code>INDUSTRY&lt;/code>, &lt;code>ANNUALREVENUE&lt;/code>, &lt;code>BILLINGSTREET&lt;/code>, &lt;code>BILLINGCITY&lt;/code>, &lt;code>BILLINGSTATE&lt;/code>, &lt;code>BILLINGPOSTALCODE&lt;/code>, &lt;code>CREATEDDATE&lt;/code>, &lt;code>LASTMODIFIEDDATE&lt;/code>, and so on. Custom fields like &lt;code>Renewal_Likelihood__c&lt;/code> appear as &lt;code>RENEWAL_LIKELIHOOD__C&lt;/code>.&lt;/p>
&lt;p>Two metadata columns are added automatically: &lt;code>ISDELETED&lt;/code> (tracking Salesforce soft deletes) and the system timestamp fields. This is important — &lt;strong>hard deletes in Salesforce are not reflected&lt;/strong>. If a record is permanently deleted in Salesforce, it won&amp;rsquo;t disappear from your Snowflake table. You&amp;rsquo;ll need to handle that in your downstream dbt models if it matters for your use case.&lt;/p>
&lt;/br>
&lt;h4 id="a-note-on-multiple-sync-frequencies">A Note on Multiple Sync Frequencies&lt;/h4>
&lt;p>Not every Salesforce object changes at the same pace. Your &lt;code>Opportunity&lt;/code> table might need hourly syncs while &lt;code>Account&lt;/code> is fine with daily. OpenFlow handles this by letting you deploy &lt;strong>multiple connector instances in the same runtime&lt;/strong> — each with its own filter list and schedule — at no additional compute cost beyond the shared runtime.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="schema-evolution-where-openflow-earns-its-keep">Schema Evolution: Where OpenFlow Earns Its Keep&lt;/h3>
&lt;/br>
&lt;p>This is the feature that matters most, and the one that would have saved Marcus three weeks.&lt;/p>
&lt;p>When someone adds a new custom field to a Salesforce object — say the sales ops team creates &lt;code>Deal_Confidence_Score__c&lt;/code> on the Opportunity object — the OpenFlow connector &lt;strong>automatically detects the new field and adds a corresponding column&lt;/strong> to the Snowflake destination table on the next sync. No configuration change. No redeployment. No Teams message at 7am asking the data team to &amp;ldquo;add that new field we made yesterday.&amp;rdquo;&lt;/p>
&lt;p>The column appears, data starts flowing into it, and your dbt models can reference it whenever you&amp;rsquo;re ready.&lt;/p>
&lt;/br>
&lt;p>For field renames, OpenFlow takes a conservative approach: the old column stays in place (with stale data from the last sync before the rename), and a new column is created under the new field name. This means you&amp;rsquo;ll have both &lt;code>Old_Field_Name__c&lt;/code> and &lt;code>New_Field_Name__c&lt;/code> in your table for a period. That&amp;rsquo;s actually useful for audit purposes — you can see exactly when the rename happened by comparing timestamps — but it does mean your downstream queries need updating.&lt;/p>
&lt;p>Here&amp;rsquo;s the honest gap: &lt;strong>new custom objects are not auto-discovered&lt;/strong>. If someone creates an entirely new custom object in Salesforce, you have to manually add it to the connector&amp;rsquo;s Filter parameter. Schema evolution handles field-level changes automatically, but object-level discovery is still a manual step.&lt;/p>
&lt;p>And field type changes — say someone changes a text field to a picklist, or a number to a formula — aren&amp;rsquo;t comprehensively documented in the current OpenFlow materials. In practice, this is a rare enough occurrence that it hasn&amp;rsquo;t been a problem in my experience, but it&amp;rsquo;s worth knowing the edge case exists.&lt;/p>
&lt;/br>
&lt;p>Compare this to what Marcus had to maintain. His custom integration used a hardcoded list of fields per object. Every time a field was added, renamed, or retyped, someone had to update the Python code, test it, deploy it, and verify the data. Usually that &amp;ldquo;someone&amp;rdquo; was Marcus, usually at an inconvenient time, and usually because nobody told him the change was coming.&lt;/p>
&lt;p>The value of automatic schema evolution isn&amp;rsquo;t technical elegance. It&amp;rsquo;s that &lt;strong>your engineers stop spending time on schema babysitting and start spending it on work that actually matters&lt;/strong> — building models, improving data quality, answering the questions the business is actually asking.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="a-word-about-fivetran">A Word About Fivetran&lt;/h3>
&lt;/br>
&lt;p>I&amp;rsquo;d be doing you a disservice if I wrote about managed Salesforce connectors without mentioning Fivetran. For most of the last decade, Fivetran has been the gold standard here, and their Salesforce connector is genuinely excellent.&lt;/p>
&lt;p>Fivetran&amp;rsquo;s schema evolution handling is more mature and more configurable than OpenFlow&amp;rsquo;s current offering. They offer three schema change policies: &lt;code>ALLOW_ALL&lt;/code> (sync everything automatically, including new tables), &lt;code>ALLOW_COLUMNS&lt;/code> (auto-add new columns but not new tables), and &lt;code>BLOCK_ALL&lt;/code> (manual opt-in for everything). They also generate a &lt;code>LOG&lt;/code> table that records every schema change event — &lt;code>create_table&lt;/code>, &lt;code>alter_table&lt;/code>, &lt;code>drop_table&lt;/code> — giving you an audit trail that&amp;rsquo;s useful for debugging and compliance. Field type widening (like &lt;code>INT&lt;/code> to &lt;code>LONG&lt;/code>) happens automatically; narrowing changes are blocked to prevent data loss.&lt;/p>
&lt;p>Their pricing model is fundamentally different from OpenFlow. Fivetran charges per Monthly Active Row (MAR) — the count of distinct rows synced per connection per month. A unique primary key counts once regardless of how many times it&amp;rsquo;s updated. Approximate pricing clusters around $500 per million MAR on the Standard plan, with a $12,000 annual minimum commitment. Since early 2025, MAR is calculated per connection rather than account-wide, which eliminated the bulk discount that previously benefited multi-connector setups.&lt;/p>
&lt;/br>
&lt;p>OpenFlow, by contrast, charges for compute — SPCS credits for the container runtime, Snowpipe Streaming costs for ingestion, and warehouse costs for merge operations. There&amp;rsquo;s no per-row fee and no separate licence. It&amp;rsquo;s included with Snowflake. The trade-off is that your management compute pool runs continuously while a deployment exists, which means you&amp;rsquo;re paying a base cost even when no data is flowing. One estimate I&amp;rsquo;ve seen puts idle costs at roughly $10 per day, though this varies with configuration.&lt;/p>
&lt;p>For a single large Salesforce org, OpenFlow&amp;rsquo;s compute-based model may end up cheaper than Fivetran&amp;rsquo;s per-row pricing — especially if your Salesforce data doesn&amp;rsquo;t churn heavily. For environments where you need 30+ connectors across different source types, Fivetran&amp;rsquo;s breadth and maturity is hard to argue with. They have 500+ connectors. OpenFlow has about 20.&lt;/p>
&lt;p>The honest framing: &lt;strong>OpenFlow is the better choice if you&amp;rsquo;re already deeply invested in Snowflake, value platform consolidation, and your primary sources are in the current connector catalogue.&lt;/strong> Fivetran is the safer choice if you need broad connector coverage, battle-tested reliability, and you&amp;rsquo;d rather pay a premium for someone else to handle the ops.&lt;/p>
&lt;p>They&amp;rsquo;re not mutually exclusive, either. I&amp;rsquo;ve seen teams run Fivetran for the long tail of small connectors while using OpenFlow for their highest-volume sources where compute-based pricing wins.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-zero-copy-alternative">The Zero-Copy Alternative&lt;/h3>
&lt;/br>
&lt;p>There&amp;rsquo;s a third path I haven&amp;rsquo;t mentioned yet, and it&amp;rsquo;s worth understanding even if it&amp;rsquo;s not right for every team: &lt;strong>zero-copy data federation&lt;/strong>.&lt;/p>
&lt;p>Salesforce Data Cloud offers a zero-copy integration with Snowflake that flips the entire model on its head. Instead of extracting data from Salesforce and loading it into your warehouse, zero-copy gives you direct read access to Salesforce data &lt;em>without moving it at all&lt;/em>. Salesforce&amp;rsquo;s query pushdown engine handles the work — when you query the federated data, Salesforce pushes the query to the source, filters and aggregates there, and returns only the results you need. No pipelines to build. No schemas to manage. No sync schedules to configure.&lt;/p>
&lt;p>It works bidirectionally, too. You can share data from Snowflake back into Salesforce Data Cloud, letting sales reps see warehouse-enriched signals directly on Account and Opportunity pages without anyone exporting a CSV or building a reverse ETL pipeline. The integration runs on Apache Iceberg under the hood, which at least means the underlying format is open.&lt;/p>
&lt;/br>
&lt;p>For handling new custom objects — the gap I flagged with OpenFlow — zero-copy sidesteps the problem entirely. There&amp;rsquo;s nothing to discover because there&amp;rsquo;s nothing to sync. The data stays where it is, and you access it in place. As your Salesforce org evolves, the federated view evolves with it.&lt;/p>
&lt;p>That&amp;rsquo;s genuinely appealing. For proof-of-concept work, rapid prototyping, or use cases where you need real-time access to Salesforce data without building infrastructure, zero-copy is fast and effective.&lt;/p>
&lt;/br>
&lt;p>But here&amp;rsquo;s where I get cautious.&lt;/p>
&lt;p>&lt;strong>It comes at a premium.&lt;/strong> Salesforce Data Cloud runs on a consumption credit model — credits are purchased in bundles (roughly $500 per 100,000 credits at list price), and every action burns them. Zero-copy federation queries consume about 70 credits per million records, which sounds reasonable until you&amp;rsquo;re running it at scale across multiple business units. The pricing model has simplified since late 2025 — Salesforce consolidated multiple credit types into a single fungible credit and made ingestion from core Salesforce products free — but it&amp;rsquo;s still a consumption model with real costs that compound quickly if you&amp;rsquo;re not watching.&lt;/p>
&lt;p>More importantly, &lt;strong>zero-copy creates a deep platform dependency&lt;/strong>. Your data access is mediated entirely through Salesforce and Snowflake&amp;rsquo;s partnership. If Salesforce changes their terms, adjusts their pricing (which they&amp;rsquo;ve done repeatedly), or if you decide to move off Snowflake to Databricks or BigQuery, that zero-copy integration doesn&amp;rsquo;t come with you. You&amp;rsquo;d need to build actual data movement infrastructure — the thing you avoided — under time pressure and with no existing pipeline to fall back on.&lt;/p>
&lt;/br>
&lt;p>I value keeping infrastructure portable. Not because I&amp;rsquo;m paranoid about vendor lock-in — I&amp;rsquo;m pragmatic about it. I&amp;rsquo;ve been in this industry long enough to see what happens when a new CEO arrives and wants to renegotiate every vendor contract, or when a cloud provider changes their pricing structure overnight, or when the company you depend on gets acquired and the roadmap shifts. Having your data physically in your own warehouse, in formats you control, with pipelines you can redirect to a different destination — that&amp;rsquo;s insurance. Not glamorous insurance, but the kind you&amp;rsquo;re grateful for when you need it.&lt;/p>
&lt;p>The balance is real, though. You always have to weigh portability against the time you spend building and maintaining pipelines. If your team is drowning in pipeline maintenance, zero-copy might buy you breathing room while you stand up proper infrastructure. Just go in with your eyes open about what you&amp;rsquo;re trading for that convenience.&lt;/p>
&lt;/br>
&lt;p>There&amp;rsquo;s a more fundamental concern, though, and it&amp;rsquo;s the one that won&amp;rsquo;t appear in any vendor comparison chart.&lt;/p>
&lt;p>&lt;strong>Zero-copy makes it dangerously easy to skip the data warehouse entirely.&lt;/strong> When you give analysts and report builders direct read access to raw Salesforce data, the temptation to build dashboards straight from it is enormous. And for a quick proof of concept or a one-off analysis, fine. But the moment those PoC dashboards become &amp;ldquo;the dashboard the VP checks every Monday,&amp;rdquo; you&amp;rsquo;ve got a problem.&lt;/p>
&lt;p>Raw source data doesn&amp;rsquo;t tell you which &lt;code>CustomerNo&lt;/code> field to use — the one from the legacy migration, the one from the 2023 CRM consolidation, or the one from the current integration. It doesn&amp;rsquo;t encode the business rule that says &amp;ldquo;Closed Won&amp;rdquo; opportunities in APAC exclude training deals under $5,000 because those are tracked separately. It doesn&amp;rsquo;t flag that the &lt;code>Last_Activity_Date&lt;/code> field is unreliable for accounts owned by the partner team because they log activities in a different system.&lt;/p>
&lt;/br>
&lt;p>That knowledge lives in your data warehouse layer. It lives in the dbt models that your team built over months of conversations with stakeholders, through change management cycles, through debugging sessions where someone finally said &amp;ldquo;oh, we stopped using that field in Q3 because&amp;hellip;&amp;rdquo; That&amp;rsquo;s not transformation logic — it&amp;rsquo;s &lt;em>institutional memory encoded as code&lt;/em>. It connects raw data across domains and gives the business a full landscape picture instead of a narrow, single-source view. It catches the behavioural anomalies unique to your organisation — the sales rep who bulk-updates 500 records every Friday afternoon, the integration that occasionally double-fires during DST transitions, the custom object that gets repurposed every time the business restructures.&lt;/p>
&lt;p>So by all means, evaluate zero-copy for the use cases where it shines. Real-time signals on Account pages for sales reps? Great fit. Federated access for a specific analytics team that knows the source data intimately? Reasonable. But don&amp;rsquo;t let it replace the warehouse layer. That layer is where the hard-won understanding of your data lives, and it&amp;rsquo;s harder to recreate than any pipeline.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-honest-trade-offs">The Honest Trade-Offs&lt;/h3>
&lt;/br>
&lt;p>I&amp;rsquo;ve been positive about OpenFlow because I think it solves a real problem well. But I&amp;rsquo;d be violating my own writing principles if I didn&amp;rsquo;t lay out the limitations clearly.&lt;/p>
&lt;p>&lt;strong>The Salesforce connector is still in preview.&lt;/strong> It works, but it&amp;rsquo;s not covered by production support SLAs. If you&amp;rsquo;re running a mission-critical pipeline that your CFO stares at every Monday morning, that matters.&lt;/p>
&lt;p>&lt;strong>Custom Salesforce domains aren&amp;rsquo;t supported.&lt;/strong> If your org uses a vanity URL (like &lt;code>mycompany.my.salesforce.com&lt;/code> with a custom domain), check the documentation carefully before committing.&lt;/p>
&lt;p>&lt;strong>Hard deletes aren&amp;rsquo;t tracked.&lt;/strong> The &lt;code>ISDELETED&lt;/code> column catches soft deletes, but records that are permanently purged from Salesforce will persist in your Snowflake tables indefinitely. You&amp;rsquo;ll need a reconciliation process if that matters for your use case.&lt;/p>
&lt;p>&lt;strong>Certain field types are silently dropped.&lt;/strong> &lt;code>location&lt;/code>, &lt;code>address&lt;/code> (compound), and &lt;code>base64&lt;/code> fields are not synced. The connector doesn&amp;rsquo;t error on these — it just ignores them. This is the kind of thing that bites you three months in when someone asks why geographic data isn&amp;rsquo;t in the warehouse.&lt;/p>
&lt;p>&lt;strong>Formula fields require full refresh.&lt;/strong> Because formula field values are computed server-side and don&amp;rsquo;t update &lt;code>SystemModstamp&lt;/code>, they can&amp;rsquo;t be synced incrementally. You need a separate connector instance running full refreshes to capture formula field changes. That&amp;rsquo;s not a bug — it&amp;rsquo;s how Salesforce works — but it&amp;rsquo;s a gotcha if you&amp;rsquo;re not expecting it.&lt;/p>
&lt;p>&lt;strong>No relationship traversal.&lt;/strong> You can&amp;rsquo;t configure the connector to follow lookups and pull related objects automatically. Each object is pulled independently. Joining them is your job in the transformation layer, which is where it belongs anyway if you&amp;rsquo;re running dbt.&lt;/p>
&lt;/br>
&lt;p>None of these are dealbreakers. But every one of them is the kind of thing that bites an engineer at 4pm on a Friday if they weren&amp;rsquo;t expecting it. Now you&amp;rsquo;re expecting it.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="getting-started">Getting Started&lt;/h3>
&lt;/br>
&lt;p>If you want to try this on your own Salesforce org, here&amp;rsquo;s my recommended order of operations:&lt;/p>
&lt;p>Start with a &lt;strong>non-production Salesforce sandbox&lt;/strong> and a Snowflake trial account if you don&amp;rsquo;t have a dev environment handy. Pick two or three standard objects — &lt;code>Account&lt;/code>, &lt;code>Contact&lt;/code>, &lt;code>Opportunity&lt;/code> — and one custom object if your org has them. Run through the three-phase setup I described above. Get data flowing, verify the table structures, and then add a custom field to one of the objects in Salesforce. Wait for the next sync. Watch the column appear automatically in Snowflake.&lt;/p>
&lt;p>That&amp;rsquo;s the moment it clicks. That&amp;rsquo;s the moment you stop thinking about managed connectors as a luxury and start thinking about custom API code as technical debt you&amp;rsquo;re choosing to carry.&lt;/p>
&lt;p>Once you&amp;rsquo;ve validated the core flow, layer on your dbt models. OpenFlow lands raw data — it doesn&amp;rsquo;t transform it. Your staging models handle naming conventions, type casting, soft delete filtering, and the join logic that stitches related objects together. This is the same pattern you&amp;rsquo;d use with Fivetran or any other EL tool: land it raw, transform it in the warehouse.&lt;/p>
&lt;/br>
&lt;p>For teams already running Fivetran for Salesforce, I&amp;rsquo;m not suggesting you rip it out tomorrow. If it&amp;rsquo;s working and the cost is acceptable, that&amp;rsquo;s a solved problem. But the next time you&amp;rsquo;re adding a new high-volume source that&amp;rsquo;s in OpenFlow&amp;rsquo;s catalogue — especially if you&amp;rsquo;re already paying for Snowflake — it&amp;rsquo;s worth running the numbers. You might find that native integration with compute-based pricing is the better deal.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-real-why">The Real Why&lt;/h3>
&lt;/br>
&lt;p>I started this article with Marcus&amp;rsquo;s story because it illustrates something I feel strongly about: &lt;strong>your engineers&amp;rsquo; time is the most expensive resource in your data organisation, and spending it on solved problems is a leadership failure&lt;/strong>.&lt;/p>
&lt;p>Building a custom Salesforce integration isn&amp;rsquo;t impressive. It was impressive in 2016, when the tooling didn&amp;rsquo;t exist. Today, it&amp;rsquo;s a choice to take on maintenance burden that a managed service handles automatically. It&amp;rsquo;s choosing to babysit schema changes instead of building the models and analyses that actually move the business forward.&lt;/p>
&lt;p>Marcus is a senior engineer now. He builds dimensional models and designs data products that directly influence how the sales team allocates resources. He doesn&amp;rsquo;t write API integration code anymore. Not because he can&amp;rsquo;t — because his time is worth more than that.&lt;/p>
&lt;/br>
&lt;p>But here&amp;rsquo;s what I keep coming back to. The tools keep getting better — OpenFlow, Fivetran, zero-copy, and yes, AI-assisted pipeline generation. Every year there&amp;rsquo;s a new thing that promises to automate another piece of the data engineering workflow. And many of them genuinely do.&lt;/p>
&lt;p>What they don&amp;rsquo;t automate is the part that actually matters.&lt;/p>
&lt;p>No tool — not OpenFlow, not Fivetran, not an AI agent — is going to sit in a room with your head of sales and ask &amp;ldquo;what keeps you up at night?&amp;rdquo; No connector is going to notice that the way your company defines &amp;ldquo;active customer&amp;rdquo; has quietly drifted from what the CRM tracks. No pipeline is going to push back and say &amp;ldquo;we could build that dashboard, but the metric you&amp;rsquo;re asking for doesn&amp;rsquo;t answer the question you actually have.&amp;rdquo;&lt;/p>
&lt;/br>
&lt;p>That&amp;rsquo;s the connective tissue between raw data and business value. It&amp;rsquo;s a person who understands the domain, who&amp;rsquo;s been present through the change management cycles, who&amp;rsquo;s built relationships with the stakeholders and knows which &lt;code>CustomerNo&lt;/code> field to use because they were in the room when the decision was made three years ago. It&amp;rsquo;s someone who asks clarifying questions instead of presuming they already know the answer.&lt;/p>
&lt;p>AI will keep getting better at the mechanical parts of data engineering. Schema detection, pipeline code generation, anomaly detection — those are well-defined problems that automation is suited for. But the assumption that any tool can skip the human step — the clarification, the context, the judgment about what&amp;rsquo;s actually worth building — that&amp;rsquo;s where I&amp;rsquo;ve seen projects go sideways. Not because the technology failed, but because nobody asked the right questions before building.&lt;/p>
&lt;/br>
&lt;p>OpenFlow isn&amp;rsquo;t perfect. The connector catalogue is young, the Salesforce connector is in preview, and Fivetran still wins on breadth and battle-hardened maturity. But OpenFlow represents something important: the data platform taking responsibility for data movement, not just data storage and compute. It&amp;rsquo;s Snowflake saying &amp;ldquo;we&amp;rsquo;ll handle getting the data in — you focus on making it useful.&amp;rdquo;&lt;/p>
&lt;p>That last part — &lt;em>making it useful&lt;/em> — is still your job. And it&amp;rsquo;s the part that no tool can do for you. The best thing a managed connector gives you isn&amp;rsquo;t fewer lines of code. It&amp;rsquo;s time back. Time to spend on the work that actually requires a human: understanding the business, building the right models, and making sure the data tells an accurate story.&lt;/p>
&lt;p>The next time someone on your team says &amp;ldquo;I&amp;rsquo;ll just build a quick Salesforce integration&amp;rdquo; — send them this article. And the next time someone says &amp;ldquo;AI can just handle the data pipeline&amp;rdquo; — ask them who&amp;rsquo;s going to sit with the stakeholders and figure out what the pipeline should actually deliver.&lt;/p>
&lt;p>That&amp;rsquo;s still you. Make sure you have the time for it.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Cloud Architecture</category><category>Data Engineering</category><category>Snowflake</category><category>OpenFlow</category><category>Salesforce</category><category>API Integration</category><category>Schema Evolution</category><category>Fivetran</category><category>Data Pipelines</category></item><item><title>Your Data Model Isn't Broken, Part II: The Refactoring Playbook</title><link>https://ghostinthedata.info/posts/2026/2026-03-28-your-data-model-isnt-broken-part-2/</link><pubDate>Sat, 28 Mar 2026 09:00:00 +1100</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-03-28-your-data-model-isnt-broken-part-2/</guid><author>Chris Hillman</author><description>Strangler Figs, Write-Audit-Publish, and the art of replacing a data warehouse one piece at a time without anyone noticing. The practical sequel to why you shouldn't rebuild from scratch.</description><content:encoded>&lt;p>In [Part I], I made the case that your legacy data model isn&amp;rsquo;t the disaster it looks like. That the strange WHERE clauses, the bridge tables nobody can explain, and the slowly-changing-dimension-within-a-slowly-changing-dimension aren&amp;rsquo;t bugs — they&amp;rsquo;re business rules earned through years of production reality. I argued that big-bang rebuilds fail at alarming rates, that the complexity you&amp;rsquo;re fighting is mostly essential rather than accidental, and that the impulse to &amp;ldquo;start from scratch&amp;rdquo; is driven more by cognitive bias than by engineering judgment.&lt;/p>
&lt;p>A few people messaged me afterward and said, roughly: &amp;ldquo;Okay, I&amp;rsquo;m convinced. But what am I supposed to &lt;em>do&lt;/em> about it?&amp;rdquo;&lt;/p>
&lt;p>Fair question. Let&amp;rsquo;s talk about the playbook.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-vine-that-eats-the-tree">The vine that eats the tree&lt;/h3>
&lt;br>
&lt;p>In 2004, Martin Fowler was visiting a rainforest in Queensland when he noticed something remarkable about the strangler fig trees. These vines start as seeds deposited in the upper branches of a host tree by birds. They grow downward, wrapping around the trunk, eventually establishing their own root system in the soil. Over years — sometimes decades — the fig gradually replaces the host tree. The original tree decays inside the fig&amp;rsquo;s lattice until there&amp;rsquo;s nothing left of it. But at no point during this process does the canopy collapse. The forest continues to function the entire time.&lt;/p>
&lt;p>Fowler looked at those trees and thought: that&amp;rsquo;s how you should replace software systems.&lt;/p>
&lt;p>The Strangler Fig pattern is, in my experience, the single most useful idea in all of software architecture — and it maps onto data warehouse migrations with almost uncomfortable precision. The concept is deceptively simple: instead of replacing the old system all at once, you build the new system around it. You intercept one piece of functionality at a time, route it through the new implementation, verify it works, and move on. The old system slowly shrinks. The new system slowly grows. And at no point does the business lose access to anything.&lt;/p>
&lt;p>Fowler updated his essay on this pattern as recently as August 2024, and his framing is worth internalising. He described the alternative — the big-bang rewrite — and noted that he&amp;rsquo;d watched it fail most of the time. The strangler fig approach, by contrast, gives value steadily and allows you to monitor progress more carefully through frequent releases.&lt;/p>
&lt;p>Now, here&amp;rsquo;s where it gets interesting for data teams specifically. Application systems have clean interfaces — APIs, endpoints, well-defined contracts. You can intercept a request, route it to the new system, and nobody downstream knows the difference. Data systems are messier. Elliott Cordo, who documented migrating three legacy data warehouses using this pattern, identified the core challenge: analytics systems have porous boundaries and poorly defined interfaces with consumers. You don&amp;rsquo;t always know who&amp;rsquo;s querying what, or how.&lt;/p>
&lt;p>Cordo&amp;rsquo;s solution — and this is the variant I&amp;rsquo;ve seen work best — is what he calls the Legacy Façade. You build a new interface that mimics the legacy system&amp;rsquo;s contract. Same table names, same column names, same output format. Behind the façade, you&amp;rsquo;ve plumbed in the new logic. Consumers don&amp;rsquo;t change anything. They don&amp;rsquo;t even need to know a migration is happening. And this, critically, eliminates the organisational bottleneck that kills most migrations: getting twenty different teams to simultaneously update their queries, dashboards, and downstream pipelines to point at the new system.&lt;/p>
&lt;p>Cordo&amp;rsquo;s experience is particularly telling because one of the three warehouses he migrated existed &lt;em>precisely because&lt;/em> a previous attempt to replace one of the others had failed — mainly due to change management. The strangler fig approach bypasses that failure mode entirely by asking consumers to do little or nothing to adopt the new system.&lt;/p>
&lt;p>I think this matters so much: the majority of data warehouse migrations don&amp;rsquo;t fail for technical reasons. They fail because you can&amp;rsquo;t convince the entire organisation to cut over on the same Tuesday. The Legacy Façade removes that dependency. You migrate at your pace, not at the pace of every team that consumes your data.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="stop-publishing-before-youve-checked">Stop publishing before you&amp;rsquo;ve checked&lt;/h3>
&lt;br>
&lt;p>I&amp;rsquo;m continually surprised by how few people have heard of this pattern. When Robin Moffatt polled data engineers on Reddit about Write-Audit-Publish, most hadn&amp;rsquo;t even encountered the term — let alone used it. That&amp;rsquo;s a problem, because WAP solves the single scariest moment in any refactoring effort: the moment you push changed logic to production and hope you didn&amp;rsquo;t break anything.&lt;/p>
&lt;p>Netflix introduced WAP at the DataWorks Summit in 2017, and the concept is elegant in its simplicity. Instead of the typical pipeline flow — transform data, write it to the production table, then maybe run some checks afterward — you split the process into three distinct stages:&lt;/p>
&lt;p>&lt;strong>Write&lt;/strong> to an isolated staging area. Not production. Not anywhere consumers can see it.&lt;/p>
&lt;p>&lt;strong>Audit&lt;/strong> the output against your quality checks. Row counts. Schema validation. Business rule assertions. Whatever gives you confidence.&lt;/p>
&lt;p>&lt;strong>Publish&lt;/strong> only if the audit passes. If it doesn&amp;rsquo;t, production never sees the bad data. Nobody&amp;rsquo;s dashboard breaks. Nobody&amp;rsquo;s quarterly report includes garbage. The corrupted data dies quietly in staging where it belongs.&lt;/p>
&lt;p>The reason this matters for refactoring specifically is that refactoring &lt;em>requires&lt;/em> you to change transformation logic while producing identical output. You&amp;rsquo;re restructuring the internals — breaking a 400-line SQL query into modular CTEs, moving logic from stored procedures into dbt models, replacing hardcoded values with reference table lookups. Each of these changes is supposed to be behaviour-preserving. WAP gives you the safety net to verify that it actually was.&lt;/p>
&lt;p>Julien Hurault described the common anti-pattern — what he calls Write-(Publish)-Audit — where if an error is detected during testing, it&amp;rsquo;s already too late because the corrupted data has already been released to downstream systems. I&amp;rsquo;ve seen this play out more times than I want to admit. Someone refactors a model, the tests pass in the development environment, the change gets deployed, and then a stakeholder notices that revenue is off by 12% in a dashboard three hours later. By that point, the bad data has propagated through six downstream models and someone&amp;rsquo;s already screenshot the dashboard for a presentation.&lt;/p>
&lt;p>WAP prevents this entirely. And with Apache Iceberg now supporting it natively through branch-based isolation, the tooling has finally caught up with the idea. You write to an audit branch, run your validations, and only fast-forward merge to main when everything checks out. It&amp;rsquo;s the data engineering equivalent of a pull request with automated tests — and honestly, it&amp;rsquo;s bizarre that we&amp;rsquo;ve accepted as normal the practice of deploying data transformations straight to production without a verification step in between.&lt;/p>
&lt;p>If you&amp;rsquo;re on Snowflake, the implementation is straightforward using CLONE and SWAP operations. Clone the production table, run your refactored pipeline against the clone, validate the output, and swap the clone into the production role. If validation fails, the clone gets dropped and production is untouched. I&amp;rsquo;ve written about this pattern &lt;a href="https://ghostinthedata.info/blog/write-audit-publish-with-iceberg">in more detail here&lt;/a> — but the key point for this article is simpler: WAP turns refactoring from a high-wire act into a methodical, reversible process.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-archaeology-of-sql">The archaeology of SQL&lt;/h3>
&lt;br>
&lt;p>How does refactoring a data model actually works in practice, because it&amp;rsquo;s less dramatic than people expect. That&amp;rsquo;s sort of the point.&lt;/p>
&lt;p>dbt Labs published a refactoring course that formalises the approach, and having used variants of this workflow on multiple teams, I think they&amp;rsquo;ve got it right. The core philosophy is one sentence: while refactoring, you&amp;rsquo;ll be moving around a lot of logic, but ideally you won&amp;rsquo;t be changing the logic.&lt;/p>
&lt;p>Read that again. You are not improving business rules during a refactor. You are not fixing data quality issues during a refactor. You are not adding new features during a refactor. You are rearranging the existing logic into a better structure while proving, at every step, that the output hasn&amp;rsquo;t changed. The improvements come later, once the structure is clean enough to support them safely.&lt;/p>
&lt;p>The workflow goes something like this. First, you take the legacy SQL — the 800-line stored procedure, the five nested subqueries, the transformation that nobody wants to touch because the last person who edited it left the company — and you port it into a dbt model &lt;em>unchanged&lt;/em>. You don&amp;rsquo;t clean it up. You don&amp;rsquo;t refactor it. You bring it over exactly as it is, warts and all, and you verify that it produces identical output.&lt;/p>
&lt;p>This is the step most people want to skip, and it&amp;rsquo;s the most important one. That ugly SQL is your baseline. It&amp;rsquo;s your source of truth. Until you&amp;rsquo;ve proved your refactored version matches it row-for-row, column-for-column, you don&amp;rsquo;t have a refactor — you have a rewrite wearing a refactor&amp;rsquo;s clothing.&lt;/p>
&lt;p>Once you&amp;rsquo;ve got your baseline, you start decomposing. Extract the source references into dbt sources. Break the monolithic query into CTEs, each one doing a single logical operation. Move shared logic into staging models. Separate business rules from data cleaning from aggregation. At each step — and this is non-negotiable — you use something like dbt&amp;rsquo;s &lt;code>audit_helper&lt;/code> package to compare the output of your refactored model against the original. If the outputs match, you move on. If they don&amp;rsquo;t, you figure out what you changed and fix it before going further.&lt;/p>
&lt;p>Tristan Handy talks about his formative experience at Casper in 2016, where his team refactored all of their existing pipelines and brought them over to dbt in a single week, delivering work at least ten times faster than they would have been able to do otherwise. That speed came not from heroics but from methodology: migrate unchanged, then decompose with verification at every step.&lt;/p>
&lt;p>I&amp;rsquo;ve watched a team at refactor over 200 dbt models with complex logic in three days using &lt;code>audit_helper&lt;/code>. Three days. Not three months. Not the eighteen-month rebuild timeline that some stakeholder approved in a steering committee. Three days, because the tooling verified every change automatically and the team could move fast &lt;em>with confidence&lt;/em>.&lt;/p>
&lt;p>That&amp;rsquo;s the real payoff of refactoring over rebuilding. Not that it&amp;rsquo;s safer — though it is. Not that it preserves business logic — though it does. The payoff is that it&amp;rsquo;s &lt;em>faster&lt;/em>. Faster to start delivering value. Faster to reach a clean architecture. Faster to build the confidence you need to actually make meaningful improvements.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="expand-then-contract">Expand, then contract&lt;/h3>
&lt;br>
&lt;p>There&amp;rsquo;s a pattern for handling breaking schema changes that I don&amp;rsquo;t see data teams use nearly enough, and it deserves more attention. Fowler calls it Parallel Change. Others call it Expand and Contract. The name doesn&amp;rsquo;t matter.&lt;/p>
&lt;p>Imagine you need to rename a column in a fact table that thirty dashboards reference. In a rebuild mindset, you&amp;rsquo;d design the new schema, build it, and then coordinate a cutover where every consumer updates their queries simultaneously. Good luck with that. In my experience, &amp;ldquo;simultaneously&amp;rdquo; means &amp;ldquo;over the course of several painful weeks, with a growing list of things that are broken in the meantime.&amp;rdquo;&lt;/p>
&lt;p>Expand and Contract takes a different approach, in three phases:&lt;/p>
&lt;p>In the &lt;strong>Expand&lt;/strong> phase, you add the new column alongside the old one. Both exist. Both contain the same data. Nothing breaks, because nothing has been removed.&lt;/p>
&lt;p>In the &lt;strong>Migrate&lt;/strong> phase, you update consumers one at a time. Dashboard A switches to the new column. Pipeline B switches. Report C switches. Each migration is independent, low-risk, and reversible. There&amp;rsquo;s no coordination required between teams.&lt;/p>
&lt;p>In the &lt;strong>Contract&lt;/strong> phase — and only after every consumer has migrated — you drop the old column. By this point, nobody is using it, so removing it is a non-event.&lt;/p>
&lt;p>Does this take more deployments? Yes. it might require up to five deployments before the change is fully in effect. But each deployment is safe. Each deployment is reversible. And at no point does anyone lose access to their data because you renamed a column in a way they weren&amp;rsquo;t expecting.&lt;/p>
&lt;p>For data models specifically, Expand and Contract shines when you&amp;rsquo;re restructuring dimensions, splitting fact tables, or changing grain. Add the new structure alongside the old one. Migrate consumers gradually. Remove the old structure once it&amp;rsquo;s no longer referenced. It&amp;rsquo;s boring. It&amp;rsquo;s methodical. It works.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-ship-that-replaced-itself">The ship that replaced itself&lt;/h3>
&lt;br>
&lt;p>If you want a single case study that demonstrates what successful incremental replacement looks like at scale, look at what Slack did with their desktop client in 2019.&lt;/p>
&lt;p>Slack&amp;rsquo;s engineering team faced a classic dilemma. Their desktop application — built on jQuery and a custom framework — had accumulated years of complexity. It was slow. It used too much memory. The architecture made certain improvements difficult or impossible. Every instinct said: rewrite it in React.&lt;/p>
&lt;p>Their engineering blog explicitly acknowledged the risk: running code knows things. Hard-won knowledge gained through billions of hours of cumulative usage and tens of thousands of bug fixes.&lt;/p>
&lt;p>So instead of a big-bang rewrite, they did something much harder and much smarter. Over two years, they replaced every component of the desktop client while shipping continuously. They called it modernising &amp;ldquo;bit by bit&amp;rdquo; — enforcing strict interfaces between existing and modern code, shipping every change to production as it was completed, and never asking users to endure a disruptive cutover.&lt;/p>
&lt;p>The results: 50% memory reduction, 33% load time improvement. Users never experienced a broken release. Had they waited until the entire application was rewritten before releasing it, their users would have had a worse experience for years while the team worked in a vacuum.&lt;/p>
&lt;p>Now, Slack is an application, not a data warehouse. But the principles translate directly. Strict interfaces between old and new code? That&amp;rsquo;s the Legacy Façade. Shipping continuously? That&amp;rsquo;s WAP combined with Expand and Contract. Never asking users to endure a disruption? That&amp;rsquo;s the entire point.&lt;/p>
&lt;p>The ancient Greeks had a thought experiment about this: if you replace every plank of a ship over time, is it still the same ship? Philosophers argue about it. Engineers don&amp;rsquo;t care. The ship still floats. The passengers still get where they&amp;rsquo;re going.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-numbers-for-the-sceptics">The numbers, for the sceptics&lt;/h3>
&lt;br>
&lt;p>I held back on the quantitative evidence in Part I because I wanted to make the argument from experience and conviction. But some people want receipts, so here they are.&lt;/p>
&lt;p>McKinsey and Oxford University studied 5,400 IT projects with initial budgets exceeding $15 million. On average, large projects ran 45% over budget and 7% over time, while delivering 56% less value than predicted. Seventeen percent were &amp;ldquo;black swans&amp;rdquo; — cost overruns exceeding 200% that threatened the company&amp;rsquo;s viability. One retailer abandoned a $1.4 billion IT modernisation, failed on a $600 million follow-up, and filed for bankruptcy.&lt;/p>
&lt;p>The Standish Group has been tracking IT project outcomes since 1994. In their original CHAOS report, only 16.2% of projects succeeded on time, on budget, with specified features. Average cost overrun: 189%. By 2020, the overall success rate had improved to 31% — but here&amp;rsquo;s the statistic that matters most: small projects succeeded approximately 90% of the time. Large projects succeeded less than 10% of the time.&lt;/p>
&lt;p>Read that again. Small projects: 90%. Large projects: under 10%.&lt;/p>
&lt;p>That&amp;rsquo;s not a technology problem. That&amp;rsquo;s a scope problem. And the single most reliable way to reduce scope is to stop rebuilding entire systems and start refactoring them one piece at a time.&lt;/p>
&lt;p>For data systems specifically, Gartner reports that more than 50% of data warehouse projects fail to reach user acceptance. Other analyses put the number closer to 80%. CDInsights reports that 70% of data warehouse modernisation projects exceed budget or fail. Integrate.io found that approximately 80% of data migration projects exceed timelines or budgets, with large-scale projects showing 50% higher failure rates than incremental approaches.&lt;/p>
&lt;p>I don&amp;rsquo;t know how much clearer the data can be. Big-bang rewrites of data platforms fail most of the time. Incremental approaches succeed most of the time. The methodology I&amp;rsquo;ve described in this article — strangler fig, WAP, dbt refactoring, expand and contract — isn&amp;rsquo;t just safer or more philosophically sound. It&amp;rsquo;s what the evidence says actually works.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="refactoring-as-a-practice-not-a-project">Refactoring as a practice, not a project&lt;/h3>
&lt;br>
&lt;p>Refactoring should not be a project. It should not have a start date and an end date and a steering committee and a Gantt chart. Refactoring should be a &lt;em>practice&lt;/em> — something your team does continuously, as a natural part of how you work with your data models.&lt;/p>
&lt;p>Every time you touch a model to add a feature or fix a bug, you leave it a little cleaner than you found it. You extract a hardcoded value into a reference table. You add a test that didn&amp;rsquo;t exist before. You rename a column from &lt;code>col_7&lt;/code> to something a human can understand. You break a 200-line CTE into two smaller, focused ones.&lt;/p>
&lt;p>None of these changes are dramatic. None of them require a business case or a migration plan. They&amp;rsquo;re small, safe, incremental improvements that compound over time. A year from now, the model is cleaner, better-tested, and easier to understand — not because you stopped everything to rebuild it, but because you improved it a little bit every time you were in the neighbourhood.&lt;/p>
&lt;p>This is what Fowler meant when he defined refactoring as changing internal structure without modifying observable behaviour. It&amp;rsquo;s not a phase. It&amp;rsquo;s a discipline. And it&amp;rsquo;s the discipline that separates data teams that thrive from data teams that periodically torch everything and start over every three years, wondering why they never seem to make lasting progress.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-closing-argument">The closing argument&lt;/h3>
&lt;br>
&lt;p>In Part I, I said your data model isn&amp;rsquo;t broken — it&amp;rsquo;s battle-scarred, and those scars are knowledge. In Part II, I&amp;rsquo;ve tried to show you what the alternative to demolition looks like: patient, methodical, verified improvement. Strangler figs instead of bulldozers. Write-Audit-Publish instead of deploy-and-pray. Expand and contract instead of coordinate-and-hope.&lt;/p>
&lt;p>It&amp;rsquo;s not glamorous work. Nobody&amp;rsquo;s going to invite you to give a conference talk about the time you renamed some columns and extracted a few staging models. You won&amp;rsquo;t get a promotion for migrating a legacy fact table so smoothly that nobody noticed it happened.&lt;/p>
&lt;p>But that&amp;rsquo;s exactly the point. The best infrastructure work is invisible. The best migration is the one where the business never had to care. The best data model is the one that quietly absorbed twenty years of business complexity and still answers questions accurately at 7am on a Monday when a Department Head needs numbers for a board meeting.&lt;/p>
&lt;p>I wrote Part I of this series last week. The technology has changed — Netscape to Snowflake, CGI scripts to dbt — but the lesson hasn&amp;rsquo;t. Old code that works is more valuable than new code that doesn&amp;rsquo;t exist yet. Embedded knowledge is harder to create than clean architecture. And the patient, unglamorous work of incremental improvement will always outperform the seductive fantasy of starting fresh.&lt;/p>
&lt;p>Your data model isn&amp;rsquo;t broken. Stop trying to replace it. Start making it better.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Data Modelling</category><category>Data Engineering</category><category>Refactoring</category><category>Data Warehousing</category><category>dbt</category><category>Snowflake</category><category>Apache Iceberg</category><category>Write-Audit-Publish</category><category>Strangler Fig</category><category>Data Quality</category></item><item><title>You Don't Need Permission to Fix Your Data</title><link>https://ghostinthedata.info/posts/2026/2026-03-21-data-quality-when-your-a-junior/</link><pubDate>Sat, 21 Mar 2026 09:00:00 +1100</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-03-21-data-quality-when-your-a-junior/</guid><author>Chris Hillman</author><description>Five battle-tested tactics junior data engineers can use to improve data quality without waiting for authority, backed by real case studies from Airbnb, Google, and Warner Bros. Discovery.</description><content:encoded>&lt;p>Let me tell you about a junior engineer called Sam.&lt;/p>
&lt;p>Sam had been on the team about four months when I noticed something in a pull request. Tucked between two routine model changes was a new &lt;code>schema.yml&lt;/code> entry — five &lt;code>accepted_values&lt;/code> tests on a column called &lt;code>customer_status&lt;/code> that had been silently accumulating fourteen different spellings of &amp;ldquo;active&amp;rdquo; for the better part of a year.&lt;/p>
&lt;p>Nobody asked Sam to do this. It wasn&amp;rsquo;t in a sprint. There was no Jira ticket. Sam had just been working in that part of the warehouse, noticed the mess, and decided to clean it up on the way through.&lt;/p>
&lt;p>I left a comment on the PR: &amp;ldquo;Nice catch.&amp;rdquo; Sam&amp;rsquo;s reply was almost apologetic — &amp;ldquo;Hope it&amp;rsquo;s okay that I added these, I wasn&amp;rsquo;t sure if I was supposed to.&amp;rdquo;&lt;/p>
&lt;p>That response stuck with me. &lt;em>I wasn&amp;rsquo;t sure if I was supposed to.&lt;/em>&lt;/p>
&lt;p>Here was someone who&amp;rsquo;d spotted a real problem, built a real solution, tested it, and shipped it — and their instinct was to wonder whether they had permission. Not whether the fix was correct. Not whether the tests were well-written. Whether they were &lt;em>allowed&lt;/em>.&lt;/p>
&lt;p>I&amp;rsquo;ve thought about that moment a lot since, because it captures something I see in almost every data team I work with. The people closest to the problems — the ones running the queries, staring at the nulls, fielding the Microsoft Teams messages when dashboards look wrong — are often the last to believe they can do anything about it. Not because they lack skill. Because somewhere along the way, they absorbed the idea that fixing things is someone else&amp;rsquo;s job.&lt;/p>
&lt;p>That belief is expensive. And it&amp;rsquo;s wrong.&lt;/p>
&lt;p>I&amp;rsquo;m telling you this because data quality — the real, unglamorous, column-by-column work of making data trustworthy — is one of those things that sounds like a senior engineering problem until you see a four-month-old team member fix something a dozen experienced engineers walked past every day. It doesn&amp;rsquo;t take authority. It takes the willingness to act before you&amp;rsquo;re told to.&lt;/p>
&lt;p>This article is about how to be more like Sam. Five tactics you can start this week — no permission required — backed by real case studies from companies like Airbnb, Google, and Warner Bros. Discovery that figured this out the hard way.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="the-129-million-reason-your-quality-problem-isnt-too-small-to-fix">The $12.9 million reason your quality problem isn&amp;rsquo;t &amp;ldquo;too small&amp;rdquo; to fix&lt;/h3>
&lt;br>
&lt;p>Before we get into tactics, let&amp;rsquo;s talk about why your &amp;ldquo;small&amp;rdquo; fix matters more than you think.&lt;/p>
&lt;p>Gartner pegs the average annual cost of poor data quality at &lt;strong>$12.9 million per organisation&lt;/strong>. IBM and Harvard Business Review put the aggregate U.S. cost at $3.1 trillion annually. These numbers feel abstract until you realise what they actually represent: knowledge workers spending roughly half their time hunting for data, correcting errors, and searching for confirmatory sources instead of doing their actual jobs.&lt;/p>
&lt;p>Monte Carlo&amp;rsquo;s 2024 State of Data Quality survey found bad data now impacts &lt;strong>31% of company revenue&lt;/strong> on average — up from 26% just two years earlier. Their 2022 survey found data engineers spend &lt;strong>40% of their time&lt;/strong> evaluating or checking data quality. Soda&amp;rsquo;s research put that number even higher: 61% of data engineers spend half or more of their time handling data issues.&lt;/p>
&lt;p>And the human cost? A DataKitchen survey of 600 data engineers found &lt;strong>87% are frequently blamed&lt;/strong> when things go wrong with data. 78% wish their job came with a therapist.&lt;/p>
&lt;p>The implication for you: the problem is so large and so costly that any measurable improvement you create has real dollar value attached to it. That null check you added last Thursday? That&amp;rsquo;s not a minor fix — it&amp;rsquo;s a tiny piece of a multi-million-dollar problem that nobody else bothered to address.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="tactic-1-write-tests-for-your-own-models-without-asking-anyone">Tactic 1: write tests for your own models without asking anyone&lt;/h3>
&lt;br>
&lt;p>The single best case study for this tactic comes from Airbnb, and it didn&amp;rsquo;t start with a VP mandate or an OKR. It started with one engineer.&lt;/p>
&lt;p>Airbnb&amp;rsquo;s engineering blog documents how they went from &amp;ldquo;an engineering culture in which most new code shipped without a single test&amp;rdquo; to one where proposing a change without tests got called out immediately. The transformation wasn&amp;rsquo;t top-down. As the blog explains: &amp;ldquo;One approach is to rule by edict, but given our culture of engineer autonomy this would have been very poorly received. The other approach is to lead by example and build a movement.&amp;rdquo;&lt;/p>
&lt;p>That single champion&amp;rsquo;s playbook was remarkably tactical. They spoke in engineering meetings. They held office hours showing teammates how to write tests. They published a &lt;strong>weekly newsletter highlighting well-written test specs&lt;/strong> — giving public credit to engineers who included tests. They revamped the testing bootcamp for new hires, effectively making each new hire a champion. They improved tooling, reducing CI build times from over an hour to roughly six minutes.&lt;/p>
&lt;p>The result? Airbnb later applied the same grassroots philosophy to data quality, building a DQ Score on a 0–100 scale across their entire warehouse. Their anomaly detection system now prevents quality issues in new pipelines before they reach production. The Minerva metrics platform manages 12,000+ metrics and 4,000+ dimensions with 200+ data producers.&lt;/p>
&lt;p>One engineer. No mandate. A newsletter and some office hours.&lt;/p>
&lt;h4 id="what-this-looks-like-for-you">What this looks like for you&lt;/h4>
&lt;p>If you&amp;rsquo;re working with dbt, you already have the tools. dbt ships with four built-in generic tests — &lt;code>not_null&lt;/code>, &lt;code>unique&lt;/code>, &lt;code>accepted_values&lt;/code>, &lt;code>relationships&lt;/code> — that you define in YAML alongside your model definitions. No framework to build. No proposal to write. Just add a few lines to your &lt;code>schema.yml&lt;/code>:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">models&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">dim_customers&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">columns&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">customer_id&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">tests&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#ae81ff">not_null&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#ae81ff">unique&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">customer_status&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">tests&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">accepted_values&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">values&lt;/span>: [&lt;span style="color:#e6db74">&amp;#34;active&amp;#34;&lt;/span>, &lt;span style="color:#e6db74">&amp;#34;inactive&amp;#34;&lt;/span>, &lt;span style="color:#e6db74">&amp;#34;churned&amp;#34;&lt;/span>, &lt;span style="color:#e6db74">&amp;#34;suspended&amp;#34;&lt;/span>]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">email&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">tests&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#ae81ff">not_null&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>That&amp;rsquo;s it. Next time the pipeline runs, those tests execute automatically. If someone pushes a change that breaks uniqueness on &lt;code>customer_id&lt;/code>, the pipeline catches it before it reaches production.&lt;/p>
&lt;p>The &lt;code>dbt-expectations&lt;/code> package goes further, porting Great Expectations&amp;rsquo; advanced tests into dbt. Unit tests — introduced in dbt v1.8 — let you validate SQL logic with static mock inputs before transformations run. Tests integrate directly into CI/CD, blocking PRs with failures from reaching production.&lt;/p>
&lt;p>Not on dbt? You can still write assertion queries. Here&amp;rsquo;s a completeness check you can run against any SQL database — no framework required:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Quick completeness check: what percentage of each critical
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- field is actually populated?
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(&lt;span style="color:#f92672">*&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> total_records,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#75715e">-- How many emails do we actually have?
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> ROUND(&lt;span style="color:#ae81ff">100&lt;/span>.&lt;span style="color:#ae81ff">0&lt;/span> &lt;span style="color:#f92672">*&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(email) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(&lt;span style="color:#f92672">*&lt;/span>), &lt;span style="color:#ae81ff">1&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> email_pct,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#75715e">-- How many phone numbers?
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> ROUND(&lt;span style="color:#ae81ff">100&lt;/span>.&lt;span style="color:#ae81ff">0&lt;/span> &lt;span style="color:#f92672">*&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(phone) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(&lt;span style="color:#f92672">*&lt;/span>), &lt;span style="color:#ae81ff">1&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> phone_pct,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#75715e">-- How many addresses?
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> ROUND(&lt;span style="color:#ae81ff">100&lt;/span>.&lt;span style="color:#ae81ff">0&lt;/span> &lt;span style="color:#f92672">*&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(address_line_1) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(&lt;span style="color:#f92672">*&lt;/span>), &lt;span style="color:#ae81ff">1&lt;/span>) &lt;span style="color:#66d9ef">AS&lt;/span> address_pct,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#75715e">-- Flag: are we below the 95% threshold on anything critical?
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span> &lt;span style="color:#66d9ef">CASE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHEN&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(email) &lt;span style="color:#f92672">&amp;lt;&lt;/span> &lt;span style="color:#66d9ef">COUNT&lt;/span>(&lt;span style="color:#f92672">*&lt;/span>) &lt;span style="color:#f92672">*&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>.&lt;span style="color:#ae81ff">95&lt;/span> &lt;span style="color:#66d9ef">THEN&lt;/span> &lt;span style="color:#e6db74">&amp;#39;BELOW THRESHOLD&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">ELSE&lt;/span> &lt;span style="color:#e6db74">&amp;#39;OK&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">END&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> email_status
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> main.customers;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Schedule that to run daily. Store the results in a logging table. You&amp;rsquo;ve just built a monitoring system — and you didn&amp;rsquo;t ask anyone&amp;rsquo;s permission.&lt;/p>
&lt;h4 id="the-numbers-back-this-up">The numbers back this up&lt;/h4>
&lt;p>Monte Carlo&amp;rsquo;s 2022 survey found that organisations conducting at least three types of data tests weekly experienced &lt;strong>46 incidents per month&lt;/strong> compared to &lt;strong>61 incidents per month&lt;/strong> for less rigorous testers — a 25% reduction. The Forrester Total Economic Impact study of data observability platforms (April 2025) found 358% ROI, 6,500+ data personnel hours reclaimed annually, and $1.5M+ in avoided lost revenue from reduced data downtime.&lt;/p>
&lt;p>Daniel Beach of Data Engineering Central offers a pragmatic path for getting started: begin with end-to-end tests first when inheriting a pipeline with no tests, then add data quality checks, then unit tests. You don&amp;rsquo;t need to boil the ocean. You need to start.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="tactic-2-document-one-thing-every-time-you-touch-it">Tactic 2: document one thing every time you touch it&lt;/h3>
&lt;br>
&lt;p>The numbers on missing documentation are genuinely staggering. According to DX, developers spend &lt;strong>3–10 hours per week&lt;/strong> searching for information that should be documented. For a 100-person engineering team, that equals 300–1,000 hours weekly — the equivalent of 8–25 full-time engineers doing nothing but hunting for answers. Stack Overflow&amp;rsquo;s developer survey found 62% of developers spend over 30 minutes daily on poorly documented issues. And 38% listed poor documentation among their top reasons for leaving a company.&lt;/p>
&lt;p>One practitioner tracked &amp;ldquo;Time to First Production Commit&amp;rdquo; at &lt;strong>28 days&lt;/strong> for new engineers — versus an industry benchmark of 3–5 days. Engineers on that team spent 30% of their time rediscovering what previous engineers had already learned. 40% of sprint tickets required inter-team clarification, averaging 12 clarifications per ticket.&lt;/p>
&lt;p>This is what documentation debt looks like. Not a gap in a wiki — a tax on every single person who touches the codebase.&lt;/p>
&lt;h4 id="documentation-fridays-dont-work-this-does">&amp;ldquo;Documentation Fridays&amp;rdquo; don&amp;rsquo;t work. This does.&lt;/h4>
&lt;p>Multiple sources confirm that big-bang documentation efforts fail. As one case study puts it: &amp;ldquo;&amp;lsquo;Okay team, let&amp;rsquo;s spend Friday afternoon documenting our processes!&amp;rsquo; Nobody wants to do that. And it doesn&amp;rsquo;t work.&amp;rdquo;&lt;/p>
&lt;p>The successful pattern is document-as-you-go. Every time you encounter an issue, you write it down. Shopify Engineering treats documentation as a product, explicitly encouraging new team members to fill documentation gaps from day one. Their philosophy: &amp;ldquo;The most effective way to encourage others to write is to actually write.&amp;rdquo;&lt;/p>
&lt;p>For data teams specifically, the approach is even simpler. If you&amp;rsquo;re using dbt, column descriptions live in &lt;code>schema.yml&lt;/code> alongside your model definitions. dbt&amp;rsquo;s column-level lineage means passthrough and renamed columns &lt;strong>automatically inherit descriptions from upstream models&lt;/strong>. Write a description once, and it flows downstream through every model that references that column.&lt;/p>
&lt;p>Here&amp;rsquo;s what this looks like in practice:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">models&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">dim_customers&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">description&lt;/span>: &amp;gt;&lt;span style="color:#e6db74">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> One row per customer. Grain: customer_id.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Source: CRM extract via Fivetran, refreshed daily at 03:00 UTC.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Owner: data-eng-team@company.com&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">columns&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">customer_id&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">description&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;Primary key. Sourced from CRM.contacts.id. Unique, never null.&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">customer_status&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">description&lt;/span>: &amp;gt;&lt;span style="color:#e6db74">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Current lifecycle status. Valid values: active, inactive, 
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> churned, suspended. Updated by CRM webhook on status change.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Known issue: legacy records from pre-2023 may contain 
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> &amp;#39;Active&amp;#39; (capitalised) — normalised in this model.&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">lifetime_value&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">description&lt;/span>: &amp;gt;&lt;span style="color:#e6db74">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Cumulative revenue attributed to this customer, in AUD.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Calculated as SUM(order_total) from fct_orders.
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> Excludes refunded orders. Updated daily.&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>That &amp;ldquo;Known issue&amp;rdquo; note on &lt;code>customer_status&lt;/code>? That&amp;rsquo;s worth its weight in gold to the next person who touches this model. It took thirty seconds to write and it&amp;rsquo;ll save hours of confusion.&lt;/p>
&lt;p>Microsoft&amp;rsquo;s engineering blogs describe the alternative bluntly: &amp;ldquo;When tribal knowledge holders leave, reverse engineering is the only way to understand what was done.&amp;rdquo; Enboarder&amp;rsquo;s 2025 research found &lt;strong>47% of organisations&lt;/strong> cite institutional knowledge loss as their top offboarding challenge. Atlan&amp;rsquo;s Modern Data Survey found data professionals waste &lt;strong>one full day per week&lt;/strong> simply trying to figure out what data to use.&lt;/p>
&lt;p>You don&amp;rsquo;t need to document the entire warehouse. You need to document the thing you touched today. Tomorrow, document the next thing. In three months, you&amp;rsquo;ll have built something genuinely valuable — and you&amp;rsquo;ll have done it one commit at a time.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="tactic-3-build-a-quality-dashboard-nobody-asked-for">Tactic 3: build a quality dashboard nobody asked for&lt;/h3>
&lt;br>
&lt;p>At Warner Bros. Discovery, the DICE team built data quality management frameworks and created a &amp;ldquo;Data Quality Forum&amp;rdquo; that became a bridge between teams. During the Olympics livestream, they used custom SQL checks to detect missing content metadata — fixing issues before they broke reporting. The forum evolved into what Pam Zirpoli called a &amp;ldquo;cultural flywheel.&amp;rdquo; Patterns surfaced in the forum were incorporated into onboarding, translated into new monitors, and used to refine their priority matrix. Teams shifted from reactive — reacting after dashboards broke — to proactive.&lt;/p>
&lt;p>At Tempus, a data engineer documented how adopting Elementary (an open-source dbt package) transformed their monitoring. Before: &amp;ldquo;test alerting relied on a dashboard in Data Studio built on top of a log sink, and the &amp;lsquo;alerts&amp;rsquo; were a daily email that you actually had to open.&amp;rdquo; After: Microsoft Teams alerts that &amp;ldquo;@-ed a user and continued to do so until they&amp;rsquo;re turned off.&amp;rdquo; The key outcome: tests no longer fail silently.&lt;/p>
&lt;h4 id="the-six-metrics-that-actually-matter">The six metrics that actually matter&lt;/h4>
&lt;p>Drawing from IBM, dbt Labs, Databricks, and Alation, the consensus metrics for data quality monitoring are:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Freshness&lt;/strong>: time since last data update. If your orders table hasn&amp;rsquo;t been updated since yesterday, something&amp;rsquo;s broken.&lt;/li>
&lt;li>&lt;strong>Volume&lt;/strong>: expected versus actual row counts. Did today&amp;rsquo;s load bring in 50,000 rows when you normally get 48,000–52,000? Fine. Did it bring 12? Problem.&lt;/li>
&lt;li>&lt;strong>Schema changes&lt;/strong>: added, deleted, or altered columns. These should never surprise you.&lt;/li>
&lt;li>&lt;strong>Completeness&lt;/strong>: null rates per column. Simple &lt;code>COUNT(*)&lt;/code> vs &lt;code>COUNT(column)&lt;/code> tells you a lot.&lt;/li>
&lt;li>&lt;strong>Distribution shifts&lt;/strong>: statistical changes in data values. If the average order value suddenly doubles, that&amp;rsquo;s worth investigating.&lt;/li>
&lt;li>&lt;strong>Test pass/fail rates&lt;/strong>: the percentage of your quality assertions that are passing. Track this over time.&lt;/li>
&lt;/ul>
&lt;p>Here&amp;rsquo;s a freshness check you can run right now — no tools, no frameworks, just SQL:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- How stale is each of your critical tables?
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">-- Run this daily and store results in a monitoring table
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e">&lt;/span>&lt;span style="color:#66d9ef">SELECT&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#e6db74">&amp;#39;fct_orders&amp;#39;&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> &lt;span style="color:#66d9ef">table_name&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">MAX&lt;/span>(updated_at) &lt;span style="color:#66d9ef">AS&lt;/span> last_record_timestamp,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> checked_at,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ROUND(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">EXTRACT&lt;/span>(EPOCH &lt;span style="color:#66d9ef">FROM&lt;/span> (&lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span> &lt;span style="color:#f92672">-&lt;/span> &lt;span style="color:#66d9ef">MAX&lt;/span>(updated_at))) &lt;span style="color:#f92672">/&lt;/span> &lt;span style="color:#ae81ff">3600&lt;/span>.&lt;span style="color:#ae81ff">0&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#ae81ff">1&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> ) &lt;span style="color:#66d9ef">AS&lt;/span> hours_since_update,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">CASE&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHEN&lt;/span> &lt;span style="color:#66d9ef">MAX&lt;/span>(updated_at) &lt;span style="color:#66d9ef">IS&lt;/span> &lt;span style="color:#66d9ef">NULL&lt;/span> &lt;span style="color:#66d9ef">THEN&lt;/span> &lt;span style="color:#e6db74">&amp;#39;CRITICAL - No data&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHEN&lt;/span> &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span> &lt;span style="color:#f92672">-&lt;/span> &lt;span style="color:#66d9ef">MAX&lt;/span>(updated_at) &lt;span style="color:#f92672">&amp;gt;&lt;/span> INTERVAL &lt;span style="color:#e6db74">&amp;#39;24 hours&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">THEN&lt;/span> &lt;span style="color:#e6db74">&amp;#39;STALE - Exceeds 24hr SLA&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">WHEN&lt;/span> &lt;span style="color:#66d9ef">CURRENT_TIMESTAMP&lt;/span> &lt;span style="color:#f92672">-&lt;/span> &lt;span style="color:#66d9ef">MAX&lt;/span>(updated_at) &lt;span style="color:#f92672">&amp;gt;&lt;/span> INTERVAL &lt;span style="color:#e6db74">&amp;#39;12 hours&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">THEN&lt;/span> &lt;span style="color:#e6db74">&amp;#39;WARNING - Approaching SLA&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">ELSE&lt;/span> &lt;span style="color:#e6db74">&amp;#39;FRESH&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">END&lt;/span> &lt;span style="color:#66d9ef">AS&lt;/span> freshness_status
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span> analytics.fct_orders;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Store that in a &lt;code>data_quality_log&lt;/code> table, connect it to Looker Studio or Metabase (both free), and you&amp;rsquo;ve got a monitoring dashboard. That&amp;rsquo;s it. No Kubernetes cluster. No proposal document.&lt;/p>
&lt;h4 id="visibility-changes-behaviour">Visibility changes behaviour&lt;/h4>
&lt;p>Anna Geller documented how sending data quality alerts to a shared Microsoft Teams channel transforms team behaviour. Her core insight: &amp;ldquo;As social creatures, we are more motivated to fix issues if other people can see the effort we put into it.&amp;rdquo; Data owners reply with root cause analysis, add checkmarks, or link tickets — creating a social proof loop that issues are no longer ignored.&lt;/p>
&lt;p>This isn&amp;rsquo;t just anecdotal. Research on the Hawthorne effect shows healthcare hand hygiene compliance increases by &lt;strong>up to 70%&lt;/strong> when workers know they&amp;rsquo;re observed, and employees increase productivity by an average of 13% when activity is monitored. Making data quality visible doesn&amp;rsquo;t just inform people — it changes their behaviour.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="tactic-4-create-a-data-bugs-channel">Tactic 4: create a #data-bugs channel&lt;/h3>
&lt;br>
&lt;p>MIT Sloan Management Review&amp;rsquo;s HelloFresh case study documents one of the best examples of how visibility catalyses organisational change. HelloFresh transitioned through three modes of data quality management. In Mode 1, issues were dealt with reactively by whoever happened to be affected. In Mode 2, a central engineering team handled cleanup but lacked understanding of how data was actually used — &amp;ldquo;data customers grew increasingly frustrated as they continued to spend large fractions of their workdays dealing with data that did not meet their needs.&amp;rdquo;&lt;/p>
&lt;p>The transition to Mode 3 — proactive, collaborative quality management — happened precisely because &amp;ldquo;quality issues became increasingly visible.&amp;rdquo; Not because someone wrote a strategy doc. Not because a VP mandated it. Because the problems became impossible to ignore.&lt;/p>
&lt;p>You can accelerate that transition. Create a Microsoft Teams channel. Call it &lt;code>#data-bugs&lt;/code> or &lt;code>#data-quality&lt;/code> or whatever feels natural. Start posting issues you find — with specifics. Not &amp;ldquo;the data looks wrong&amp;rdquo; but &amp;ldquo;fct_orders has 3,247 rows with null customer_id from the 2026-02-15 load, affecting the weekly revenue dashboard.&amp;rdquo;&lt;/p>
&lt;p>That&amp;rsquo;s it. You&amp;rsquo;ve just created the visibility mechanism that HelloFresh spent years arriving at.&lt;/p>
&lt;h4 id="the-broken-windows-effect-is-real--and-its-been-measured">The broken windows effect is real — and it&amp;rsquo;s been measured&lt;/h4>
&lt;p>The Pragmatic Programmer introduced the &amp;ldquo;broken windows&amp;rdquo; software analogy decades ago: &amp;ldquo;Don&amp;rsquo;t leave &amp;lsquo;broken windows&amp;rsquo; (bad designs, wrong decisions, or poor code) unrepaired. Fix each one as soon as it is discovered.&amp;rdquo;&lt;/p>
&lt;p>This isn&amp;rsquo;t just metaphor anymore. A 2024 arXiv paper by Diomidis Spinellis et al. analysed &lt;strong>2 million code commits&lt;/strong> across &lt;strong>122 projects&lt;/strong> comprising &lt;strong>5.5 million lines of code&lt;/strong>. The key finding: &amp;ldquo;History matters — developers behave differently depending on some aspects of the code quality they encounter.&amp;rdquo; Developers tailor the quality of their commits based on the quality of the file they&amp;rsquo;re committing to. Low quality begets lower quality. High quality attracts higher quality.&lt;/p>
&lt;p>Mat Ryer extends this to everyday engineering: &amp;ldquo;If you work on a project that has flaky tests, then you&amp;rsquo;re more likely to add more flaky tests. If there is a hacky design, you&amp;rsquo;re more likely to hack more in.&amp;rdquo;&lt;/p>
&lt;p>Fixing one visible data quality issue — or even just making it visible — can interrupt the decay spiral.&lt;/p>
&lt;h4 id="why-juniors-dont-speak-up-and-why-thats-exactly-the-problem">Why juniors don&amp;rsquo;t speak up (and why that&amp;rsquo;s exactly the problem)&lt;/h4>
&lt;p>Here&amp;rsquo;s the uncomfortable part. A Culture Shift UK survey found &lt;strong>junior colleagues were twice as likely (54%) as senior leaders (27%)&lt;/strong> to say speaking up about issues is &amp;ldquo;pointless.&amp;rdquo; 37% said speaking up &amp;ldquo;isn&amp;rsquo;t worth the personal risk.&amp;rdquo; A Blind survey found 58% of tech workers experience imposter syndrome.&lt;/p>
&lt;p>Google&amp;rsquo;s Project Aristotle studied 180+ teams over two years and found psychological safety was &amp;ldquo;by far the most important&amp;rdquo; of five dynamics for team effectiveness. Amy Edmondson&amp;rsquo;s foundational research contains a striking finding: while studying hospital teams, she expected high-performing teams to report fewer errors. Instead, &lt;strong>better teams reported MORE errors&lt;/strong>. Her insight: &amp;ldquo;Maybe the good teams don&amp;rsquo;t make more mistakes, maybe they report more.&amp;rdquo;&lt;/p>
&lt;p>This is the paradox. The teams that look like they have the most data quality issues are often the healthiest — because they&amp;rsquo;ve created an environment where problems surface instead of festering.&lt;/p>
&lt;p>Etsy pioneered blameless postmortems under CTO John Allspaw. Engineers send company-wide emails confessing mistakes. Etsy even gives out an annual &amp;ldquo;three-armed sweater&amp;rdquo; award to the employee who made the most surprising error. As a Hadley Wickham–curated Stanford reading list notes, while blameless postmortems are standard in DevOps, they &amp;ldquo;have not yet infiltrated standard data science practices.&amp;rdquo; That&amp;rsquo;s a gap you can help fill.&lt;/p>
&lt;p>Creating a &lt;code>#data-bugs&lt;/code> channel isn&amp;rsquo;t just about tracking issues. It&amp;rsquo;s about establishing the norm that &lt;strong>finding problems is good work&lt;/strong>, not a sign that something&amp;rsquo;s wrong with you.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="tactic-5-ship-an-example-dont-write-a-proposal">Tactic 5: ship an example, don&amp;rsquo;t write a proposal&lt;/h3>
&lt;br>
&lt;p>Marcus Blankenship&amp;rsquo;s article on learned helplessness in software engineering — 96,000+ reads, 900+ Reddit comments — contains a story about a coworker called Milind. Milind wasn&amp;rsquo;t a manager. He didn&amp;rsquo;t have a fancy title. But he was an informal leader through pure action. He asked questions about why particular approaches were taken. He admitted ignorance in group settings. He insisted the team understand root causes before moving on.&lt;/p>
&lt;p>Blankenship&amp;rsquo;s observation: &amp;ldquo;Milind changed our environment by his actions&amp;hellip; He showed me that I had much more power than I thought, if I would only stop expecting to be spoon-fed everything. His actions showed me that &lt;strong>true leadership isn&amp;rsquo;t granted, it&amp;rsquo;s grasped.&lt;/strong>&amp;rdquo;&lt;/p>
&lt;p>A junior engineer on Irina Stanescu&amp;rsquo;s newsletter shared something similar: &amp;ldquo;Even as a junior engineer, I remember influencing my chief architect and brought Terraform into my organisation 7 years ago by simply pointing out a common pain point and how much longer things took without managing our infrastructure in code.&amp;rdquo; They added: &amp;ldquo;Being able to influence without authority is a very underrated skill, and I believe this is what helped me get promoted to senior engineer even more than my technical skills.&amp;rdquo;&lt;/p>
&lt;p>The pattern is consistent. Don&amp;rsquo;t propose a data quality framework. Don&amp;rsquo;t write a 15-page RFC. Build the thing. Make it work. Show people.&lt;/p>
&lt;h4 id="why-working-examples-spread-faster-than-proposals">Why working examples spread faster than proposals&lt;/h4>
&lt;p>Everett Rogers&amp;rsquo; Diffusion of Innovations theory identifies five characteristics that determine how fast an idea spreads: relative advantage, compatibility, complexity (lower is better), &lt;strong>trialability&lt;/strong>, and &lt;strong>observability&lt;/strong>. A working example maximises both trialability and observability — the two factors most under your control as a junior engineer.&lt;/p>
&lt;p>A proposal says &amp;ldquo;we should do this.&amp;rdquo; A working example says &amp;ldquo;look, I already did this, and here&amp;rsquo;s what happened.&amp;rdquo; One requires people to imagine the benefit. The other puts the benefit directly in front of them.&lt;/p>
&lt;p>Addy Osmani, after 14 years at Google, captured the complementary lesson: &amp;ldquo;Early in my career, I believed great work would speak for itself. I was wrong. Code sits silently in a repository.&amp;rdquo; You need to build the example &lt;em>and&lt;/em> make its impact visible. Post in Microsoft Teams. Show the before and after. Quantify what changed.&lt;/p>
&lt;h4 id="the-junior-advantage-nobody-talks-about">The junior advantage nobody talks about&lt;/h4>
&lt;p>Ant Weiss&amp;rsquo;s widely-read article on organisational learned helplessness identifies common symptoms: &amp;ldquo;This is how we were told to work,&amp;rdquo; &amp;ldquo;These are the limitations of the system,&amp;rdquo; &amp;ldquo;We don&amp;rsquo;t have the resources.&amp;rdquo; His core observation: &amp;ldquo;No engineer starts their career willing to be unproductive, pessimistic and stressed out. But the companies they work at teach them to.&amp;rdquo;&lt;/p>
&lt;p>The cure, drawn from Seligman&amp;rsquo;s original research, is critical: &amp;ldquo;Threats, rewards, and observed demonstrations had no effect on the &amp;lsquo;helpless&amp;rsquo; dogs. The dogs had to be physically moved through the escape action at least twice before they would do it themselves.&amp;rdquo; You can&amp;rsquo;t just tell engineers things can be better. You have to walk them through the experience of making something better.&lt;/p>
&lt;p>Here&amp;rsquo;s the part that should give you confidence. JazzTeam reports a remarkable finding: &amp;ldquo;To prove that any problem can be solved, our team many times gave a task to Junior and even Intern Engineers who did not have psychological fears and tunnel thinking.&amp;rdquo; Juniors, unencumbered by years of &amp;ldquo;that&amp;rsquo;s impossible&amp;rdquo; conditioning, were able to tackle problems seniors had declared unsolvable.&lt;/p>
&lt;p>Your fresh eyes aren&amp;rsquo;t a weakness. They&amp;rsquo;re an asset. You haven&amp;rsquo;t yet learned that &amp;ldquo;it can&amp;rsquo;t be done.&amp;rdquo; So go do it.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="putting-it-all-together">Putting it all together&lt;/h3>
&lt;br>
&lt;p>None of these tactics require a committee meeting. None require architectural approval. None require you to be senior.&lt;/p>
&lt;p>Here&amp;rsquo;s what they do require: the willingness to act before you&amp;rsquo;re told to.&lt;/p>
&lt;p>Write a test this week. Add a column description tomorrow. Build a freshness check and schedule it. Create a Slack channel and post the first issue. Turn one of your quality checks into a reusable example your team can copy.&lt;/p>
&lt;p>DORA&amp;rsquo;s research — surveying 39,000+ professionals — finds that with focused effort and the right tooling, teams typically see measurable improvements within &lt;strong>1–3 months&lt;/strong>. A GitHub engineer who went from junior to mid-level in 2.5 years described the compound effect: &amp;ldquo;If I got stuck on some undocumented functionality, I made sure to update the docs and let the team know. Tackling a tricky bug requiring cross-team collaboration? I&amp;rsquo;d summarise everything we discovered so that it would be easier for others later. Post about it in Microsoft Teams and highlight its impact.&amp;rdquo;&lt;/p>
&lt;p>Kevin Workman, a former Google engineer, captured the mindset that makes all of this work: &amp;ldquo;Realising that I had agency and even a responsibility to advocate for my ideas is one of the most important lessons I learned on my journey to becoming a &amp;lsquo;senior&amp;rsquo; software engineer&amp;hellip; Some of the most interesting and most successful moments of my career have come from advocating for changes I thought nobody would agree to.&amp;rdquo;&lt;/p>
&lt;p>I think about Sam sometimes — that four-month-old team member whose instinct was to apologise for fixing something. The tests Sam added that day caught three data quality issues in the following month. Other engineers started adding their own tests alongside model changes. Six months later, the team had 200+ automated quality checks running on every pipeline, and the weekly &amp;ldquo;data looks wrong&amp;rdquo; Microsoft Teams messages had dropped to nearly zero.&lt;/p>
&lt;p>Sam never got a mandate. Never wrote a proposal. Never waited for Q3.&lt;/p>
&lt;p>The data quality problem in your organisation is a $12.9 million issue wearing a &amp;ldquo;we&amp;rsquo;ll get to it in Q3&amp;rdquo; disguise. And you — yes, you, the person who&amp;rsquo;s been here five months and noticed fourteen spellings of &amp;ldquo;active&amp;rdquo; — are exactly the right person to start fixing it.&lt;/p>
&lt;p>Not because someone gave you permission. Because the data doesn&amp;rsquo;t care about your title.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Data Quality</category><category>Career Development</category><category>Data Quality</category><category>dbt</category><category>SQL</category><category>Testing</category><category>Documentation</category><category>Junior Engineer</category><category>Career Growth</category><category>Psychological Safety</category></item><item><title>Your Friends Will Be There for You. Your Work Won't.</title><link>https://ghostinthedata.info/posts/2026/2026-03-18-friendship/</link><pubDate>Wed, 18 Mar 2026 09:00:00 +1100</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-03-18-friendship/</guid><author>Chris Hillman</author><description>The longest study of adult happiness found relationships matter more than career achievement. Yet most of us keep cancelling on our friends. Here's what that costs — and what being a good friend actually looks like.</description><content:encoded>&lt;p>There&amp;rsquo;s a table in a rented house in Queenscliff — nothing fancy, just whatever furniture came with the place — with a hand-drawn map spread across it, dice scattered at the edges, and snacks slowly migrating toward the centre of the board.&lt;/p>
&lt;p>Once a year, Steve, one of our friends, organises for us to make a trip down to Queenscliff to play tabletop RPGs for a weekend. More recently, we&amp;rsquo;ve started playing monthly or fortnightly at home too. Nothing glamorous about any of it. We&amp;rsquo;re not Glass Cannon Network or Critical Role — no cameras, no production values, no audience watching us roleplay with solemn gravity.&lt;/p>
&lt;p>We&amp;rsquo;re just a bunch of friends. Laughing at the missed sword swing. Losing our minds when someone makes a completely unreasonable jump and somehow lands the killing blow on a copper golem with a natural twenty. Rolling dice and arguing about the rules and eating too much while the world outside keeps spinning without us.&lt;/p>
&lt;p>I&amp;rsquo;ve been thinking about why it matters so much to me.&lt;/p>
&lt;p>It&amp;rsquo;s not the game, really — or at least, the game isn&amp;rsquo;t the point. It&amp;rsquo;s what the game creates. A ritual. An excuse to sit in the same room for a weekend and actually be present with people I care about. A context where you ask &amp;ldquo;how are you going?&amp;rdquo; and actually mean it, and stay long enough to hear the real answer. The dragon we&amp;rsquo;re fighting is almost beside the point. What matters is being there for each other.&lt;/p>
&lt;p>When you&amp;rsquo;re up at 2am resolving a critical data pipeline failure, your employer will not remember that next year. They might not remember it next month. The organisation absorbs it, thanks you if you&amp;rsquo;re lucky, and moves on to the next incident.&lt;/p>
&lt;p>But fighting dragons with your friends — showing up for a weekend in Queenscliff year after year — those become stories. Stories that become memories. Memories that become the fabric of who you are to each other, and who they are to you.&lt;/p>
&lt;p>Your friends will be there for you. Your work won&amp;rsquo;t.&lt;/p>
&lt;p>And yet.&lt;/p>
&lt;p>When was the last time someone you loved lost someone, and you climbed into bed next to them — not to fix it, not to say the right thing, just to be there in the dark with them? When something hard was happening in your own life, did you pick up the phone and call a friend? Or did the thought cross your mind — &amp;ldquo;they&amp;rsquo;re busy, they don&amp;rsquo;t have time to listen to me&amp;rdquo; — and you put the phone down and carried it alone?&lt;/p>
&lt;p>Do you have a friend you&amp;rsquo;d call when you got the promotion — someone who&amp;rsquo;d hear the news and mean it when they say &amp;ldquo;that&amp;rsquo;s great&amp;rdquo;? Not quietly threatened by it. Not performing enthusiasm. Just genuinely proud of you, because your win is their win. That kind of friend is rarer than it should be.&lt;/p>
&lt;p>Those moments. That&amp;rsquo;s what friendship actually is. Not the catch-up drinks you schedule six weeks out and cancel twice. The showing up in the hard moments. The willingness to burden someone with your struggle, and the willingness to carry theirs.&lt;/p>
&lt;p>We don&amp;rsquo;t build trust by offering help. We build trust by asking for it. The act of calling someone when things are falling apart — not suffering quietly, not performing fine — is the thing that makes the friendship real on both sides. It&amp;rsquo;s not a burden. For the right person, it&amp;rsquo;s an honour.&lt;/p>
&lt;p>This is something I&amp;rsquo;m still learning. I want to talk about what it actually means to be a good friend — because I think most of us, if we&amp;rsquo;re honest, have let that muscle go slack.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="are-you-actually-a-good-friend">Are You Actually a Good Friend?&lt;/h3>
&lt;br>
&lt;p>Most people, if you ask them directly, will say yes. Of course they&amp;rsquo;re a good friend. They&amp;rsquo;re there for the people they care about.&lt;/p>
&lt;p>But ask yourself: do you call your friends on their birthday — actually call, sing happy birthday — or do you post on their Facebook because you saw the notification and everyone else was doing it? When a friend is going through something hard, do you show up? Not message. Show up. Do you say &amp;ldquo;I love you&amp;rdquo; to the people in your life that you love?&lt;/p>
&lt;p>Most of us, if we&amp;rsquo;re being honest, have outsourced friendship to the low-effort channel. The emoji response. The &amp;ldquo;we should catch up soon!&amp;rdquo; comment thread that leads nowhere. We maintain the social graph without maintaining the relationships.&lt;/p>
&lt;p>But when you&amp;rsquo;re in a dark place, who do you call? And more confronting — who would call you?&lt;/p>
&lt;p>The answer to that second question is the real measure. Friendship isn&amp;rsquo;t what you feel for someone. It&amp;rsquo;s what you demonstrate over time. It&amp;rsquo;s the accumulated weight of showing up — for the birthday dinner that falls on a work night, for the hospital waiting room at 7am, for the phone call that starts &amp;ldquo;I just need to talk to someone&amp;rdquo; at an inconvenient hour.&lt;/p>
&lt;p>Those moments don&amp;rsquo;t happen automatically. They have to be built. And they have to be built before you need them, not when you&amp;rsquo;re reaching for them in a crisis.&lt;/p>
&lt;p>When our daughter Natasha was born, she wouldn&amp;rsquo;t latch on, and breastfeeding wasn&amp;rsquo;t working. Lena and I tried everything, exhausted and running on no sleep, seriously weighing up whether to just switch to formula. At some point we gave up trying to solve it ourselves and called our friends Tom and Rachel. We ended up talking for somewhere between four and six hours straight. We didn&amp;rsquo;t need advice, really — we needed to not be alone with it. And something about that phone call, about calming down, settled Natasha too. By the end of it she was feeding. I still think about that night. We&amp;rsquo;d been treating a human problem like a technical problem, trying to fix it in isolation, when all we&amp;rsquo;d actually needed was to pick up the phone.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="what-women-already-know">What Women Already Know&lt;/h3>
&lt;br>
&lt;p>Women are better at friendship than men. Not universally, not as a stereotype — but as a pattern that shows up consistently enough to be worth talking about honestly.&lt;/p>
&lt;p>The research on women&amp;rsquo;s friendship patterns tells part of the story. Women&amp;rsquo;s friendships tend to orient toward direct emotional disclosure — what psychologists call face-to-face friendship. Personal sharing, vulnerability, actual conversation about what&amp;rsquo;s happening in your life. Men&amp;rsquo;s friendships tend to be shoulder-to-shoulder — bonding through shared activity, talking while doing something else together. Neither mode is better. But when the activity disappears — when the sport stops, when the shared project ends, when the workplace changes — shoulder-to-shoulder friendship has nothing to stand on. Face-to-face friendship survives the context change because the context was never the point.&lt;/p>
&lt;p>What women understand, and what most men have to learn later and harder, is that the friendship is the thing. Not the activity. The person.&lt;/p>
&lt;p>The good news is that the shoulder-to-shoulder model doesn&amp;rsquo;t have to disappear. The Queenscliff table is shoulder-to-shoulder. The hack is building ritual around it — making the activity a recurring container for the friendship, rather than the friendship being a side effect of the activity. When you have the ritual, the friendship survives when everything else changes.&lt;/p>
&lt;p>I saw this firsthand when I was working in India. Some colleagues (Sarasa, Nidhi, Abhilash, Deepak) took me out to play badminton — I&amp;rsquo;d never played before, and I was genuinely terrible. But what struck me wasn&amp;rsquo;t the sport. It was that this group did it every week, without fail. They weren&amp;rsquo;t there to compete or keep score. They were there to be together, to unwind, to laugh at each other and at themselves. What I remember is leaving that court feeling like I&amp;rsquo;d glimpsed something. These people had built a ritual, and the ritual had built them.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="what-the-harvard-longevity-study-actually-found">What the Harvard Longevity Study Actually Found&lt;/h3>
&lt;br>
&lt;p>Harvard physician Arlie Bock began following the lives of a few hundred Harvard sophomores. The study is now 85 years old, has followed over 1,300 people including the children of original participants, and is the longest-running study of adult development in history.&lt;/p>
&lt;p>The fourth director of the study, psychiatrist Robert Waldinger, gave what became one of the ten most-watched TED talks ever. And the headline finding, distilled from nearly nine decades of data, is deceptively simple:&lt;/p>
&lt;p>&lt;strong>Good relationships lead to health and happiness. Not wealth. Not professional achievement. Not optimised nutrition or fitness metrics. Relationships.&lt;/strong>&lt;/p>
&lt;p>Waldinger and his co-author Marc Schulz published the full findings in &lt;em>The Good Life&lt;/em> in 2023. Two things from that research hit particularly hard.&lt;/p>
&lt;p>First: relationship satisfaction at age 50 was a better predictor of physical health at 80 than cholesterol levels. That&amp;rsquo;s the kind of finding that should make every analytically-minded person stop and recalibrate. We spend enormous energy optimising for things we can measure while ignoring the metric that turns out to be most predictive of how we age and how long we live.&lt;/p>
&lt;p>Second: when participants reached their 80s, their biggest regret was almost universally the same. Too much time at work. Not enough time with the people they loved. And their proudest achievements were almost entirely relational — being a good parent, a good friend, a good partner, a good mentor.&lt;/p>
&lt;p>George Vaillant, who directed the study for over three decades, summarised it: &amp;ldquo;When the study began, nobody cared about empathy or attachment. But the key to healthy aging is relationships, relationships, relationships.&amp;rdquo;&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="what-we-actually-traded">What We Actually Traded&lt;/h3>
&lt;br>
&lt;p>Reflecting on how high achievers use the word &amp;ldquo;sacrifice&amp;rdquo;. When we say we&amp;rsquo;ve sacrificed something for our career, we shouldn&amp;rsquo;t be afraid to put a name to who that sacrifice was. Because often it was the people in our lives that we call friends.&lt;/p>
&lt;p>Put a name to it.&lt;/p>
&lt;p>We&amp;rsquo;re not sacrificing some abstract concept. We&amp;rsquo;re sacrificing specific people. Former colleagues whose numbers we keep but never dial. School friends who live in the same city and see us once a year if they&amp;rsquo;re lucky. The people who would drop everything if we called — but who we don&amp;rsquo;t call, because we&amp;rsquo;re heads-down on the next deliverable.&lt;/p>
&lt;p>So many of us have cancelled on friends because a meeting came up, telling ourselves they&amp;rsquo;ll understand. Yet the reverse almost never happens — we wouldn&amp;rsquo;t reschedule a meeting for a friend. We&amp;rsquo;ve quietly trained ourselves to treat friendship as the flexible commitment, the one that can move.&lt;/p>
&lt;p>And the friendship debt is unlike technical debt — you can&amp;rsquo;t see it slowing you down, it doesn&amp;rsquo;t show up in velocity metrics or retrospectives. It just quietly accumulates until something breaks. The incident you pushed through alone. The difficult conversation at home that had nowhere to go because there were no friends to process it with. The creeping sense that despite being deeply capable and professionally respected, you&amp;rsquo;re not quite sure who you&amp;rsquo;d call.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="ubuntu-a-different-operating-system">Ubuntu: A Different Operating System&lt;/h3>
&lt;br>
&lt;p>There&amp;rsquo;s a Zulu/Nguni philosophy called Ubuntu. The phrase &amp;ldquo;Umuntu ngumuntu ngabantu&amp;rdquo; translates roughly as &amp;ldquo;a person is a person through other people.&amp;rdquo; You&amp;rsquo;ll also hear it rendered as &amp;ldquo;I am because we are.&amp;rdquo;&lt;/p>
&lt;p>Desmond Tutu described Ubuntu this way: &amp;ldquo;In our African worldview, we need other human beings for us to learn how to be human. For none of us comes fully formed into the world. &lt;strong>The solitary human being is a contradiction in terms.&lt;/strong>&amp;rdquo;&lt;/p>
&lt;p>He contrasted it explicitly with Descartes: &amp;ldquo;It is not &amp;lsquo;I think therefore I am.&amp;rsquo; It says rather: &amp;lsquo;I am human because I belong. I participate. I share.&amp;rsquo;&amp;rdquo;&lt;/p>
&lt;p>Most of the professional identity we build in data engineering runs on Descartes. I architect, therefore I am. I optimise the pipeline, therefore I am. I resolved the incident, therefore I am. We define ourselves by what we produce, alone, with headphones on, in a flow state that the rest of the world is just interrupting.&lt;/p>
&lt;p>Ubuntu is a different operating system. One where your identity is constituted by your relationships — where the question &amp;ldquo;who are you?&amp;rdquo; is answered not by your job title or your GitHub commit history, but by who you show up for and who shows up for you.&lt;/p>
&lt;p>The longest-running study of adult development in history arrived at the same conclusion Tutu was describing. Every major piece of research on social connection and mortality found it too. The architecture most of us are running on isn&amp;rsquo;t wrong in a subtle way. It&amp;rsquo;s wrong in the foundational way.&lt;/p>
&lt;p>Mark Shuttleworth — South African entrepreneur — named the Ubuntu Linux operating system after this philosophy in 2004, explicitly to emphasise community, open-source contribution, and the idea that the work is better when people build it together. Engineers understood that intuitively, because it maps to how the best software actually gets built. It also maps, as it turns out, to how the best lives get built.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-dragon-will-be-there-next-year">The Dragon Will Be There Next Year&lt;/h3>
&lt;br>
&lt;p>There&amp;rsquo;s a table in a rented house in Queenscliff, with a hand-drawn map on it and dice scattered at the edges.&lt;/p>
&lt;p>Every year, we make the trip. We gather around a table and roll dice and laugh at the missed sword swing and the impossible jump. The dragons change. The characters level up. The rules get argued over and amended and house-ruled into something nobody outside the group would recognise.&lt;/p>
&lt;p>But the people are constant. And the ritual is the point.&lt;/p>
&lt;p>It is more amazing to have an extraordinary experience with someone than by yourself. You can go somewhere alone and say &amp;ldquo;look what I did&amp;rdquo; — versus &amp;ldquo;do you remember that time we did that?&amp;rdquo;&lt;/p>
&lt;p>The 2am pipeline failure you resolved will be forgotten. By you. By your employer. By the Jira ticket that gets closed and archived.&lt;/p>
&lt;p>The dragon you fought with your friends — the copper golem that went down to a natural twenty on an impossible jump — that becomes a story. The story becomes a memory. The memory becomes part of who you are to each other and who they are to you.&lt;/p>
&lt;p>That&amp;rsquo;s not sentimentality. That&amp;rsquo;s the mechanism by which good lives actually get built. Eighty-five years of longitudinal data says so.&lt;/p>
&lt;p>Researcher Robin Dunbar found that building a close friendship takes roughly 200 hours of time together. Those hours don&amp;rsquo;t come from nowhere. They come from showing up, again and again, at the table. From protecting that time the same way you&amp;rsquo;d protect a critical production window. From understanding that the maintenance cost of friendship is the highest-ROI investment available to you as a human being.&lt;/p>
&lt;p>Your work will absorb your best years and move on when it&amp;rsquo;s convenient. Your friends will still be at the table.&lt;/p>
&lt;p>Put a name to what you&amp;rsquo;re protecting.&lt;/p>
&lt;p>Then protect it.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Leadership</category><category>Career Development</category><category>Personal Development</category><category>Leadership</category><category>Career Development</category><category>Mental Health</category><category>Burnout</category><category>Wellbeing</category><category>Friendship</category><category>Work-Life Balance</category></item><item><title>Your Data Model Isn't Broken, Part I: Why Refactoring Beats Rebuilding</title><link>https://ghostinthedata.info/posts/2026/2026-03-14-your-data-model-isnt-broken-part-1/</link><pubDate>Sat, 14 Mar 2026 09:00:00 +1100</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-03-14-your-data-model-isnt-broken-part-1/</guid><author>Chris Hillman</author><description>That fact table with 200 columns? Those bridge tables nobody understands? They're not bugs — they're reality encoded. Why the 'let's rebuild' instinct destroys more data teams than technical debt ever will.</description><content:encoded>&lt;p>In the early 2000&amp;rsquo;s - Netscape&amp;rsquo;s decision to rewrite their browser from scratch was the single worst strategic mistake a software company could make.&lt;/p>
&lt;/br>
&lt;p>At the time, Netscape was &lt;em>winning&lt;/em>. They had the dominant browser. They had market share. They had momentum. And then they decided the codebase was too messy, too tangled, too hard to work with — so they threw it all away and started over. Navigator 4.0 became the foundation for a rewrite that would eventually ship as version 6.0. There was no 5.0. Three years of development. No shipping product. And while Netscape&amp;rsquo;s engineers were busy building their beautiful new browser in a vacuum, Internet Explorer ate their lunch, their dinner, and most of their market share.&lt;/p>
&lt;p>It&amp;rsquo;s haunted me ever since: old code isn&amp;rsquo;t ugly because it&amp;rsquo;s bad. Old code is ugly because it &lt;em>works&lt;/em>. Every strange condition, every seemingly redundant check, every patch that makes a new developer wince — those are battle scars. Each one represents a bug that took weeks to find in production, a customer workflow nobody anticipated, or an edge case that only surfaces on the third Tuesday of months ending in &amp;ldquo;R.&amp;rdquo;&lt;/p>
&lt;p>I think about that every time I hear a data team say &amp;ldquo;let&amp;rsquo;s just rebuild it.&amp;rdquo;&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-sentence-that-starts-every-failed-data-project">The sentence that starts every failed data project&lt;/h3>
&lt;br>
&lt;p>I&amp;rsquo;ve been on data teams long enough to recognise the pattern. It always starts the same way. Someone — usually someone new, often someone senior — opens a dimensional model they didn&amp;rsquo;t build, scrolls through a few hundred lines of transformation logic, and says the words that should make every data leader&amp;rsquo;s blood run cold:&lt;/p>
&lt;p>&amp;ldquo;This is a mess. We need to start from scratch.&amp;rdquo;&lt;/p>
&lt;p>And look, I get it. I&amp;rsquo;ve felt that impulse myself, and likely have said the words myself. You open a fact table with 200 columns. You find bridge tables that reference other bridge tables. You discover a slowly-changing dimension nested inside another slowly-changing dimension, and you think: who built this? What were they &lt;em>thinking&lt;/em>?&lt;/p>
&lt;p>But here&amp;rsquo;s what I&amp;rsquo;ve learned the hard way, across multiple teams and more warehouse migrations than I care to count: they were thinking about the business. That bizarre WHERE clause filtering out both &amp;ldquo;Unknown&amp;rdquo; and &amp;ldquo;unknown&amp;rdquo;? That&amp;rsquo;s not sloppy code. That&amp;rsquo;s a case-sensitivity bug someone found in production data from a source system that nobody controlled. The seemingly redundant join that adds three seconds to your query? It handles a quarterly reconciliation edge case that cost the finance team two days of manual work before someone encoded the fix.&lt;/p>
&lt;p>That fact table with 200 columns isn&amp;rsquo;t a design failure. It&amp;rsquo;s an accurate representation of a business that has 200 things it needs to measure. The &amp;ldquo;clean&amp;rdquo; replacement model will eventually have 200 columns too — they&amp;rsquo;ll just have different names and it&amp;rsquo;ll take you eighteen months to figure out why you need them all.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="what-data-teams-keep-forgetting">What data teams keep forgetting&lt;/h3>
&lt;br>
&lt;p>This isn&amp;rsquo;t about browsers or even about code quality. It was about &lt;em>knowledge&lt;/em>. When you throw away a codebase and start fresh, you&amp;rsquo;re not just discarding syntax. You&amp;rsquo;re discarding years of accumulated understanding about how the real world actually behaves — understanding that was earned through production incidents, user complaints, and painful debugging sessions.&lt;/p>
&lt;p>Every collected bug fix in that old code represents something learned. Each fix might be just one line, a couple of characters even, but a lot of work and time went into figuring out those two characters were needed. And that knowledge — the &lt;em>why&lt;/em> behind the fix — almost never makes it into documentation. It lives in the code itself, or it lives nowhere.&lt;/p>
&lt;p>This is doubly true for data systems. A software application has unit tests, integration tests, user acceptance testing. A data model has&amp;hellip; what exactly? Row counts? Spot checks? A business user who eyeballs the dashboard and says &amp;ldquo;yeah, that looks about right&amp;rdquo;? The knowledge embedded in a mature data warehouse is far more fragile than application code, because the testing infrastructure around it is almost always weaker. Throw it away and you&amp;rsquo;re not just rebuilding a codebase — you&amp;rsquo;re rebuilding an institutional memory that was never written down in the first place.&lt;/p>
&lt;p>Fred Brooks saw this coming fifty years ago. His &amp;ldquo;Second System Effect&amp;rdquo; from &lt;em>The Mythical Man-Month&lt;/em> describes what happens when an engineer builds their second version of something: they over-design it. All the features they wisely deferred from the first version, all the architectural improvements they dreamed about, all the &amp;ldquo;if only we&amp;rsquo;d done it this way&amp;rdquo; ideas — they dump everything into the replacement. Brooks noted that a designer&amp;rsquo;s first system tends to be spare and clean because they know their limitations. The second system becomes a dumping ground for ambition.&lt;/p>
&lt;p>Sound familiar? Every data warehouse rebuild I&amp;rsquo;ve witnessed follows this arc. The team doesn&amp;rsquo;t just want to replicate what exists — they want to add a semantic layer, implement a medallion architecture, introduce data contracts, add real-time streaming, switch to Data Vault, build a self-serve analytics platform, and migrate to a new cloud provider. All at once. In a single initiative.&lt;/p>
&lt;p>And then they wonder why, two years later, they&amp;rsquo;ve delivered nothing and the business is still running reports off the &amp;ldquo;legacy&amp;rdquo; warehouse that was supposed to be decommissioned eighteen months ago.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-teradata-to-snowflake-reality-check">The Teradata-to-Snowflake reality check&lt;/h3>
&lt;br>
&lt;p>If you want to see the rebuild-vs-refactor debate play out in real time, watch a Teradata-to-Snowflake migration. The pattern is remarkably consistent.&lt;/p>
&lt;p>The pitch is compelling: Snowflake is cheaper, scales elastically, separates compute from storage, and runs standard SQL. Moving from Teradata should be straightforward — convert the SQL, move the data, validate the results. Easy, right?&lt;/p>
&lt;p>Roland Wenzlofsky, a Snowflake Solutions Architect, wrote about this in early 2026 and his observations line up perfectly with what I&amp;rsquo;ve seen. Most organisations approach these migrations as a translation exercise. The technical migration succeeds. Then the first quarterly cloud bill arrives at double the projected budget. Dashboards that loaded in seconds now queue for twenty minutes during batch windows. Pipelines that ran in 45 minutes on Teradata take five hours on Snowflake.&lt;/p>
&lt;p>The problem isn&amp;rsquo;t Snowflake. The problem is that Teradata professionals carry assumptions into Snowflake that aren&amp;rsquo;t just incomplete — they&amp;rsquo;re actively counterproductive. Distribution keys, join strategies, indexing patterns — all the hard-won optimisation knowledge from Teradata becomes a liability in a platform built on fundamentally different architecture.&lt;/p>
&lt;p>One of the largest Teradata-to-Snowflake migrations in North America involved 1.5 petabytes across 600 databases and roughly 45,000 objects. Before they could even begin, the team had to inventory every orphan object and document every behavioural difference between platforms. That&amp;rsquo;s not a &amp;ldquo;lift and shift.&amp;rdquo; That&amp;rsquo;s an archaeological dig.&lt;/p>
&lt;p>But here&amp;rsquo;s what I find most telling: even when the migration succeeds technically, the data quality problems survive the move perfectly intact. William Flaiz documented a healthcare organisation that spent $1.8 million migrating from Siebel to Salesforce. Not a single record lost. Zero downtime. Perfect data mapping. Post-migration, the sales team still couldn&amp;rsquo;t run accurate pipeline reports. The &amp;ldquo;opportunity stage&amp;rdquo; field contained 89 different values for what should have been six standard stages — including 1,247 records with the typo &amp;ldquo;Closeing Soon.&amp;rdquo; That typo migrated perfectly to the new $2.3 million platform. As Flaiz put it: when performance is bad, we assume the technology is the limiting factor. Data quality problems are messier. They implicate people, processes, training gaps, and years of accumulated shortcuts.&lt;/p>
&lt;p>New platform, same chaos. Because the platform was never the problem.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="chestertons-fence-or-why-you-shouldnt-delete-that-where-clause">Chesterton&amp;rsquo;s Fence, or: why you shouldn&amp;rsquo;t delete that WHERE clause&lt;/h3>
&lt;br>
&lt;p>There&amp;rsquo;s a principle in philosophy that every data engineer should have tattooed somewhere visible. G.K. Chesterton wrote it in 1929, and it goes roughly like this: if you come across a fence in the middle of a field and can&amp;rsquo;t see what purpose it serves, don&amp;rsquo;t tear it down. Go away and figure out why someone built it. Once you understand the reason, &lt;em>then&lt;/em> you can decide whether it still needs to be there.&lt;/p>
&lt;p>In data engineering, the fences are everywhere. That &lt;code>WHERE status NOT IN ('Unknown', 'unknown', 'UNKNOWN')&lt;/code> clause. That filter excluding records from a specific date range in 2019. The join to a reference table that only has 12 rows and hasn&amp;rsquo;t been updated in three years. They all look pointless until you remove one and discover that the finance reconciliation breaks, or that a regulatory report starts including test transactions that were supposed to be filtered out, or that a dashboard starts showing a revenue spike from a data migration artifact that happened four years ago.&lt;/p>
&lt;p>Hyrum Wright — formerly at Google, now at Adobe — formalised a related idea that&amp;rsquo;s become known as Hyrum&amp;rsquo;s Law: with enough users of an API, every observable behaviour will be depended on by somebody, regardless of what the documentation promises. For data systems, this is an absolute nightmare during rebuilds. Your downstream consumers don&amp;rsquo;t just depend on documented schemas. They depend on output ordering, null handling patterns, timestamp precision, and format quirks that were never specified anywhere. You rebuild the pipeline, and suddenly a report that&amp;rsquo;s worked for three years breaks — not because the data is wrong, but because the columns come back in a different order and someone&amp;rsquo;s Excel macro was hardcoded to column positions.&lt;/p>
&lt;p>I&amp;rsquo;ve seen teams spend more time debugging these &amp;ldquo;invisible contracts&amp;rdquo; after a rebuild than they would have spent refactoring the original system over two years.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-psychology-that-makes-us-do-it-anyway">The psychology that makes us do it anyway&lt;/h3>
&lt;br>
&lt;p>So if the evidence against big-bang rewrites is this overwhelming — and it is, from Spolsky to Brooks to McKinsey studies showing that large IT projects run 45% over budget and deliver 56% less value than predicted — why do smart, experienced data teams keep proposing them?&lt;/p>
&lt;p>Because the impulse to rebuild isn&amp;rsquo;t rational. It&amp;rsquo;s psychological. And once you understand the cognitive biases at play, you start seeing them everywhere.&lt;/p>
&lt;p>&lt;strong>The planning fallacy&lt;/strong> is the first culprit. Kahneman and Tversky showed that humans systematically underestimate how long tasks will take, even when they have direct experience with similar tasks that took longer than expected. In controlled studies, only 13% of subjects finished their project by the time they&amp;rsquo;d assigned a 50% probability of completion. The Sydney Opera House was estimated at AU$7 million and four years; it was delivered at AU$102 million and fourteen years. Every data warehouse rebuild proposal I&amp;rsquo;ve seen exhibits this same delusional optimism. &amp;ldquo;Six months, maybe nine.&amp;rdquo; It&amp;rsquo;s never six months.&lt;/p>
&lt;p>And it gets worse, because of what one writer calls the Achilles Paradox of rewrites: while you&amp;rsquo;re building the new version, features keep getting added to the old one. The business doesn&amp;rsquo;t stop generating requirements just because you&amp;rsquo;ve decided to rebuild. So the target keeps moving. The new system is perpetually six months from matching the functionality of the old one.&lt;/p>
&lt;p>&lt;strong>Not Invented Here syndrome&lt;/strong> is the second bias, and it&amp;rsquo;s the one that nobody wants to admit. Katz and Allen&amp;rsquo;s research on R&amp;amp;D project groups found that teams with stable composition develop a decreasing ability to absorb ideas from outside the group over time. They start believing they possess a monopoly on knowledge in their domain. The diagnostic sign? What researchers call &amp;ldquo;thought-terminating clichés&amp;rdquo; — phrases like &amp;ldquo;we already tried that&amp;rdquo; or &amp;ldquo;our situation is different.&amp;rdquo;&lt;/p>
&lt;p>In data teams, NIH manifests as building custom orchestration frameworks instead of using Airflow. Writing bespoke quality checks instead of adopting Great Expectations or dbt tests. Designing proprietary transformation layers instead of using tools that thousands of other teams have battle-tested. And most relevantly: insisting that the existing data model is fundamentally broken and needs to be replaced with something designed in-house from the ground up.&lt;/p>
&lt;p>&lt;strong>The new leader rewrite&lt;/strong> is the third pattern, and it might be the most destructive. A new head of data arrives. They look at the legacy warehouse they&amp;rsquo;ve inherited. They don&amp;rsquo;t understand the history behind any of the design decisions. They don&amp;rsquo;t know about the regulatory edge case that explains the weird date filter, or the source system quirk that necessitates the redundant join. All they see is complexity that they didn&amp;rsquo;t create — and complexity you didn&amp;rsquo;t create always looks worse than complexity you did.&lt;/p>
&lt;p>So they propose a rebuild. It&amp;rsquo;s partly strategic — they want to put their stamp on the architecture. And it&amp;rsquo;s partly genuine — they really do think they can do better. But they&amp;rsquo;re falling prey to the same illusion Spolsky identified: reading code is harder than writing it, and unfamiliar code always looks worse than it is.&lt;/p>
&lt;p>There&amp;rsquo;s a Stack Overflow survey finding that sits underneath all of this: &amp;ldquo;feeling unproductive&amp;rdquo; was the number one cause of developer unhappiness at 45%. Working with legacy systems is slow. It&amp;rsquo;s frustrating. Progress is incremental and often invisible. A greenfield rebuild, by contrast, feels &lt;em>amazing&lt;/em> — for the first few months. You&amp;rsquo;re making decisions, building things, moving fast. The architecture is clean and the tests pass and the world makes sense. And then reality seeps in, the edge cases accumulate, and eighteen months later you&amp;rsquo;ve got a codebase that looks suspiciously like the one you replaced.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-complexity-was-always-essential">The complexity was always essential&lt;/h3>
&lt;br>
&lt;p>Fred Brooks drew a distinction that I think about constantly. He separated &lt;strong>essential complexity&lt;/strong> — complexity that&amp;rsquo;s inherent to the problem domain and can&amp;rsquo;t be removed by better engineering — from &lt;strong>accidental complexity&lt;/strong>, which comes from poor implementation choices and can theoretically be eliminated.&lt;/p>
&lt;p>The truth about mature data warehouses: most of the complexity is essential. It&amp;rsquo;s not there because the original engineers were incompetent. It&amp;rsquo;s there because the business is genuinely that complicated.&lt;/p>
&lt;p>The transformation handling fifteen date formats? That&amp;rsquo;s because you have fifteen source systems and none of them agree on how to represent a date. The slowly-changing dimension with Type 2 tracking on attributes that seem trivial? Someone in compliance needed audit trails on those fields after a regulatory inquiry. The bridge table that connects customers to accounts through an intermediate entity that nobody can explain? It models a many-to-many relationship that emerged when the company acquired a subsidiary with a different customer hierarchy.&lt;/p>
&lt;p>A rebuild won&amp;rsquo;t make this complexity disappear. It&amp;rsquo;ll just redistribute it. Instead of one tangled fact table, you&amp;rsquo;ll have twelve microservices each handling a piece of the logic. Instead of one confusing WHERE clause, you&amp;rsquo;ll have business rules scattered across a semantic layer, a transformation layer, and a data quality framework. The total complexity will be identical — or worse, because now it&amp;rsquo;s distributed across more systems with more failure modes.&lt;/p>
&lt;p>Ward Cunningham — the person who coined the term &amp;ldquo;technical debt&amp;rdquo; — has said he wishes he&amp;rsquo;d used the word &amp;ldquo;opportunity&amp;rdquo; instead. His original metaphor wasn&amp;rsquo;t about sloppy code at all. It was about the gap between your code&amp;rsquo;s current model of the problem domain and your team&amp;rsquo;s evolved understanding of that domain. Technical debt, properly understood, is &lt;em>learning&lt;/em> that hasn&amp;rsquo;t been applied yet. The legacy data model doesn&amp;rsquo;t represent failure. It represents the team&amp;rsquo;s best understanding at the time it was built, plus every correction that production reality demanded afterward.&lt;/p>
&lt;p>Refactoring respects this. It asks: which parts of this complexity are essential (keep them, clarify them, test them) and which parts are accidental (remove them incrementally, one safe step at a time)? It preserves institutional knowledge while improving structure. It delivers value continuously instead of asking the business to wait years for a payoff that — statistically, based on everything we know about large-scale IT projects — probably won&amp;rsquo;t arrive.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="so-what-do-you-do-instead">So what do you do instead?&lt;/h3>
&lt;br>
&lt;p>I&amp;rsquo;m going to save the detailed playbook for Part II of this series, but the short version is this: you refactor. Methodically. Incrementally. With tests.&lt;/p>
&lt;p>Martin Fowler and Pramod Sadalage proved that database refactoring is possible back in 2006, when conventional wisdom said it couldn&amp;rsquo;t be done. Their principle was simple: make each change as small as possible, because the pain of integration increases exponentially with the size of the integration.&lt;/p>
&lt;p>dbt has turned this into a practical workflow for analytics engineering teams. Bring your legacy SQL in unchanged. Wrap it in a model. Verify the output matches. Then — and only then — start decomposing. Extract CTEs. Introduce staging layers. Build tests. Refactor one model at a time, auditing each change against the original output. It&amp;rsquo;s not glamorous. It doesn&amp;rsquo;t let you redesign everything from first principles. But it works, and it works without asking the business to lose access to their reports for six months.&lt;/p>
&lt;p>The Strangler Fig pattern, the Write-Audit-Publish pattern, the Expand and Contract pattern — these are all variations on the same theme: replace incrementally, verify continuously, and never throw away working logic until you&amp;rsquo;ve proved the replacement handles every case the original did.&lt;/p>
&lt;p>I&amp;rsquo;ll dig into each of these in Part II. For now, the point is simpler.&lt;/p>
&lt;hr>
&lt;/br>
&lt;/br>
&lt;h3 id="the-scars-are-data">The scars are data&lt;/h3>
&lt;br>
&lt;p>When I look at a legacy data model — really look at it, with patience and without ego — I don&amp;rsquo;t see a mess anymore. I see a record of everything a business has learned about itself. Every weird join is a relationship somebody fought to understand. Every cryptic transformation is a business rule that was discovered through painful experience. Every seemingly arbitrary filter is a production incident that someone fixed at some point, probably at an hour they&amp;rsquo;d rather not remember.&lt;/p>
&lt;p>The Netscape rewrite remains one of the most studied failures in software history. And yet here we are, in 2026, watching data teams propose the same mistake — just with different technology names. Teradata becomes Snowflake. Informatica becomes dbt. The on-prem warehouse becomes the cloud lakehouse. The pitch changes, but the impulse doesn&amp;rsquo;t: tear it down, start over, do it right this time.&lt;/p>
&lt;p>My view? Your data model isn&amp;rsquo;t broken. It&amp;rsquo;s battle-tested. It&amp;rsquo;s messy because reality is messy. And the right response to inherited complexity isn&amp;rsquo;t demolition — it&amp;rsquo;s archaeology. Understand what&amp;rsquo;s there. Understand &lt;em>why&lt;/em> it&amp;rsquo;s there. Then improve it, one careful change at a time.&lt;/p>
&lt;p>The scars in your data model aren&amp;rsquo;t flaws. They&amp;rsquo;re knowledge. Treat them accordingly.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Data Modelling</category><category>Data Engineering</category><category>Refactoring</category><category>Data Warehousing</category><category>Technical Debt</category><category>Snowflake</category><category>dbt</category><category>Legacy Systems</category><category>Data Quality</category></item><item><title>12 Steps to Better Data Engineering</title><link>https://ghostinthedata.info/posts/2026/2026-03-07-twelve-steps-to-better-data-engineering/</link><pubDate>Sat, 07 Mar 2026 09:00:00 +1100</pubDate><guid>https://ghostinthedata.info/posts/2026/2026-03-07-twelve-steps-to-better-data-engineering/</guid><author>Chris Hillman</author><description>A quick scoring framework to assess your data team's engineering maturity. Twelve yes-or-no questions, each with concrete benchmarks for what good and amazing look like using dbt, Snowflake, GitHub Actions, and AWS.</description><content:encoded>&lt;p>Let me tell you about the moment I stopped trusting architecture diagrams.&lt;/p>
&lt;p>I was three days into a new role, getting up to speed with the data team. Smart people. Modern stack. On paper, everything looked right. They walked me through a beautiful data platform diagram: clean lines, labelled layers, colour-coded domains. It looked like something you&amp;rsquo;d see in a data conference.&lt;/p>
&lt;p>Then I asked a question that changed everything: &amp;ldquo;Can you rebuild your finance table from scratch right now?&amp;rdquo;&lt;/p>
&lt;p>The room went quiet. One person started explaining that it was &amp;ldquo;mostly possible&amp;rdquo; but there were &amp;ldquo;a few manual steps&amp;rdquo; and &amp;ldquo;some seeds that someone uploads&amp;rdquo;. Someone else mentioned an incremental model that &amp;ldquo;hasn&amp;rsquo;t been full-refreshed in about eight months because last time it broke.&amp;rdquo;&lt;/p>
&lt;p>That gap — between the diagram on the wall and the reality in the warehouse — is where I&amp;rsquo;ve spent most of my career. And over the years, working across team after team, I started noticing something. The teams that were struggling weren&amp;rsquo;t missing knowledge. They had smart engineers who knew about testing, CI/CD, documentation, data contracts. They&amp;rsquo;d read the blog posts. They&amp;rsquo;d watched the conference talks. What they didn&amp;rsquo;t have was a way to honestly assess where they stood and what to fix first.&lt;/p>
&lt;p>I kept wishing someone would build a simple diagnostic. Maybe they just needed a handful of honest questions that would tell you, in ten minutes, whether your team was engineering or firefighting.&lt;/p>
&lt;p>The moment that turned wishing into doing came about six months later, with a team I was leading. We implemented just two practices — CI testing on pull requests and mandatory code review — and I watched the compound effect over twelve weeks. Our incidents dropped. The analysts stopped asking &amp;ldquo;which table do I use?&amp;rdquo; A new hire shipped a model on day four. One of the senior engineers pulled me aside and said, &amp;ldquo;I wish we&amp;rsquo;d known how bad things were before you got here. We thought we were fine because nothing was on fire.&amp;rdquo;&lt;/p>
&lt;p>Nothing was on fire. But everything was smouldering.&lt;/p>
&lt;p>That&amp;rsquo;s why I built this test. Not because the world needs another framework — it doesn&amp;rsquo;t. But because the gap between &lt;em>thinking you&amp;rsquo;re fine&lt;/em> and &lt;em>knowing where you stand&lt;/em> is where data teams lose months of momentum. And the only way to close that gap is to ask questions honest enough that the answers sting a little.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="the-twelve-questions">The twelve questions&lt;/h3>
&lt;br>
&lt;p>Before we get into the detail, here they are. Score yourself. One point for each &amp;ldquo;yes.&amp;rdquo;&lt;/p>
&lt;ol>
&lt;li>Can you rebuild any table from raw data in one command?&lt;/li>
&lt;li>Do you have a data catalog that people actually use?&lt;/li>
&lt;li>Can a new analyst find the data they need without asking you?&lt;/li>
&lt;li>Do you test transformations before deploying them?&lt;/li>
&lt;li>Do you fix data quality issues before building new pipelines?&lt;/li>
&lt;li>Do you have SLAs for your critical tables?&lt;/li>
&lt;li>Do you have a single source of truth for business definitions?&lt;/li>
&lt;li>Can you explain the lineage of any metric in under 2 minutes?&lt;/li>
&lt;li>Do data producers know when they break downstream consumers?&lt;/li>
&lt;li>Do you do code review on SQL and dbt models?&lt;/li>
&lt;li>Do new hires build a real pipeline in their first week?&lt;/li>
&lt;li>Do you regularly talk to the people who actually use your data?&lt;/li>
&lt;/ol>
&lt;p>&lt;strong>A score of 12 is unicorn territory.&lt;/strong> I&amp;rsquo;ve never seen it in the wild. Most teams I&amp;rsquo;ve worked with score 3 or 4. Anything below 6 and you&amp;rsquo;re not doing data engineering — you&amp;rsquo;re doing data triage.&lt;/p>
&lt;p>Now let&amp;rsquo;s break each one down. For every question, I&amp;rsquo;ll show you what &lt;em>good&lt;/em> looks like and what &lt;em>amazing&lt;/em> looks like — with concrete implementations using dbt, Snowflake, GitHub Actions, and AWS. Because &amp;ldquo;yes&amp;rdquo; isn&amp;rsquo;t binary in practice. There&amp;rsquo;s a massive gap between &amp;ldquo;sort of&amp;rdquo; and &amp;ldquo;absolutely.&amp;rdquo;&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="1-can-you-rebuild-any-table-from-raw-data-in-one-command">1. Can you rebuild any table from raw data in one command?&lt;/h3>
&lt;br>
&lt;p>This is the foundation. If you can&amp;rsquo;t reproduce your outputs from your inputs, you don&amp;rsquo;t have a pipeline — you have a prayer.&lt;/p>
&lt;p>The core principle here is &lt;strong>idempotency&lt;/strong>: running the same operation twice produces identical results. dbt is designed around this idea, but achieving it requires deliberate choices, especially with incremental models. Too many teams build incremental models that silently drift from what a full refresh would produce, and they don&amp;rsquo;t discover the discrepancy until something breaks spectacularly.&lt;/p>
&lt;h4 id="what-good-looks-like">What good looks like&lt;/h4>
&lt;p>Individual tables rebuild cleanly with &lt;code>dbt run --full-refresh --select model_name&lt;/code>. Your incremental models are tested against full refreshes during development. You&amp;rsquo;ve configured &lt;code>on_schema_change: sync_all_columns&lt;/code> so schema evolution doesn&amp;rsquo;t silently break things. And you&amp;rsquo;ve got a CI pipeline in GitHub Actions that runs &lt;code>dbt build&lt;/code> on pull requests.&lt;/p>
&lt;p>That&amp;rsquo;s good. That&amp;rsquo;s better than most. But it&amp;rsquo;s still table-by-table.&lt;/p>
&lt;h4 id="what-amazing-looks-like">What amazing looks like&lt;/h4>
&lt;p>Your entire warehouse rebuilds from raw data using a &lt;strong>Write-Audit-Publish (WAP) pattern&lt;/strong>. If you&amp;rsquo;re not familiar with WAP, I wrote a &lt;a href="https://ghostinthedata.info/posts/2025/2025-05-18-wap-data-pipelines/" target="_blank" rel="noopener">deep dive on implementing it with Airflow&lt;/a> — the core idea is that new data gets written to a staging environment, audited against quality checks, and only published to production once it passes. It&amp;rsquo;s the data engineering equivalent of a preflight checklist: nothing reaches consumers until it&amp;rsquo;s been verified. And if you&amp;rsquo;re on Snowflake with Iceberg tables, the game has changed — &lt;a href="https://ghostinthedata.info/posts/2026/2026-02-27-wap-iceberg-branching/" target="_blank" rel="noopener">WAP with Iceberg branching&lt;/a> gives you git-like isolation at the table level, which means you can stage, audit, and publish without creating separate schemas or databases at all.&lt;/p>
&lt;p>The practical implementation: dbt always targets a staging schema or database (or an Iceberg branch), runs all models and tests, and only on success does the data get promoted to production. Combined with Snowflake&amp;rsquo;s &lt;strong>zero-copy cloning&lt;/strong>, this becomes practical even at scale — cloning a multi-terabyte database takes seconds and costs nothing in storage. You&amp;rsquo;re not duplicating data; you&amp;rsquo;re creating metadata pointers.&lt;/p>
&lt;p>And when things &lt;em>do&lt;/em> go wrong — and they will — the question becomes how quickly you can recover. If your incremental models have drifted, or a source system silently changed its schema three weeks ago, you need a backfill strategy that doesn&amp;rsquo;t involve rebuilding everything from scratch. I wrote about the &lt;a href="https://ghostinthedata.info/posts/2026/2026-02-07-self-healing/" target="_blank" rel="noopener">pitfalls of day-by-day backfills and how to heal tables properly&lt;/a> — it&amp;rsquo;s the companion piece to idempotency, because reproducibility isn&amp;rsquo;t just about building forward. It&amp;rsquo;s about being able to go back.&lt;/p>
&lt;p>The really mature teams take it further with &lt;strong>Slim CI&lt;/strong>. Instead of rebuilding everything on every PR, they store the production &lt;code>manifest.json&lt;/code> in S3 and compare it against the new build. Only modified models and their downstream dependencies get rebuilt:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># .github/workflows/dbt-ci.yml&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">on&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">pull_request&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">types&lt;/span>: [&lt;span style="color:#ae81ff">opened, reopened, synchronize]&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">steps&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">Download production manifest&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">run&lt;/span>: &lt;span style="color:#ae81ff">aws s3 cp s3://dbt-artifacts/manifest.json ./state/manifest.json&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">Run dbt build (modified models only)&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">run&lt;/span>: &lt;span style="color:#ae81ff">dbt build -s &amp;#39;state:modified+&amp;#39; --defer --state ./state --target ci&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>One team I worked with reported a &lt;strong>30% reduction in compute costs&lt;/strong> just by splitting their monolithic CI job into targeted workflows.&lt;/p>
&lt;blockquote>
&lt;p>&lt;strong>Watch out for this:&lt;/strong> Protect massive tables with &lt;code>full_refresh = false&lt;/code> in your dbt config. One accidental &lt;code>--full-refresh&lt;/code> on a billion-row fact table can ruin your morning. Override it with a variable when you actually mean it: &lt;code>full_refresh = var(&amp;quot;force_full_refresh&amp;quot;, false)&lt;/code>.&lt;/p>&lt;/blockquote>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="2-do-you-have-a-data-catalog-that-people-actually-use">2. Do you have a data catalog that people actually use?&lt;/h3>
&lt;br>
&lt;p>The emphasis is on &amp;ldquo;actually use.&amp;rdquo; I&amp;rsquo;ve lost count of how many teams have shown me a data catalog, only for me to discover that nobody&amp;rsquo;s opened it in months. The data catalog is the gym membership of the data world — everyone has one, nobody goes.&lt;/p>
&lt;p>The failure mode is almost always the same: a team selects a tool, loads metadata, declares victory, and moves on. Six months later it&amp;rsquo;s a ghost town of stale descriptions and orphaned tables.&lt;/p>
&lt;h4 id="what-good-looks-like-1">What good looks like&lt;/h4>
&lt;p>dbt Docs generated and hosted somewhere accessible. Your major tables have descriptions in &lt;code>schema.yml&lt;/code>. The lineage graph exists and gets pulled up occasionally during incident response or onboarding. It&amp;rsquo;s not perfect, but it&amp;rsquo;s there.&lt;/p>
&lt;h4 id="what-amazing-looks-like-1">What amazing looks like&lt;/h4>
&lt;p>The catalog is the &lt;strong>default entry point&lt;/strong> for data questions — not Microsoft Teams, not the data engineer sitting three desks over.&lt;/p>
&lt;p>Spotify&amp;rsquo;s internal tool Lexikon pushed data scientist adoption from 75% to 95% by doing something clever: they added &lt;em>personalised dataset recommendations&lt;/em>, people/team pages showing who uses each dataset, common joins, and popular fields. It became a top-five internal tool because it was genuinely faster than asking a colleague.&lt;/p>
&lt;p>Airbnb&amp;rsquo;s Metis serves over 1,000 data users weekly. Their secret? A Google-like search interface that works for every skill level, from SQL-fluent engineers to product managers who&amp;rsquo;ve never written a query.&lt;/p>
&lt;p>The pattern that actually works for smaller teams: author descriptions in whatever UI is easiest, auto-generate a PR back to dbt&amp;rsquo;s &lt;code>schema.yml&lt;/code> daily, review in Git. This keeps your definitions version-controlled without making people learn Git just to document a table.&lt;/p>
&lt;p>And here&amp;rsquo;s the leading indicator that your catalog is working: &lt;strong>track the decline in Microsoft Teams questions about data.&lt;/strong> Productboard used exactly this metric. Fewer &amp;ldquo;which table do I use for revenue?&amp;rdquo; messages meant the catalog was earning its keep.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="3-can-a-new-analyst-find-the-data-they-need-without-asking-you">3. Can a new analyst find the data they need without asking you?&lt;/h3>
&lt;br>
&lt;p>This is the catalog question&amp;rsquo;s evil twin. A catalog helps, but self-service discovery is really about the cumulative effect of naming conventions, documentation-as-code, semantic layers, and project structure.&lt;/p>
&lt;p>dbt Labs published guidance that should be printed and taped to every data engineer&amp;rsquo;s monitor: &lt;strong>&amp;ldquo;Assume your end-user will have no other context than the model name.&amp;rdquo;&lt;/strong> Model names persist across databases, BI tools, DAGs, and docs. Folder names and schemas don&amp;rsquo;t follow the data the same way.&lt;/p>
&lt;h4 id="what-good-looks-like-2">What good looks like&lt;/h4>
&lt;p>You&amp;rsquo;re using the canonical dbt naming convention: &lt;code>stg_&lt;/code> for staging, &lt;code>int_&lt;/code> for intermediate, &lt;code>fct_&lt;/code> for facts, &lt;code>dim_&lt;/code> for dimensions. Your project follows the three-layer structure — staging (1:1 with sources, materialised as views), intermediate (composable business logic), and marts (final business-ready datasets organised by domain). A reasonably technical person can navigate the DAG and find what they need within ten minutes.&lt;/p>
&lt;h4 id="what-amazing-looks-like-2">What amazing looks like&lt;/h4>
&lt;p>A new analyst finds and understands any dataset &lt;strong>within minutes&lt;/strong>, without messaging anyone.&lt;/p>
&lt;p>Every model has both table-level and column-level descriptions. Schema field search works across all tables — &amp;ldquo;find every table with &lt;code>customer_id&lt;/code>&amp;rdquo; returns results instantly. Personalised dataset recommendations surface relevant tables based on role.&lt;/p>
&lt;p>And the crown jewel: a &lt;strong>semantic layer&lt;/strong> that lets business users query metrics in plain language without writing SQL. The dbt Semantic Layer (powered by MetricFlow) centralises metric definitions in YAML alongside your models — entities, dimensions, and measures declared once, then consumed via API in Tableau, Power BI, Looker, Python notebooks, and increasingly, AI agents.&lt;/p>
&lt;p>Bilt Rewards reported an 80% decrease in data costs after implementing it for embedded analytics. And dbt Labs&amp;rsquo; testing showed 83% of natural-language questions answered correctly when routed through the semantic layer. That last number matters more than you&amp;rsquo;d think — it&amp;rsquo;s the difference between AI tooling that actually works and AI tooling that confidently gives the wrong answer.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="4-do-you-test-transformations-before-deploying-them">4. Do you test transformations before deploying them?&lt;/h3>
&lt;br>
&lt;p>This is probably the single highest-impact practice on the list. If I could only get a team to adopt one of these twelve, it would be this one.&lt;/p>
&lt;p>But here&amp;rsquo;s what most teams get wrong: they think &amp;ldquo;testing&amp;rdquo; means slapping &lt;code>not_null&lt;/code> on a few columns and calling it done. Real testing means understanding the &lt;a href="https://ghostinthedata.info/posts/2025/2025-11-24-data-quality-framework/" target="_blank" rel="noopener">dimensions of data quality&lt;/a> — completeness, uniqueness, timeliness, validity, accuracy, consistency — and building checks that cover each one deliberately. I&amp;rsquo;ve written about &lt;a href="https://ghostinthedata.info/posts/2025/2025-11-22-data-quality/" target="_blank" rel="noopener">why data quality is a deeper problem than most teams realise&lt;/a>, and the short version is this: testing in CI is where quality becomes a habit rather than a hope. But only if your tests are actually measuring the things that matter.&lt;/p>
&lt;p>The consensus across every practitioner I&amp;rsquo;ve studied is clear: &lt;strong>CI for dbt is the most impactful thing you can do for data quality.&lt;/strong> Three complementary testing layers form a comprehensive strategy: generic dbt tests (&lt;code>not_null&lt;/code>, &lt;code>unique&lt;/code>, &lt;code>accepted_values&lt;/code>, &lt;code>relationships&lt;/code>), unit tests (introduced in dbt v1.8 for testing complex business logic with static inputs), and data diffing for value-level comparison between production and development.&lt;/p>
&lt;h4 id="what-good-looks-like-3">What good looks like&lt;/h4>
&lt;p>Basic dbt tests on key models. Tests run as part of scheduled production jobs using &lt;code>dbt build&lt;/code> (which runs tests immediately after each model, not in a separate pass). Pull requests exist for dbt changes, and someone eyeballs them before merging.&lt;/p>
&lt;h4 id="what-amazing-looks-like-3">What amazing looks like&lt;/h4>
&lt;p>Full Slim CI with per-PR isolated Snowflake schemas, data diffing integrated into PR comments, and automated linting that catches style issues before a human ever looks at the code.&lt;/p>
&lt;p>The numbers from real teams are staggering. Thumbtack — 50-plus analysts, five data engineers, over 100 PRs per month — previously spent one to two hours manually validating each pull request with SQL queries and spreadsheets. After integrating data diffing into their GitHub CI pipeline, they saved over 200 hours per month. That&amp;rsquo;s not a rounding error. That&amp;rsquo;s a full-time engineer&amp;rsquo;s worth of capacity recovered.&lt;/p>
&lt;p>Dutchie caught timezone corruption on &lt;code>created_at&lt;/code> fields and case-when logic errors that were silently shifting 20% of data between columns. Nutrafol caught a transformation that would have shown net revenue plummeting — before it reached production and before the CFO&amp;rsquo;s Monday morning dashboard refreshed.&lt;/p>
&lt;p>Zscaler went even further. They built PRISM, a multi-agent AI PR review system that reduced manual review time by 90% — auto-approving conformant PRs and posting targeted comments on complex logic changes.&lt;/p>
&lt;p>Here&amp;rsquo;s what the mature pipeline looks like in practice:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># On every PR: lint, compile, build modified models, diff against prod&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">steps&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">SQLFluff lint&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">run&lt;/span>: &lt;span style="color:#ae81ff">sqlfluff lint models/ --dialect snowflake&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">dbt compile&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">run&lt;/span>: &lt;span style="color:#ae81ff">dbt compile&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">Slim CI build&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">run&lt;/span>: &lt;span style="color:#ae81ff">dbt build -s &amp;#39;state:modified+&amp;#39; --defer --state ./state --target ci&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">Data Diff&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">run&lt;/span>: &lt;span style="color:#ae81ff">datafold ci submit --ci-config-id ${{ secrets.DATAFOLD_CI_CONFIG }}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="5-do-you-fix-data-quality-issues-before-building-new-pipelines">5. Do you fix data quality issues before building new pipelines?&lt;/h3>
&lt;br>
&lt;p>Here&amp;rsquo;s an uncomfortable stat: Monte Carlo&amp;rsquo;s 2023 State of Data Quality report found that &lt;strong>74% of organisations reported that business stakeholders identify data quality issues first&lt;/strong> — up from 47% the prior year. Most data teams learn about broken data from the people who are supposed to trust it. That&amp;rsquo;s not a technology problem. That&amp;rsquo;s a culture problem.&lt;/p>
&lt;h4 id="what-good-looks-like-4">What good looks like&lt;/h4>
&lt;p>Basic dbt tests catch obvious issues. Tests run in production jobs. The team has an informal sense of which data matters most, usually earned through painful experience — someone got burned by bad revenue numbers in an executive review, and now those tables have tests.&lt;/p>
&lt;h4 id="what-amazing-looks-like-4">What amazing looks like&lt;/h4>
&lt;p>Data quality SLAs enforced with &lt;strong>error budgets&lt;/strong>, borrowed from the SRE playbook. A 99.5% data availability target allows approximately 3.6 hours of acceptable downtime per month. When the error budget is consumed, reliability work takes priority over feature delivery. Full stop. No new pipelines until you&amp;rsquo;ve earned back your quality margin.&lt;/p>
&lt;p>Warner Bros. Discovery created something they call a &lt;strong>Data Quality Forum&lt;/strong> — not a reactive incident-review meeting, but an operational bridge between teams. They used observability tooling to surface anomalies early, developed a priority matrix (P0/P1/P2), and made the forum&amp;rsquo;s patterns part of new hire onboarding. During Olympics livestreaming, custom SQL checks detected missing content metadata before it could break reporting.&lt;/p>
&lt;p>Snowflake now offers native &lt;strong>Data Metric Functions&lt;/strong> — &lt;code>FRESHNESS&lt;/code>, &lt;code>NULL_COUNT&lt;/code>, &lt;code>DUPLICATE_COUNT&lt;/code>, &lt;code>UNIQUE_COUNT&lt;/code>, &lt;code>ROW_COUNT&lt;/code> — that can be scheduled to run on DML changes or time intervals. Results land in &lt;code>SNOWFLAKE.LOCAL.DATA_QUALITY_MONITORING_RESULTS&lt;/code>. It&amp;rsquo;s not a replacement for dbt tests, but it provides a warehouse-level safety net that catches issues your transformation layer might miss.&lt;/p>
&lt;p>The cultural shift matters more than the tooling. HelloFresh&amp;rsquo;s VP of Data drove a transformation through three stages: ad-hoc individual fixes, organised cleanup by a central team, and finally proactive quality at the source. The final stage required embedding data product owners within business domains and running data literacy programs. It&amp;rsquo;s not glamorous work. But it&amp;rsquo;s the work that changes outcomes.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="6-do-you-have-slas-for-your-critical-tables">6. Do you have SLAs for your critical tables?&lt;/h3>
&lt;br>
&lt;p>&amp;ldquo;The dashboard is stale&amp;rdquo; is the Slack message that launches a thousand fire drills. SLAs turn that reactive scramble into a measured, prioritised response.&lt;/p>
&lt;p>The breakthrough insight from practitioners: &lt;strong>don&amp;rsquo;t aim for 100%.&lt;/strong> A 99.5% target gives you a realistic buffer for maintenance, edge cases, and the occasional Snowflake service hiccup. Different data products warrant different targets — which means you need a tiered classification system.&lt;/p>
&lt;h4 id="what-good-looks-like-5">What good looks like&lt;/h4>
&lt;p>Your team knows which tables are critical, usually because they&amp;rsquo;ve been burned before. Some &lt;code>dbt source freshness&lt;/code> checks are running. Basic Slack alerts fire when jobs fail. It&amp;rsquo;s reactive, but at least you know when things break.&lt;/p>
&lt;h4 id="what-amazing-looks-like-5">What amazing looks like&lt;/h4>
&lt;p>A formal tiered classification published in the data catalog with specific, measurable commitments:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Tier 1 (Gold):&lt;/strong> ML systems, revenue reporting → PagerDuty on-call, must have &lt;code>unique&lt;/code> + &lt;code>not_null&lt;/code> tests + assigned owner, freshness SLA under 4 hours&lt;/li>
&lt;li>&lt;strong>Tier 2 (Silver):&lt;/strong> Executive dashboards, KPI reports → Slack team channel alerts, owner assigned, freshness SLA under 12 hours&lt;/li>
&lt;li>&lt;strong>Tier 3 (Bronze):&lt;/strong> Ad-hoc analytics, exploration tables → weekly digest, freshness SLA under 24 hours&lt;/li>
&lt;/ul>
&lt;p>The implementation lives in dbt &lt;code>meta&lt;/code> config:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">models&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">fct_revenue&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">meta&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">owner&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;finance-data-team&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">criticality&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;tier_1&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">sla_freshness&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;4 hours&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">tests&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">unique&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">column_name&lt;/span>: &lt;span style="color:#ae81ff">order_id&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">not_null&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">column_name&lt;/span>: &lt;span style="color:#ae81ff">order_id&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>SLA compliance dashboards track attainment percentage, breach counts, and trends over time. When the error budget is exhausted, feature development freezes — same principle as practice #5. Dedicated Snowflake warehouses isolate critical workloads from exploratory queries so that someone&amp;rsquo;s &lt;code>SELECT *&lt;/code> on a billion-row table doesn&amp;rsquo;t cause your revenue report to miss its 8 AM deadline.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="7-do-you-have-a-single-source-of-truth-for-business-definitions">7. Do you have a single source of truth for business definitions?&lt;/h3>
&lt;br>
&lt;p>Everyone&amp;rsquo;s had this conversation: &amp;ldquo;Which revenue number is right?&amp;rdquo; Two dashboards, two numbers, two teams who each think the other is wrong. This is what happens when metric definitions live inside BI tools, tribal knowledge, and the head of that one analyst who&amp;rsquo;s been here since 2019. If you&amp;rsquo;ve ever worked with dimensional models, you&amp;rsquo;ll recognise this as the conformed dimension problem — I covered why &lt;a href="https://ghostinthedata.info/posts/2025/2025-11-07-effective-data-modelling/" target="_blank" rel="noopener">dimensional modeling still matters&lt;/a> and how conformed dimensions are the integration backbone that prevents exactly this kind of mess.&lt;/p>
&lt;h4 id="what-good-looks-like-6">What good looks like&lt;/h4>
&lt;p>Key metrics are defined in dbt docs or a wiki. The team knows which mart table is the canonical source for revenue, orders, or whatever your core entities are. When someone asks &amp;ldquo;which table do I use?&amp;rdquo;, there&amp;rsquo;s a consistent answer — even if it&amp;rsquo;s only communicated verbally.&lt;/p>
&lt;h4 id="what-amazing-looks-like-6">What amazing looks like&lt;/h4>
&lt;p>A centralised &lt;strong>semantic layer&lt;/strong> enforced as the only path to metrics. All definitions are version-controlled in dbt YAML, reviewed via PR, and validated in CI.&lt;/p>
&lt;p>Here&amp;rsquo;s what a metric definition looks like in MetricFlow:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">metrics&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">cancellation_rate&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">description&lt;/span>: &lt;span style="color:#e6db74">&amp;#34;Percentage of orders cancelled within 24 hours of placement&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">type&lt;/span>: &lt;span style="color:#ae81ff">ratio&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">type_params&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">numerator&lt;/span>: &lt;span style="color:#ae81ff">cancellations&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">denominator&lt;/span>: &lt;span style="color:#ae81ff">order_total&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">filter&lt;/span>: |&lt;span style="color:#e6db74">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#e6db74"> {{ TimeDimension(&amp;#39;metric_time&amp;#39;, &amp;#39;day&amp;#39;) }}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>That definition gets consumed identically whether someone queries it from Tableau, Power BI, a Python notebook, or an AI agent. One definition. One answer. Everywhere.&lt;/p>
&lt;p>Whatnot — the livestream marketplace growing at breakneck speed — solved a different version of this problem. They used Protobuf schemas as a single source of truth for event definitions, consolidating hundreds of chaotic Snowflake tables down to two clean &amp;ldquo;exposure&amp;rdquo; tables: &lt;code>backend_events&lt;/code> and &lt;code>frontend_events&lt;/code>. Brutal simplification. But it worked because everyone could find the data.&lt;/p>
&lt;p>&lt;code>dbt sl validate&lt;/code> in CI catches breaking changes to semantic definitions before merge. That&amp;rsquo;s the key differentiator — your metric definitions aren&amp;rsquo;t just documented, they&amp;rsquo;re &lt;em>enforced&lt;/em>.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="8-can-you-explain-the-lineage-of-any-metric-in-under-2-minutes">8. Can you explain the lineage of any metric in under 2 minutes?&lt;/h3>
&lt;br>
&lt;p>When your CFO asks &amp;ldquo;where does this number come from?&amp;rdquo;, you need an answer faster than &amp;ldquo;let me check and get back to you.&amp;rdquo; Two minutes is generous. In most incident situations, you&amp;rsquo;ve got about thirty seconds before people start making assumptions.&lt;/p>
&lt;h4 id="what-good-looks-like-7">What good looks like&lt;/h4>
&lt;p>The dbt DAG is viewable, hosted on S3 or an internal site. Your team can trace major mart tables back to their sources. During incident response, someone pulls up the lineage graph and walks through the chain. It takes some squinting, but the information is there.&lt;/p>
&lt;h4 id="what-amazing-looks-like-7">What amazing looks like&lt;/h4>
&lt;p>&lt;strong>Column-level lineage&lt;/strong> automated across the full stack — from source systems through dbt transformations to the BI dashboards consumers see.&lt;/p>
&lt;p>dbt Cloud provides column-level lineage showing whether each column is &amp;ldquo;transformed&amp;rdquo; versus &amp;ldquo;passthrough/rename.&amp;rdquo; On the Snowflake side, the &lt;code>ACCESS_HISTORY&lt;/code> view (Enterprise Edition) tracks both read and write operations with column-level mappings. Snowflake Labs published an open-source adapter that converts &lt;code>ACCESS_HISTORY&lt;/code> data into OpenLineage JSON format, making cross-platform lineage possible.&lt;/p>
&lt;p>But here&amp;rsquo;s where it gets practical. The open-source tool &lt;strong>Recce&lt;/strong> compares two dbt environments and produces a &amp;ldquo;Lineage Diff&amp;rdquo; in CI. It categorises changes as breaking, partial-breaking, or non-breaking. PR reviewers get an instant risk assessment: &amp;ldquo;this change affects 47 downstream models, including three Tier 1 dashboards&amp;rdquo; versus &amp;ldquo;this change is isolated to a staging model with no downstream consumers.&amp;rdquo;&lt;/p>
&lt;p>PRs get annotated automatically with impact analysis. Cross-project lineage via dbt Mesh connects multiple dbt projects. Lineage isn&amp;rsquo;t a pretty graph that nobody looks at — it&amp;rsquo;s used daily for onboarding, incident response, impact analysis, and cost optimisation.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="9-do-data-producers-know-when-they-break-downstream-consumers">9. Do data producers know when they break downstream consumers?&lt;/h3>
&lt;br>
&lt;p>This is where data engineering starts borrowing seriously from software engineering. In a microservices world, you wouldn&amp;rsquo;t deploy an API change without knowing who calls your endpoint. But in data, teams routinely change source schemas, modify column semantics, or sunset tables with zero awareness of who&amp;rsquo;s consuming them downstream.&lt;/p>
&lt;p>&lt;strong>Data contracts&lt;/strong> change that dynamic.&lt;/p>
&lt;h4 id="what-good-looks-like-8">What good looks like&lt;/h4>
&lt;p>dbt sources are defined with &lt;code>dbt source freshness&lt;/code> running and alerting. The team is aware of major upstream dependencies. When Fivetran&amp;rsquo;s sync breaks or a source system changes its schema, someone notices within a few hours — usually because a test fails.&lt;/p>
&lt;h4 id="what-amazing-looks-like-8">What amazing looks like&lt;/h4>
&lt;p>dbt &lt;strong>model contracts&lt;/strong> (v1.5+) enforced on all gold-layer and public models:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">models&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">dim_users&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">config&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">contract&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">enforced&lt;/span>: &lt;span style="color:#66d9ef">true&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">columns&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">user_id&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">data_type&lt;/span>: &lt;span style="color:#ae81ff">int&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">constraints&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">type&lt;/span>: &lt;span style="color:#ae81ff">primary_key&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">email&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">data_type&lt;/span>: &lt;span style="color:#ae81ff">varchar&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">constraints&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">type&lt;/span>: &lt;span style="color:#ae81ff">not_null&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>This is a &lt;strong>preflight check&lt;/strong>: dbt verifies the model&amp;rsquo;s compiled SQL returns columns matching the contract before building. Breaking changes — removed columns, changed data types, modified constraints — are caught by &lt;code>state:modified&lt;/code> in CI. Model access levels (&lt;code>public&lt;/code>, &lt;code>protected&lt;/code>, &lt;code>private&lt;/code>) control cross-project visibility.&lt;/p>
&lt;p>Whatnot went further with Protobuf schemas enforced via Buf linter in CI. Every event producer runs against a common testing harness. Post-deployment monitoring catches semantic drift — when a field still exists but its meaning has changed.&lt;/p>
&lt;blockquote>
&lt;p>&lt;strong>Gotcha worth knowing:&lt;/strong> dbt contracts validate the compiled SQL, not the actual Snowflake table. If tools like Fivetran add columns directly to your raw layer, dbt won&amp;rsquo;t know until the next run. Monitor schema changes externally using Snowflake&amp;rsquo;s &lt;code>INFORMATION_SCHEMA&lt;/code> or &lt;code>ACCESS_HISTORY&lt;/code> to catch drift between runs.&lt;/p>&lt;/blockquote>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="10-do-you-do-code-review-on-sql-and-dbt-models">10. Do you do code review on SQL and dbt models?&lt;/h3>
&lt;br>
&lt;p>I&amp;rsquo;m constantly surprised by how many data teams still deploy SQL changes without review. In software engineering, unreviewed code going to production would be considered reckless. In data engineering, it&amp;rsquo;s Tuesday.&lt;/p>
&lt;h4 id="what-good-looks-like-9">What good looks like&lt;/h4>
&lt;p>Pull requests are required for all dbt changes — branch protection is enabled on main. At least one human reviewer looks at each PR. There&amp;rsquo;s a basic PR description explaining what changed and why.&lt;/p>
&lt;h4 id="what-amazing-looks-like-9">What amazing looks like&lt;/h4>
&lt;p>Automated CI running SQLFluff lint + &lt;code>dbt compile&lt;/code> + &lt;code>dbt build&lt;/code> + data diff on every PR. Pre-commit hooks catch issues before code is even committed. PR templates with structured sections: description, linked tickets, impact zone, testing evidence.&lt;/p>
&lt;p>The recommended pre-commit stack combines SQLFluff for linting with dbt-checkpoint for structural validation:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># .pre-commit-config.yaml&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">repos&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">repo&lt;/span>: &lt;span style="color:#ae81ff">https://github.com/sqlfluff/sqlfluff&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">hooks&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">id&lt;/span>: &lt;span style="color:#ae81ff">sqlfluff-lint&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">args&lt;/span>: [--&lt;span style="color:#ae81ff">dialect, snowflake]&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">repo&lt;/span>: &lt;span style="color:#ae81ff">https://github.com/dbt-checkpoint/dbt-checkpoint&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">hooks&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">id&lt;/span>: &lt;span style="color:#ae81ff">check-model-has-tests&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">id&lt;/span>: &lt;span style="color:#ae81ff">check-model-columns-have-desc&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">id&lt;/span>: &lt;span style="color:#ae81ff">check-source-has-freshness&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>That last hook — &lt;code>check-model-columns-have-desc&lt;/code> — is quietly revolutionary. It means documentation isn&amp;rsquo;t optional. You literally cannot merge a model without column descriptions. The &amp;ldquo;we&amp;rsquo;ll document it later&amp;rdquo; excuse dies right there in the CI pipeline.&lt;/p>
&lt;p>Surfline (700+ dbt models) reported that after integrating SQLFluff, SQL consistency improved dramatically and reviewer burden dropped. New engineers learned &amp;ldquo;good SQL&amp;rdquo; from day one because the linter enforced it automatically. Another team, Markerr, found that automated style enforcement freed up review time to focus on deeper logic questions rather than arguing about capitalisation and trailing commas.&lt;/p>
&lt;p>The &lt;code>dbt_project_evaluator&lt;/code> package is worth mentioning too — it audits your entire DAG structure against dbt Labs&amp;rsquo; published best practices. Run it in CI and it&amp;rsquo;ll flag things like models that reference sources directly (bypassing staging), duplicate sources, or models with no tests.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="11-do-new-hires-build-a-real-pipeline-in-their-first-week">11. Do new hires build a real pipeline in their first week?&lt;/h3>
&lt;br>
&lt;p>The speed at which a new hire becomes productive tells you everything about the state of your documentation, tooling, and team culture. If it takes three weeks before someone can make a meaningful contribution, you don&amp;rsquo;t have an onboarding problem — you have a platform problem. I wrote a &lt;a href="https://ghostinthedata.info/posts/2025/2025-02-23-landed-the-role/" target="_blank" rel="noopener">guide to navigating your first 90 days as a data engineer&lt;/a> — and the teams that make those first 90 days count are almost always the ones who&amp;rsquo;ve invested in the infrastructure described here.&lt;/p>
&lt;h4 id="what-good-looks-like-10">What good looks like&lt;/h4>
&lt;p>New hire has access to tools within a day or two. Some documentation exists. A buddy or mentor is assigned. They can run the dbt project locally by end of week one, even if they haven&amp;rsquo;t contributed anything yet.&lt;/p>
&lt;h4 id="what-amazing-looks-like-10">What amazing looks like&lt;/h4>
&lt;p>Pre-arrival setup is complete before Day 1: machine provisioned, logins ready, Snowflake roles assigned, dbt Cloud account active. No engineer should spend their first morning installing things.&lt;/p>
&lt;p>Each developer gets a &lt;strong>personal Snowflake sandbox&lt;/strong> — &lt;code>dbt_&amp;lt;username&amp;gt;&lt;/code> schema in the dev database, with write access only to dev and read-only access to raw/staging. Snowflake&amp;rsquo;s zero-copy cloning makes production data available instantly without duplicating storage costs. Productboard open-sourced &lt;code>dbt-snowflake-sandbox&lt;/code> — a set of dbt macros that create isolated sandboxes by cloning only the specific model dependencies a developer needs.&lt;/p>
&lt;p>The first-week project is structured and progressive:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Day 1-2:&lt;/strong> Run the existing dbt project locally. Explore the DAG. Read the key model documentation.&lt;/li>
&lt;li>&lt;strong>Day 3:&lt;/strong> Add a new source or staging model for a real (but low-risk) dataset.&lt;/li>
&lt;li>&lt;strong>Day 4:&lt;/strong> Build a staging model with tests and documentation.&lt;/li>
&lt;li>&lt;strong>Day 5:&lt;/strong> Open a PR, go through the review process, and get it merged.&lt;/li>
&lt;/ol>
&lt;p>By Friday, they&amp;rsquo;ve shipped something real. They understand the workflow. They&amp;rsquo;ve experienced CI, code review, and deployment. And critically, they &lt;em>feel&lt;/em> like a contributor rather than a tourist.&lt;/p>
&lt;p>7shifts estimated cutting onboarding time by over a week per new hire just by having documentation and data questions accessible in their data catalog.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="12-do-you-regularly-talk-to-the-people-who-actually-use-your-data">12. Do you regularly talk to the people who actually use your data?&lt;/h3>
&lt;br>
&lt;p>This is the question that separates data teams who build for their portfolio from data teams who build for their business. You can nail every technical practice on this list and still fail if you&amp;rsquo;re building the wrong things. I&amp;rsquo;ve written before about &lt;a href="https://ghostinthedata.info/posts/2025/2025-02-15-data-impact/" target="_blank" rel="noopener">how to maximise your data team&amp;rsquo;s impact&lt;/a> — and the through-line is always the same: the teams that create real value are the ones who stay close to the people consuming their work.&lt;/p>
&lt;h4 id="what-good-looks-like-11">What good looks like&lt;/h4>
&lt;p>A Slack channel exists for data questions. Communication happens when issues arise. There are occasional stakeholder meetings, usually prompted by something breaking or a new request coming in.&lt;/p>
&lt;h4 id="what-amazing-looks-like-11">What amazing looks like&lt;/h4>
&lt;p>&lt;strong>Data office hours.&lt;/strong> Weekly or bi-weekly open sessions where anyone in the organisation can bring data questions. Not a presentation. Not a status update. An open door.&lt;/p>
&lt;p>Holistics published a detailed playbook for running what they call &amp;ldquo;data clinics.&amp;rdquo; The core principles: teach, don&amp;rsquo;t serve. Show business users how to self-serve rather than just answering their question and sending them on their way. Montreal Analytics reported that after implementing data clinics, the number of self-serve business users on their BI tool grew tenfold. &lt;em>Tenfold.&lt;/em>&lt;/p>
&lt;p>Pair that with a monthly &lt;strong>Data NPS survey&lt;/strong> — a single question: &amp;ldquo;How likely are you to recommend our data team&amp;rsquo;s products to a colleague?&amp;rdquo; It sounds corporate, but without this metric, you have no way to quantify whether your consumers are satisfied or just silently building workarounds in Excel.&lt;/p>
&lt;p>dbt &lt;strong>exposures&lt;/strong> formalise the connection between data models and their consumers:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">exposures&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">weekly_revenue_dashboard&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">type&lt;/span>: &lt;span style="color:#ae81ff">dashboard&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">owner&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">name&lt;/span>: &lt;span style="color:#ae81ff">Finance Team&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">email&lt;/span>: &lt;span style="color:#ae81ff">finance@company.com&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#f92672">depends_on&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#ae81ff">ref(&amp;#39;fct_revenue&amp;#39;)&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> - &lt;span style="color:#ae81ff">ref(&amp;#39;dim_date&amp;#39;)&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>When a model changes, the PR shows which exposures are affected. Stakeholder impact becomes visible in code review. You can proactively notify the finance team that their revenue dashboard&amp;rsquo;s upstream model is changing before they discover it themselves.&lt;/p>
&lt;p>Frame your communication in business terms. Not &amp;ldquo;model precision is up 12%&amp;rdquo; but &amp;ldquo;the sales team can now identify high-value leads 40% faster.&amp;rdquo; Nobody outside your team cares about your DAG. They care about whether they can trust the numbers. If you&amp;rsquo;re not sure how to bridge that gap, I wrote a guide on &lt;a href="https://ghostinthedata.info/posts/2025/2025-02-08-breaking-down-business-context/" target="_blank" rel="noopener">breaking down business context&lt;/a> — because the hardest part of talking to stakeholders isn&amp;rsquo;t the talking, it&amp;rsquo;s knowing what they actually need to hear.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="scoring-your-team">Scoring your team&lt;/h3>
&lt;br>
&lt;p>Here&amp;rsquo;s the scoring framework, and I want you to be ruthless with yourself:&lt;/p>
&lt;table>
 &lt;thead>
 &lt;tr>
 &lt;th>Score&lt;/th>
 &lt;th>What it means&lt;/th>
 &lt;/tr>
 &lt;/thead>
 &lt;tbody>
 &lt;tr>
 &lt;td>&lt;strong>10-12&lt;/strong>&lt;/td>
 &lt;td>You&amp;rsquo;re in elite territory. Your data platform is a competitive advantage.&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>7-9&lt;/strong>&lt;/td>
 &lt;td>You&amp;rsquo;ve got strong foundations with clear areas to improve.&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>4-6&lt;/strong>&lt;/td>
 &lt;td>You&amp;rsquo;re functional but fragile. One bad incident away from a crisis of trust.&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>1-3&lt;/strong>&lt;/td>
 &lt;td>You&amp;rsquo;re firefighting, not engineering. The business tolerates your team — it doesn&amp;rsquo;t trust it.&lt;/td>
 &lt;/tr>
 &lt;/tbody>
&lt;/table>
&lt;p>Most teams land at 3-4. That&amp;rsquo;s not a criticism — it&amp;rsquo;s where the industry is. The gap between knowing these practices exist and actually implementing them is where the real work lives.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="where-to-start">Where to start&lt;/h3>
&lt;br>
&lt;p>If you&amp;rsquo;re staring at a low score and feeling overwhelmed, here&amp;rsquo;s what I&amp;rsquo;d suggest: &lt;strong>start with practices 4 and 10.&lt;/strong> CI testing and code review.&lt;/p>
&lt;p>Here&amp;rsquo;s why. Once dbt changes go through pull requests with automated checks, you have the infrastructure to enforce everything else. Contracts? They&amp;rsquo;re a CI check. Quality gates? CI check. Documentation requirements? CI check. SLA validation? CI check. You&amp;rsquo;re not adopting twelve practices — you&amp;rsquo;re building one pipeline that gates on twelve things.&lt;/p>
&lt;p>The canonical architecture appears across virtually every mature team I&amp;rsquo;ve studied: GitHub Actions triggering Slim CI with &lt;code>dbt build -s state:modified+&lt;/code> against per-PR Snowflake schemas, with production manifests stored in S3 for state comparison. Start there. Layer on the remaining practices as your team&amp;rsquo;s maturity grows.&lt;/p>
&lt;p>Three patterns emerged from the research that I think are worth calling out explicitly:&lt;/p>
&lt;p>&lt;strong>&amp;ldquo;As code&amp;rdquo; wins everywhere.&lt;/strong> Documentation-as-code, style-guides-as-code, permissions-as-code, quality-checks-as-code, contracts-as-code. Every manual process that gets codified becomes version-controlled, reviewable, and enforceable. Every one that stays manual eventually drifts.&lt;/p>
&lt;p>&lt;strong>Snowflake&amp;rsquo;s zero-copy cloning is the force multiplier.&lt;/strong> Isolated PR environments, sandbox onboarding, rebuild testing — all of these become cheap and instant with cloning. If you&amp;rsquo;re on Snowflake and not using this feature aggressively, you&amp;rsquo;re leaving the most powerful tool in the shed.&lt;/p>
&lt;p>&lt;strong>Culture matters more than tooling.&lt;/strong> Catalog adoption fails without meeting users in their existing workflows. Data quality requires executive sponsorship and error budgets, not just more tests. Producer-consumer contracts are first and foremost a cultural change. You can buy every tool on this list and still score a 3 if the organisation doesn&amp;rsquo;t value the practices behind them.&lt;/p>
&lt;hr>
&lt;br>
&lt;br>
&lt;h3 id="the-diagram-and-the-warehouse">The diagram and the warehouse&lt;/h3>
&lt;br>
&lt;p>Remember that team — the one with the beautiful diagram and the revenue table nobody could rebuild?&lt;/p>
&lt;p>I caught up with one of the engineers about four months after I&amp;rsquo;d moved on. They&amp;rsquo;d started with exactly what I&amp;rsquo;d pushed for: CI testing and code review. Then they added contracts on their gold-layer models. Then SLAs on their Tier 1 tables. Then a first-week onboarding project for new hires.&lt;/p>
&lt;p>Their score when I first asked these twelve questions? Three. Four months later? Eight. Not perfect. But the difference wasn&amp;rsquo;t really the number. The difference was that when they pulled up that architecture diagram now, it matched what was actually running in Snowflake. The gap between the wall and the warehouse had closed.&lt;/p>
&lt;p>That&amp;rsquo;s what this test is really measuring. Not whether you have the right tools — you probably do. Not whether you know the right practices — you&amp;rsquo;re reading this, so clearly you care. It&amp;rsquo;s measuring whether there&amp;rsquo;s a gap between what you think your data platform looks like and what it actually looks like.&lt;/p>
&lt;p>These twelve questions just make you say it out loud.&lt;/p>
&lt;p>So print it. Score yourself honestly. Share it with your team. And the next time someone asks you how your data platform is doing, give them a number between 1 and 12.&lt;/p>
&lt;p>That number is worth more than any architecture diagram.&lt;/p>
&lt;p>&lt;br>&lt;br>&lt;/p></content:encoded><category>Data Engineering</category><category>Data Quality</category><category>Leadership</category><category>Data Engineering</category><category>dbt</category><category>Snowflake</category><category>GitHub Actions</category><category>AWS</category><category>Data Quality</category><category>CI/CD</category><category>Data Contracts</category></item></channel></rss>