Why Code-Based Data Pipelines are Coming Back
CIOREVIEW >> Business Intelligence >> NEWS

Blick Art Materials

Farbod Sedaei, Director of Business Intelligence and Pricing Intelligence

Why Code-Based Data Pipelines are Coming Back

Farbod Sedaei, Director of Business Intelligence and Pricing Intelligence
Farbod Sedaei, Director of Business Intelligence and Pricing Intelligence, Blick Art Materials

Farbod Sedaei

Data Modernization Authority

When Code Becomes the Faster Choice

We had a set of pipelines first built more than 15 years ago, modified continuously ever since. New source systems, changed schemas, revised business logic, each change layered onto the last by whoever was available that quarter. The result was a process that worked, that nobody fully understood, and that no one wanted to touch. It ran for about thirty minutes.

We put AI to work reading it. The output was a dbt implementation that runs in under five minutes, with documentation explaining what each transformation actually does. The logic had not changed. What changed was that 15 years of accumulated intent was finally written down, and inefficiencies hiding inside that complexity became visible once something could read the whole thing at once.

That project is why I think the modernization conversation has genuinely changed, and not for the reasons usually given.

The demand was structural. Application portfolios fragmented. A function that once lived in a single ERP now spans a SaaS CRM, a SaaS HR platform, a billing service, and two internal systems. Every boundary is a place where data has to move, be reconciled, and be reshaped. Teams responded the only way they could, which was by building faster, and low-code platforms were a rational answer to that pressure. 

Choosing a low-code platform over pipeline code is a specific trade, usually made without stating it out loud. What you get is speed. A visual tool ships integration in an afternoon, with no repository, testing framework or deployment process to set up. A junior analyst can build something usable, which matters enormously when the queue is long and the team is small.

What you give up is everything that makes software maintainable. Logic on a canvas is hard to review, because there is no diff to read. Hard to test, because the platform's testing story is thin or absent. Hard to search, so when a business rule changes, finding every instance means clicking through pipelines one at a time.

That trade was worth making for years, because writing code was genuinely slower. It meant hand-writing transformations, mapping fields, building incremental load logic, handling retries and failures. Unglamorous, repetitive work and the reason an afternoon in a visual tool became a sprint in code.

Once the speed advantage is gone, what code was always better at is all that remains to compare. It lives in version control, so every change has an author and a reason. It gets reviewed before it merges. It runs through CI, so tests execute on every change.

Our 15-year-old process was hard to understand, not because the logic was sophisticated, but because it had no single readable representation. Once it existed as dbt models in a repository, problems invisible for years were simply there on the page.

AI is also Raising the Value of Data You Used to Discard 

A second effect runs in the same direction. For as long as analytics has existed, humans have been the bottleneck on how many variables a question can hold. We pick the dozen fields we believe matter and drop the rest, because a person can only hold so many relationships in mind at once.

Models don't have that ceiling. Event timestamps, device metadata, support-ticket text and operational logs: all of it can carry a signal when something looks across the whole set at once. Fields reasonably classified as noise a few years ago now surface relationships nobody thought to look for.

  When every pipeline is in use, the question isn't what to retire but what is actually costing you.  

If the useful surface area of your data is expanding, every AI initiative is a data integration initiative wearing a different hat. The model is rarely the hard part. Getting clean, current data to it is, and that is pipeline work.

The pipelines built a decade ago are the ones expected to feed AI workloads, and were never designed for it. What changed is that AI is good at refactoring, not only greenfield work. Logic trapped in a legacy tool can be read, explained, and translated far faster than it can be rewritten. More significantly, the parts of modernization that always got cut for time are now cheap enough to keep: test cases proving the new pipeline matches the old, automated testing in CI, built-in validation for row counts and freshness, and documentation of how each number is derived.

The Obvious Caveat

AI-generated code is not always correct. Anyone who has used it seriously has examples. Translated logic can be subtly wrong in ways that pass a smoke test, and a pipeline producing plausible numbers is more dangerous than one that fails loudly.

I still consider it a net gain because the same tooling generates the tests, validations, and documentation that catch those errors. You get the review apparatus and the code together. Parity testing isn't optional, and a senior engineer still owns what ships. But the labor of building that safety net, which is what made thorough refactoring unaffordable, has largely collapsed.

Where I Would Start 

Not with a platform decision, and not with a plan to modernize everything. Seven hundred aging pipelines is a multi-year program, and treating it as one guarantees it never begins.

Start by sorting them. When every pipeline is in use, the question isn't what to retire but what is actually costing you: the slow ones, the fragile ones, the ones only one person understands, the ones feeding analytics you expect to grow. Those lists overlap, and the overlap is your starting set.

Then modernize one, end-to-end, with tests and documentation. The point of the first one is not savings. It's a proven pattern and a defensible estimate for the next hundred. Modernization has been on everyone's roadmap for years and has kept losing to operational  load. The difference now is that the effort required has dropped while the cost of deferring has  risen. Both moved in the same year, which is rare.

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.