惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Recent Announcements
Recent Announcements
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Application and Cybersecurity Blog
Application and Cybersecurity Blog
N
News | PayPal Newsroom
P
Proofpoint News Feed
L
Lohrmann on Cybersecurity
S
Security @ Cisco Blogs
K
Kaspersky official blog
A
Arctic Wolf
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Project Zero
Project Zero
L
LINUX DO - 最新话题
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
The Last Watchdog
The Last Watchdog
T
The Exploit Database - CXSecurity.com
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Security Archives - TechRepublic
Security Archives - TechRepublic
V
V2EX
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
H
Hackread – Cybersecurity News, Data Breaches, AI and More
爱范儿
爱范儿
F
Full Disclosure
I
Intezer
Schneier on Security
Schneier on Security
AWS News Blog
AWS News Blog
C
Cybersecurity and Infrastructure Security Agency CISA
博客园 - 聂微东
M
MIT News - Artificial intelligence
P
Privacy & Cybersecurity Law Blog
Attack and Defense Labs
Attack and Defense Labs
量子位
Google DeepMind News
Google DeepMind News
T
Threat Research - Cisco Blogs
Last Week in AI
Last Week in AI
Google Online Security Blog
Google Online Security Blog
博客园 - 三生石上(FineUI控件)
WordPress大学
WordPress大学
Microsoft Security Blog
Microsoft Security Blog
Scott Helme
Scott Helme
C
Check Point Blog
N
Netflix TechBlog - Medium
博客园 - Franky
SecWiki News
SecWiki News
Know Your Adversary
Know Your Adversary
Engineering at Meta
Engineering at Meta
F
Fortinet All Blogs
Blog — PlanetScale
Blog — PlanetScale
S
Securelist

Databricks

Why Talent Transformation Is the Missing Focus of Enterprise AI Public Health Intelligence Shouldn't Require a Data Scientist Mean Time to Detect Is a Data Access Problem First-party audience data is the ad sales relationship now Rethinking Distributed Systems for Serverless Performance and Reliability The AI Scaling Gap Hiding in Digital Native Companies 10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks AI success starts with clean data, not just better models How nOps Rebuilt Their Cloud Optimization Platform on Databricks Lakebase, and Why Other ISVs Should Too Peril Predicts: Precision Payouts for a Volatile World The foundation of AI scalability: one team, one platform, one operating model The Federal Data Paradox: Rich in Data, Poor in Access Driving Budapest Forward: How BKK Uses Databricks to Transform City Mobility LLM Vs AI: A Practical Guide to Differences, Use Cases, and Tools Model Risk Governance Is Not the Same as Risk Intelligence Generative AI for Business: A Complete Strategy and Implementation Guide Data Science vs Data Engineering: Choosing Analysis or Infrastructure AI Applications: Tools, Use Cases, and Platforms MLOps vs DevOps: A Practical Guide for Data Scientists and IT Teams Top Data Warehouse Tools For Modern Data Analytics Unlocking SAP Business Context in Databricks with Semantic Metadata Delta Sharing The marketing activation gap has a fix: Databricks and Stitch partner to turn data infrastructure into marketing performance Alert Fatigue Is a Business Risk Backstage with Lakebase Shipping Faster isn’t Learning Faster Why Your OEE Dashboard Is Lying to You The Turbine That Tried to Tell You It Was Failing Predicting Readmissions Isn't Enough. Acting in Time Is. Clinical Trials Run Longer Than They Have To. That's a Patient Problem Network Quality Is a Revenue Problem, Not a Technical One Shelf Availability Starts with Better Demand Visibility When Predicting the Next Hit Requires More Than Intuition Approximate Answers, Exact Decisions: New Sketch Functions for Analytics Companies Winning with AI Built the Data Layer First Rethinking SQL ETL for modern data platforms Stripe data now available on Databricks via Databricks Marketplace Databricks and Stripe Projects: Infrastructure Built for Agents Agents are ready but your architecture probably isn't Interoperability Between Unity Catalog and Google BigQuery via Catalog Federation Built In, Not Bolted On: What AI-Native Actually Means in Cybersecurity Operationalizing AI for public sector fraud prevention From months to minutes: Building real-time clinical data pipelines with natural language Agentic Data Engineering with Genie Code and Lakeflow Securely send first-party conversion signals with Snapchat Conversions API on Databricks Marketplace How leading tech companies are killing the builder’s tax with Lakebase Inside one of the first production deployments of Lakebase: LangGuard's agentic workflow governance engine The next generation of Databricks Genie Model Risk Management in 2026: A Banker’s Guide to the Revised Interagency Guidance OpenAI GPT-5.5 now available on Databricks, fully-governed through Unity AI Gateway Operational databases: How they work and when to use them Databricks partners with OpenAI on GPT-5.5 Announcing the Public Preview of Lakeflow Designer Are LLM agents good at join order optimization? How conversational analytics removes the BI bottleneck How to transform document activation workflows with Genie and Agent Bricks Beyond the spreadsheet: how Databricks is delivering the modern CFO in Financial Services AI App Development: Guide To Building AI-Powered Apps IoT in Manufacturing: Strategy, Components, Use Cases, and Challenges Stop Hand-Coding Change Data Capture Pipelines Multimodal Data Integration: Production Architectures for Healthcare AI Personalization Strategies for Media Companies A Modern AI Risk Management Framework Introducing the Databricks Excel Add-in for Business Users Real-Time Decisioning for AI Agents: Why you Need a Customer Context Layer First A Practical Guide to LLM Fine Tuning AI Data Transformation Guide for Data Engineers and Data Scientists Concurrency Control in DBMS: How Locking, MVCC and Optimistic Strategies Keep Data Consistent Bridging data science and marketing: Databricks unveils Delta Sharing integration for Adobe Experience Platform and agentic marketing workflows Take Control: Customer-Managed Keys for Lakebase Postgres Get hands on with agents, vibe coding and more at Data+ AI Summit Mercedes-Benz Builds a Cross-Cloud Data Mesh with Delta Sharing and Intelligent Replication, Cutting Costs by 66% What Is a Transactional Database? Introducing Genie Agent Mode Governing coding agent sprawl with Unity AI Gateway Governing Coding Agent Sprawl with Unity AI Gateway What is pgvector? Banks Don’t Have an AI Problem – They Have a Data Platform Problem Open Platform, Unified Pipelines: Why dbt on Databricks is Accelerating Why Your Agents Can’t Read Enterprise Documents — and How to Fix It Building with Databricks Document Intelligence and Lakeflow Databricks on Google Cloud: Innovate Faster. Smarter. Together. Introducing the Databricks Connector for Google Sheets: Real-Time, Governed Lakehouse Data in the Sheets Users Love Unity AI Gateway: How to connect agents to external MCPs securely Expanding agent governance with Unity AI Gateway Agentic reasoning in practice: Making sense of structured and unstructured data Agent Bricks: The Governed Enterprise Agent Platform 8 AI and data trends shaping financial services in 2026 Building real-time product search on Databricks Lovable + Databricks: Build Data-Driven Apps at the Speed of Thought Memory scaling for AI agents Powering clinical research innovation: How TriNetX uses Databricks to accelerate drug development Database Branching in Postgres: Git-Style Workflows with Databricks Lakebase How Zalando built a unified data foundation for AI and analytics on Databricks The next era of the open lakehouse: Apache Iceberg™ v3 in Public Preview on Databricks How FSIs eliminate silos between clients, operations, and finance How MakeMyTrip achieved millisecond personalization at scale with Databricks A multi-agent approach to audience intelligence AiChemy: Next-generation agent with MCP, skills and custom data for drug discovery Accelerate business insights with Lakeflow Connect, now with a Free Tier Unlocking Next-Gen Customer Experiences with Data Intelligence for Marketing
What is data pipeline architecture?
Databricks Staff · 2026-06-17 · via Databricks

Data pipeline architecture is the end-to-end design of how data is collected, processed, stored and delivered from source systems to the people, applications and models that use it. The word “architecture” refers to the blueprint, not the pipeline itself. It covers the choices about how data flows, where it gets transformed and which tools handle each step along the way.

Good architecture is matched to the use case rather than picked off a shelf. A data pipeline built for real-time fraud detection looks very different from one that produces a nightly sales report, even though both move data from source to destination. This glossary page covers the core layers every pipeline shares, the common stage models, the major architectural patterns and the best practices that keep pipelines reliable as they scale.

How does data pipeline architecture work?

A data pipeline moves data through a series of stages, and each stage has a specific job: gather the data, clean it up, store it and make it usable. Architecture is the plan for how those stages connect. It defines what happens to the data at each step, in what order and under what rules.

Architecture decisions sit at two levels. The logical design defines which stages exist and what each one does: this is “the what.” The physical design defines which specific tools and infrastructure run each stage: this is “the how.” Orchestration (the automatic scheduling and coordination of each step) and monitoring don’t belong to any single stage. They run across the whole pipeline. Modern platforms have also collapsed an old divide. With Lakeflow, Databricks unifies batch and streaming pipelines on a single foundation, so teams don’t have to build and maintain two parallel systems.

The core layers of a data pipeline

Regardless of the pattern a team chooses, every data pipeline is built on the same four layers. Each layer answers a different question about the data: how it gets in, how it becomes useful, where it lives and who consumes it.

Ingestion

Ingestion pulls data into the pipeline from source systems: databases, applications, APIs, files in cloud storage, event streams and sensors. Data ingestion comes in two flavors. Batch ingestion pulls data on a schedule, such as every hour or every night. Streaming ingestion captures data continuously as events happen. Many pipelines also use change data capture (CDC), a method that tracks row-level changes in a source database so the pipeline moves only what’s new or updated instead of reloading everything.

Processing and transformation

This layer is where raw data gets cleaned, reshaped, enriched and prepared for use. Typical work includes fixing missing values, standardizing formats, joining datasets and applying business logic, the same tasks at the heart of ETL. Processing follows the same split as ingestion. Batch processing works on large chunks of data together, while stream processing handles records one at a time or in tiny micro-batches as they arrive.

Storage

Storage is where processed data lands so it can be queried, analyzed or fed to models. The destination is typically a data lake, a data warehouse or a lakehouse, a single system that combines the strengths of both. Format matters as much as location. Open formats like Lakehouse Storage and Apache Iceberg let multiple tools read the same data without copying it from system to system. Delta Lake also adds reliability features such as ACID transactions (a guarantee that writes either fully succeed or fully fail, preventing corruption) and time travel (the ability to query older versions of a table).

Serving and consumption

The final layer delivers prepared data to the people and systems that need it: analysts running SQL queries, business users working in dashboards, data scientists training models and applications calling APIs. Destinations range from BI tools to ML platforms to operational systems, with a data warehouse often sitting at the center of analytics workloads. Across all four layers, orchestration and observability do the connective work: scheduling jobs, tracking data quality and raising alerts when something breaks.

How many stages are in a data pipeline? (3 vs. 4 vs. 5)

Different sources describe data pipelines as having three, four or five stages, which causes plenty of confusion. The reality is simpler. All three models describe the same underlying work at different levels of detail.

ModelStagesWhen you'll see it used
3-stageSources → Processing → DestinationHigh-level explanations, executive overviews, intro-level content
4-stageIngestion → Processing → Storage → ServingMost common in modern data engineering. Balances clarity and detail
5-stageCollection → Ingestion → Processing → Storage → AnalysisDetailed technical breakdowns. Splits “getting data” into collection (from the source) and ingestion (into the pipeline)

The number of stages is a labeling choice. The work the pipeline performs is the same.

Common data pipeline architecture patterns

Architectural patterns are the established designs teams choose from when building pipelines. The right one depends on latency requirements, data volume and how the data will be used downstream.

Batch architecture

Batch architecture processes data in scheduled chunks: every hour, every night or every week. It fits reporting, historical analysis, ML training data and any use case where minutes or hours of delay are acceptable. Batch pipelines are simpler to build, cheaper to run and easier to debug than their streaming counterparts. The trade-off is freshness. When decisions depend on what happened seconds ago, batch can’t keep up.

Streaming architecture

Streaming architecture processes data continuously, record by record, as it’s generated. It serves use cases where sub-minute response matters: fraud detection, real-time personalization and IoT monitoring. The trade-off is cost. Streaming pipelines typically cost more to run and operate than batch pipelines because they require always-on infrastructure.

Lambda architecture

Lambda architecture runs two parallel paths. A batch path delivers accurate historical data, a streaming path delivers fast, fresh data and a serving layer merges the results. The design works, but it carries a well-known downside. Maintaining two pipelines means duplicate code, duplicate logic and double the operational burden.

Kappa architecture

Kappa architecture simplifies Lambda by using a single streaming pipeline for everything. When historical analysis is needed, the stream is replayed from the beginning. Kappa suits teams that want streaming-grade freshness without the cost of maintaining two parallel systems.

Medallion architecture (lakehouse pattern)

Medallion architecture is a popular pattern on lakehouse platforms that organizes data into three quality tiers: Bronze (raw, as ingested), Silver (cleaned and conformed) and Gold (curated, business-ready). As Databricks documentation puts it, “the medallion architecture uses three layers: bronze, silver, and gold, each serving a distinct purpose in the pipeline.” Each tier can run as its own pipeline, which makes scheduling, monitoring and troubleshooting easier because problems stay isolated to a single layer.

ETL vs. ELT: how transformation order shapes architecture

ETL and ELT differ in when data gets transformed. ETL (extract, transform, load) transforms data before loading it into storage. ELT (extract, load, transform) loads raw data first and transforms it inside the destination. Modern cloud platforms such as Databricks, Snowflake and BigQuery have made ELT the dominant pattern because cloud storage and compute are now cheap and elastic enough to transform data in place. For a deeper comparison, see ETL vs. ELT.

 ETLELT
OrderExtract → Transform → LoadExtract → Load → Transform
Where transformation happensIn a separate processing tool, before storageInside the destination (lakehouse or warehouse)
Typical use caseLegacy on-prem warehouses, strict pre-load validationModern cloud lakehouses and warehouses
StrengthsCleaner data lands in storage. Predictable schemasFlexible, scalable, keeps raw data available for reprocessing
Trade-offsLess flexible. Harder to reuse raw data laterRequires capable compute at the destination

Is ETL the same as a data pipeline?

No. ETL is one type of data pipeline, but not every data pipeline is ETL. A data pipeline is the broad category: any system that moves data from one place to another. ETL is a specific approach within that category, defined by transforming data before it lands in storage. Pipelines can also be ELT, streaming, replication-only (moving data with no transformation at all) or reverse ETL (sending warehouse data back into operational systems).

Best practices for data pipeline architecture

These 10 design principles separate pipelines that scale from pipelines that break.

  1. Separate ingestion from transformation. Keep raw data landing and data cleaning in different stages so issues in one don’t cascade into the other.
  2. Design for idempotency. A pipeline should be safe to re-run without creating duplicate records or corrupting results. This is critical for handling failures and backfills.
  3. Build in data quality checks. Strong data quality checks validate schema, value ranges, null counts and freshness at each stage, and they fail loudly when something is wrong rather than letting bad data flow downstream.
  4. Plan for schema drift. Source systems change. Pipelines should detect when columns are added, removed or renamed and handle the change gracefully instead of breaking.
  5. Use open storage formats. Formats like Delta Lake and Apache Iceberg prevent lock-in and let multiple tools read the same data without copies.
  6. Decouple pipeline layers. Splitting medallion tiers (Bronze, Silver and Gold) into separate pipelines makes each one easier to schedule, monitor and troubleshoot independently.
  7. Version control everything. Store pipeline code and configuration in Git so changes are reviewed, traceable and reversible.
  8. Treat governance as a first-class concern. Apply consistent permissions, lineage tracking and audit controls across every stage with a tool like Unity Catalog, rather than bolting them on at the end.
  9. Right-size streaming vs. batch. Use streaming only where freshness genuinely matters, and default to batch everywhere else to control cost.
  10. Monitor end to end. Track data freshness, volume, quality and pipeline run times so problems are caught before downstream users notice them.

Why data pipeline architecture matters

Pipeline architecture determines whether teams can trust their data, whether decisions rest on fresh information and whether AI and ML projects make it from prototype to production. It’s the difference between a data platform that compounds in value and one that generates support tickets.

Brittle architecture creates real costs: stale dashboards, conflicting metrics, failed ML deployments and engineers who spend more time firefighting than building. The modern lakehouse approach addresses the root cause. By unifying batch and streaming, analytics and AI, and governance on a single platform like the Databricks Platform, teams remove the fragile handoffs between systems that make traditional architectures break.

Data pipeline architecture on Databricks

Databricks delivers every layer of pipeline architecture in one platform. Lakeflow Connect handles ingestion from databases, SaaS applications, file sources and event streams. Lakeflow Spark Declarative Pipelines builds batch and streaming ETL pipelines with data quality checks built in, and Lakeflow Jobs orchestrates and schedules pipeline runs across the platform. Underneath, Delta Lake provides the open storage format along with reliability features like ACID transactions and time travel, while Unity Catalog applies governance, lineage and access control across every stage.

Because batch and streaming pipelines run on the same engine and write to the same storage, teams don’t need to maintain Lambda-style parallel systems. One pipeline definition can serve both the nightly report and the real-time dashboard.

Frequently asked questions

What is data pipeline architecture in simple terms?

It’s the plan for how data gets from where it’s created to where it’s useful. The plan covers how data is collected, how it’s cleaned and prepared, where it’s stored and how it’s delivered to the people and applications that need it.

What is the difference between Lambda and Kappa architecture?

Lambda runs two parallel pipelines, one batch and one streaming, and merges their results in a serving layer. Kappa uses a single streaming pipeline for everything and replays the stream when historical analysis is needed. Kappa is simpler to operate, while Lambda persists in environments where batch and streaming paths evolved separately.

When should you use batch vs. streaming pipelines?

Use streaming when the value of data drops within seconds or minutes, as in fraud detection, live personalization or equipment monitoring. Use batch for everything else, including reporting, historical analysis and ML training data. Batch is simpler and cheaper, so it’s the sensible default until a use case proves it needs real-time data.

What’s the difference between logical and physical pipeline architecture?

Logical architecture defines the stages of a pipeline and what each one does, independent of any tool. Physical architecture maps those stages onto specific technologies and infrastructure. Teams usually settle the logical design first, then choose the platforms that implement it.

Match your architecture to the job

Data pipeline architecture is the design behind how data moves and becomes useful. The right architecture is the one that balances freshness, cost and reliability for the specific job at hand, whether that’s a nightly sales report or a fraud check that runs in milliseconds.

See how Databricks unifies batch and streaming pipelines, storage and governance on one platform.