惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Jina AI
Jina AI
I
Intezer
F
Fortinet All Blogs
S
SegmentFault 最新的问题
罗磊的独立博客
V
Visual Studio Blog
V
V2EX
大猫的无限游戏
大猫的无限游戏
The Cloudflare Blog
J
Java Code Geeks
美团技术团队
B
Blog
U
Unit 42
F
Full Disclosure
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
P
Privacy International News Feed
G
Google Developers Blog
雷峰网
雷峰网
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
P
Privacy & Cybersecurity Law Blog
T
Tor Project blog
酷 壳 – CoolShell
酷 壳 – CoolShell
量子位
GbyAI
GbyAI
S
Schneier on Security
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Google DeepMind News
Google DeepMind News
D
Darknet – Hacking Tools, Hacker News & Cyber Security
L
LINUX DO - 热门话题
Recorded Future
Recorded Future
D
Docker
博客园 - 聂微东
Project Zero
Project Zero
Know Your Adversary
Know Your Adversary
P
Palo Alto Networks Blog
K
Kaspersky official blog
Martin Fowler
Martin Fowler
H
Hackread – Cybersecurity News, Data Breaches, AI and More
L
Lohrmann on Cybersecurity
A
Arctic Wolf
T
The Blog of Author Tim Ferriss
Microsoft Security Blog
Microsoft Security Blog
T
Threat Research - Cisco Blogs
T
The Exploit Database - CXSecurity.com
V
Vulnerabilities – Threatpost
Simon Willison's Weblog
Simon Willison's Weblog
Cisco Talos Blog
Cisco Talos Blog
T
Threatpost
Hugging Face - Blog
Hugging Face - Blog
博客园_首页

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The Great Data Debate: Should You Build Your Warehouse Top-Down or Bottom-Up?
Lawrence Mur · 2026-05-12 · via DEV Community

Introduction

Imagine you have a massive, disorganized garage. You need to clean it up so you can actually find things. You have two ways to tackle this.
The first way is to take every single item out of the garage, build a perfect, custom-sized shelving unit for the entire space, categorize every loose screw and tool into a master list, and then put everything in its exact, permanent place.
The second way is to just clean out the corner where you keep your gardening tools because it’s spring and that’s what you need right now. Later, when winter comes, you can clean out the corner for your snow shovels.
This is exactly how the data engineering world looks at building a Data Warehouse. The clean the whole garage first method is the Inmon approach. The clean corner by corner method is the Kimball approach.
If your company wants to store data to make smart business decisions, you will inevitably bump into these two names; Bill Inmon and Ralph Kimball.
This article looks at how the architectures work, the good, the bad, and which one you should actually use.

The Inmon Architecture(The Top-Down Master Plan)

Bill Inmon is often called the father of the data warehouse. His philosophy is that a data warehouse should be the single, ultimate source of truth for the entire business.

How It works

Inmon uses a top-down approach. You start by looking at the entire company, pull data from all the different software systems (sales, HR, finance) and clean it up. Then, you store all of it in one massive, highly organized central database.
Because of this design, the Inmon approach requires that business requirements are defined first. You must have a complete understanding of the enterprise's overarching data needs before building the model. Furthermore, it relies on strong governance, meaning there are strict, centralized rules controlling data quality, security, and standardization across the board.
Inmon uses a normalized structure. This means data is stored without any duplication. If a customer's name changes, you only have to update it in one single place.
Building a centralized warehouse first is the core of this method. Once this giant central warehouse is built, you carve out smaller pieces of it, Data Marts, for specific departments to use. Each department gets their own data mart, but that data mart is fed strictly by the central warehouse.
Below is a flowchart showing multiple source systems feeding into a single Staging Area, flows into a large central Enterprise Data Warehouse, which then splits into smaller Data Marts pointing to the end users.
Source → ETL → Data Warehouse → Data Marts → Reports
Inmon approach

Pros

- Single source of truth - Because everything flows from one central hub, the different teams will never have conflicting numbers.
- High consistency - Due to strong governance and a centralized structure, definitions and metrics mean the exact same thing across the entire enterprise.
- Good for large organizations - The robust, highly structured foundation is capable of handling vast amounts of complex, enterprise-wide data efficiently over the long term.
- Easy to update - Since data isn't duplicated, updating records or fixing errors is very clean and simple.
- Built for the future - If the company grows or adds new departments, the foundation is already solid.

Cons

- Slow to implement - Designing a perfect system for an entire enterprise takes months, sometimes years, before anyone sees real value.
- It’s expensive - You need highly specialized database experts and a massive upfront budget to build and maintain the central hub.
- Hard for business users to read - The normalized database is great for computers, but very confusing for a regular business person trying to run a report.
- Hard to change - Because the entire enterprise is highly integrated and normalized, pivoting the architecture to accommodate new, unforeseen business models is difficult and time-consuming.

The Kimball Architecture(The Bottom-Up Quick Win)

Ralph Kimball felt Inmon method was slow and expensive and decided to craft a better method. His philosophy is that a data warehouse focus on business processes and answer specific business questions as quickly as possible.

How it works

Kimball uses a bottom-up approach prioritized around fast delivery. Instead of building a giant central warehouse first, you start by building individual Data Marts.
For example, if the sales team needs a report urgently, you pull data just for the sales team, run it through ETL and build a Sales Data Mart. Then later, you build an HR Data Mart.
Kimball uses a denormalized structure, known as the Star Schema. This means he doesn't care if data is duplicated. He organizes data into Facts (numbers such as sales amount) and Dimensions (context such as time, location, or customer name).
Rather than being isolated silos, these individual Data Marts are eventually linked together to form an Integrated Warehouse. To keep things from getting chaotic, Kimball uses conformed dimensions (an enterprise bus). This is a strict rule that says if both the Sales mart and the HR mart use a Date or a Customer, they must use the exact same definition, allowing the data marts to connect logically for company-wide reporting.
Below is a flowchart flowchart showing source systems feeding into an ETL process, which builds independent Data Marts(Star Schemas) first. These marts are linked together by shared conformed dimensions to form a logical Integrated Warehouse, which is then used for End-User Reports.
Source → ETL → Data Marts → Integrated Warehouse → Reports

The Pros

- Faster implementation - You can get a single department up and running with data in a matter of weeks, delivering immediate ROI.
- Cheaper to start - You don't need a massive upfront budget.
- Business-friendly - The Star Schema is incredibly easy for regular business users to understand. They can drag and drop fields in software like Tableau or PowerBI easily.
- Flexible - It is much easier to add new data marts or modify existing ones as business needs change without breaking a massive central database.

The Cons:
- Data duplication - Because data is stored in multiple different marts, you use up more storage space.
- Harder to update - Because Kimball favors speed and query performance over strict organization, the same piece of data is intentionally stored in multiple places. For example, if a customer's address changes, you might have to update it in five different data marts.
- Risk of inconsistency - If you aren't strictly enforcing conformed dimensions, your data marts will drift apart. Because data is duplicated across different marts, sales and finance might end up reporting different total revenue numbers.
- Integration challenges - Because the system is built piece-by-piece rather than centrally planned from the start, tying all the disparate data marts together into a unified, integrated warehouse later on can become technically complex and messy.
For example, if Sales mart is built in January and the HR mart in July, the teams might design their databases differently. A user trying to generate a combined report showing Sales Revenue vs. Employee Training Costs might realize that Sales measures time in Weeks, while HR measures time in Months. Trying to join the two data marts together to answer enterprise-wide questions thus becomes technologically complex.

Which Architecture is better?

If you ask a room full of data engineers this question, you will probably start an argument. But realistically, neither is better. It entirely depends on what your company needs.

You should use Inmon if

  • You work in a highly regulated industry (like banking, insurance, or healthcare) where data accuracy and audit trails are more important than speed.
  • You have a large budget, a big team of data engineers, and plenty of time.
  • Your company's data is incredibly complex and changes constantly.

You should use Kimball if

  • You are a startup, a retail business, or a fast-moving company that needs data right now.
  • You want your non-technical business teams to build their own reports without asking IT for help every time.
  • You are on a tight budget and need to prove the value of the data warehouse to your boss quickly.

The Modern Reality

It is worth mentioning that technology has changed a lot since Inmon and Kimball wrote their books in the 1990s.
Back then, computer storage was incredibly expensive and Inmon’s method of not duplicating data saved money.
Today, cloud storage is incredibly cheap. Because storage is cheap, many companies lean heavily toward Kimball's Star Schema because the cost of duplicating data just doesn't matter much anymore.
Furthermore, new hybrid approaches have popped up. The Data Vault architecture (by Dan Linstedt) is becoming very popular. It essentially takes the best of Inmon’s strict central storage and pairs it with Kimball’s easy-to-read data marts.

The Bottom Line

When it comes to building a data warehouse, don't get caught up in treating Inmon or Kimball like a religion. You aren't building a monument but a tool to help your company make money.
If your company has the patience to build a bulletproof foundation, go top-down with Inmon. If your company needs answers tomorrow to keep the lights on, go bottom-up with Kimball.
Pick the approach that fits your business reality, not the one that looks prettiest on a whiteboard.