惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
Security Archives - TechRepublic
Security Archives - TechRepublic
V
Vulnerabilities – Threatpost
C
CXSECURITY Database RSS Feed - CXSecurity.com
P
Proofpoint News Feed
G
GRAHAM CLULEY
P
Privacy International News Feed
The Hacker News
The Hacker News
Forbes - Security
Forbes - Security
U
Unit 42
N
News and Events Feed by Topic
D
Darknet – Hacking Tools, Hacker News & Cyber Security
C
Cyber Attacks, Cyber Crime and Cyber Security
C
Cisco Blogs
A
About on SuperTechFans
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
D
Docker
I
Intezer
Spread Privacy
Spread Privacy
The Last Watchdog
The Last Watchdog
V2EX - 技术
V2EX - 技术
S
Security @ Cisco Blogs
F
Full Disclosure
S
Secure Thoughts
M
MIT News - Artificial intelligence
Microsoft Security Blog
Microsoft Security Blog
G
Google Developers Blog
aimingoo的专栏
aimingoo的专栏
W
WeLiveSecurity
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Project Zero
Project Zero
Recorded Future
Recorded Future
Cyberwarzone
Cyberwarzone
S
Security Affairs
AWS News Blog
AWS News Blog
H
Help Net Security
The GitHub Blog
The GitHub Blog
Hacker News: Ask HN
Hacker News: Ask HN
Vercel News
Vercel News
P
Proofpoint News Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
The Register - Security
The Register - Security
S
Schneier on Security
F
Fortinet All Blogs
C
CERT Recently Published Vulnerability Notes
L
LINUX DO - 最新话题
T
Tor Project blog
T
The Exploit Database - CXSecurity.com
MongoDB | Blog
MongoDB | Blog
Webroot Blog
Webroot Blog

Databricks

How lakebase architecture delivers 5x faster Postgres writes Why Talent Transformation Is the Missing Focus of Enterprise AI Public Health Intelligence Shouldn't Require a Data Scientist Mean Time to Detect Is a Data Access Problem First-party audience data is the ad sales relationship now Rethinking Distributed Systems for Serverless Performance and Reliability The AI Scaling Gap Hiding in Digital Native Companies 10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks AI success starts with clean data, not just better models How nOps Rebuilt Their Cloud Optimization Platform on Databricks Lakebase, and Why Other ISVs Should Too Peril Predicts: Precision Payouts for a Volatile World The foundation of AI scalability: one team, one platform, one operating model The Federal Data Paradox: Rich in Data, Poor in Access Driving Budapest Forward: How BKK Uses Databricks to Transform City Mobility LLM Vs AI: A Practical Guide to Differences, Use Cases, and Tools Model Risk Governance Is Not the Same as Risk Intelligence Generative AI for Business: A Complete Strategy and Implementation Guide Data Science vs Data Engineering: Choosing Analysis or Infrastructure AI Applications: Tools, Use Cases, and Platforms MLOps vs DevOps: A Practical Guide for Data Scientists and IT Teams Top Data Warehouse Tools For Modern Data Analytics Unlocking SAP Business Context in Databricks with Semantic Metadata Delta Sharing The marketing activation gap has a fix: Databricks and Stitch partner to turn data infrastructure into marketing performance Alert Fatigue Is a Business Risk Backstage with Lakebase Shipping Faster isn’t Learning Faster Why Your OEE Dashboard Is Lying to You The Turbine That Tried to Tell You It Was Failing Predicting Readmissions Isn't Enough. Acting in Time Is. Clinical Trials Run Longer Than They Have To. That's a Patient Problem Network Quality Is a Revenue Problem, Not a Technical One Shelf Availability Starts with Better Demand Visibility When Predicting the Next Hit Requires More Than Intuition Approximate Answers, Exact Decisions: New Sketch Functions for Analytics Companies Winning with AI Built the Data Layer First Rethinking SQL ETL for modern data platforms Stripe data now available on Databricks via Databricks Marketplace Databricks and Stripe Projects: Infrastructure Built for Agents Agents are ready but your architecture probably isn't Interoperability Between Unity Catalog and Google BigQuery via Catalog Federation Built In, Not Bolted On: What AI-Native Actually Means in Cybersecurity Operationalizing AI for public sector fraud prevention From months to minutes: Building real-time clinical data pipelines with natural language Agentic Data Engineering with Genie Code and Lakeflow Securely send first-party conversion signals with Snapchat Conversions API on Databricks Marketplace How leading tech companies are killing the builder’s tax with Lakebase Inside one of the first production deployments of Lakebase: LangGuard's agentic workflow governance engine The next generation of Databricks Genie Model Risk Management in 2026: A Banker’s Guide to the Revised Interagency Guidance OpenAI GPT-5.5 now available on Databricks, fully-governed through Unity AI Gateway Operational databases: How they work and when to use them Databricks partners with OpenAI on GPT-5.5 Announcing the Public Preview of Lakeflow Designer Are LLM agents good at join order optimization? How conversational analytics removes the BI bottleneck How to transform document activation workflows with Genie and Agent Bricks Beyond the spreadsheet: how Databricks is delivering the modern CFO in Financial Services AI App Development: Guide To Building AI-Powered Apps IoT in Manufacturing: Strategy, Components, Use Cases, and Challenges Stop Hand-Coding Change Data Capture Pipelines Multimodal Data Integration: Production Architectures for Healthcare AI Personalization Strategies for Media Companies A Modern AI Risk Management Framework Introducing the Databricks Excel Add-in for Business Users Real-Time Decisioning for AI Agents: Why you Need a Customer Context Layer First A Practical Guide to LLM Fine Tuning AI Data Transformation Guide for Data Engineers and Data Scientists Concurrency Control in DBMS: How Locking, MVCC and Optimistic Strategies Keep Data Consistent Bridging data science and marketing: Databricks unveils Delta Sharing integration for Adobe Experience Platform and agentic marketing workflows Take Control: Customer-Managed Keys for Lakebase Postgres Get hands on with agents, vibe coding and more at Data+ AI Summit Mercedes-Benz Builds a Cross-Cloud Data Mesh with Delta Sharing and Intelligent Replication, Cutting Costs by 66% What Is a Transactional Database? Introducing Genie Agent Mode Governing coding agent sprawl with Unity AI Gateway Governing Coding Agent Sprawl with Unity AI Gateway What is pgvector? Banks Don’t Have an AI Problem – They Have a Data Platform Problem Open Platform, Unified Pipelines: Why dbt on Databricks is Accelerating Why Your Agents Can’t Read Enterprise Documents — and How to Fix It Building with Databricks Document Intelligence and Lakeflow Databricks on Google Cloud: Innovate Faster. Smarter. Together. Introducing the Databricks Connector for Google Sheets: Real-Time, Governed Lakehouse Data in the Sheets Users Love Unity AI Gateway: How to connect agents to external MCPs securely Expanding agent governance with Unity AI Gateway Agentic reasoning in practice: Making sense of structured and unstructured data Agent Bricks: The Governed Enterprise Agent Platform 8 AI and data trends shaping financial services in 2026 Building real-time product search on Databricks Lovable + Databricks: Build Data-Driven Apps at the Speed of Thought Memory scaling for AI agents Powering clinical research innovation: How TriNetX uses Databricks to accelerate drug development How Zalando built a unified data foundation for AI and analytics on Databricks The next era of the open lakehouse: Apache Iceberg™ v3 in Public Preview on Databricks How FSIs eliminate silos between clients, operations, and finance How MakeMyTrip achieved millisecond personalization at scale with Databricks A multi-agent approach to audience intelligence AiChemy: Next-generation agent with MCP, skills and custom data for drug discovery Accelerate business insights with Lakeflow Connect, now with a Free Tier Unlocking Next-Gen Customer Experiences with Data Intelligence for Marketing
Database Branching in Postgres: Git-Style Workflows with Databricks Lakebase
Susan Pierce · 2026-04-10 · via Databricks

The database is the last bottleneck in your dev workflow

Database branching is the missing primitive in modern development workflows. Every other part of the stack has evolved to support fast iteration. Code has Git. Infrastructure has Terraform. Deploys have CI/CD pipelines that run in minutes. But relational databases still work the way they did ten years ago.

Most teams share a single staging database. Within days of being set up, that database drifts out of sync with production. Schemas diverge as developers apply migrations in different orders. Sequence values no longer match. Test data accumulates and pollutes results. Someone eventually reseeds the whole thing, and the cycle starts over.

Setting up a new environment is worse. The standard approach is to run pg_dump against production, wait for it to finish (minutes to hours depending on database size), load it into a new instance, configure access, and hope the result actually reflects what is running in production. For a 500GB database, this means a 500GB copy operation, plus the compute and storage to keep it running.

The result is predictable. Teams avoid creating new environments because they are too expensive and too slow. Developers share a single mutable staging database. Migrations get tested against stale data, or not tested at all. Preview deployments run against empty fixtures instead of realistic schemas. CI tests share state and produce flaky results.

The database becomes the part of the stack that developers are afraid to touch.

Databricks Lakebase Postgres changes this with database branching.

What database branching actually is

A database branch is not a database copy. This distinction matters because it changes the economics of isolated environments entirely.

When you copy a database, you duplicate all of its data and schema into a new, independent instance. The time and cost scale linearly with the size of the database. Every copy is a full clone, and every clone starts going stale the moment it is created.

A branch works differently. When you create a branch in Lakebase, you get a new, fully isolated Postgres environment that:

  • Starts from the exact schema and data of its parent at a specific point in time
  • Shares the same underlying storage instead of duplicating it
  • Only writes new data when you actually make changes

This is called copy-on-write. As long as two branches have not diverged, they reference the same stored data. When you run a migration, insert rows, or modify tables on a branch, only those changes are written separately. Everything else is shared with the parent.

Database copy vs. database branch

 

Database copy (pg_dump, RDS snapshot)

Database branch (Lakebase)

Time to create

Minutes to hours, scales with database size

Seconds, constant regardless of database size

Storage cost

Full duplicate of all data

Proportional to changes only (copy-on-write)

Isolation

Full, but expensive to maintain

Full, with independent compute and connection strings

Freshness

Stale from the moment it is created

Starts from the exact state of the parent at branch time

Cleanup

Manual teardown of instances and storage

Delete the branch; compute and storage are reclaimed automatically

In practical terms, this means:

  • Branch creation takes seconds, regardless of database size. A 10GB database and a 2TB database branch in the same amount of time.
  • Storage cost is proportional to changes, not total data size. A branch that modifies 50MB of data in a 500GB database uses roughly 50MB of additional storage.
  • Each branch gets its own Postgres connection string and compute endpoint. Branches are fully isolated from each other and from their parent.
  • Idle branches automatically scale compute to zero. You only pay for active compute when a branch is actually being used.

Branches are designed to be created, used, and discarded freely. By developers, by CI pipelines, by AI agents, by automation. They are not precious environments that need to be maintained. They are disposable, like Git branches.

The architecture that makes database branching possible

Traditional managed Postgres (RDS, Azure Database for PostgreSQL) ties compute and storage together. The database process and its data live on the same instance, and the data is stored as a single mutable filesystem. That is why copying is the only option for creating a second environment: you have to duplicate the filesystem.

But a lakebase is built different. It separates compute from storage completely. All data is written to a distributed, versioned storage engine that records every change as a new version rather than overwriting existing data. This log-structured architecture is what makes database branching possible as a primitive rather than as a feature layered on top.

Because storage is versioned, multiple branches can reference the same underlying data without risk of conflict. Because compute is independent, each branch runs its own Postgres process and scales on its own. Non-production branches that sit idle scale down to zero automatically and restart in milliseconds when a connection comes in.

Not all database branching implementations are equal. Some platforms create full instance copies and call them branches. Others branch only the schema, without data. Lakebase branches include both schema and data, use copy-on-write at the storage layer to avoid duplication, and provide independent, autoscaling compute per branch. This is what makes it practical to create branches freely and at scale, without provisioning additional infrastructure.

This architecture also enables time travel. Because every version of the data is retained within a configurable restore window, you can create a branch from any point in the past, not just from the current state. This is what powers instant point-in-time recovery: instead of replaying WAL logs or restoring a backup, you create a branch at the timestamp you need and read directly from it.

What database branching unlocks for your team

Once database branching is a fast, cheap primitive instead of an expensive copy operation, new workflows become practical. Here is a summary of the most common patterns. (We cover each of these in detail in the next post in this series.)

One branch per developer. Every engineer gets their own isolated environment with production-like data. No more stepping on each other's changes in a shared dev database. When a branch drifts too far from production, reset it in one command to pull in the latest schema and data. Because branches scale to zero when idle, this pattern stays affordable even on large teams.

One branch per pull request. Automate branch creation when a PR opens and deletion when it merges or closes. Preview deployments on Vercel or Netlify each get their own database branch, so your frontend preview is backed by realistic, isolated data. Migrations run against real data shapes and constraints, not empty test fixtures. This is the workflow that teams adopt first, and it tends to be the one that convinces them to adopt database branching across the board.

One branch per test run. CI pipelines get a fresh, isolated database for every run. No leftover state from previous tests. No waiting for an empty container image to spin up and then be seeded with fake data. No flaky results caused by shared data or test ordering dependencies. Every run starts from the same baseline. For tests that require deterministic data, you can create branches from a fixed point in time or a specific Log Sequence Number (LSN).

Instant recovery. Create a branch from any point in time within your restore window. Inspect dropped tables, debug failed migrations, or audit historical data, all without touching production. Use schema diff to compare the state before and after a change. Export what you need from the recovery branch and then delete it. The whole process takes seconds, not the hours or days that traditional PITR requires.

Ephemeral environments for AI agents. AI agents can provision databases programmatically via the Lakebase API, use them for the duration of a task, and shut them down when done. Platforms can build versioning on top of snapshots: every agent action creates a checkpoint, and users can jump between versions instantly. If an agent runs a bad migration or corrupts data, rolling back is a single API call.

Getting started

Database branching in Databricks Lakebase turns your Postgres database from the slowest part of your development workflow into the fastest.

You can create your first branch in under a minute using the console, CLI, or API. Here is what it looks like from the CLI:

That is it. You now have an isolated Postgres environment with the full schema and data from production, ready to use.

If you are building on Postgres and tired of the overhead that comes with managing database environments, start with a single dev branch. Then try one per PR. Most teams that start with one database branching workflow quickly adopt the rest.

Databricks Lakebase is serverless Postgres built for agents and apps. Learn more at databricks.com/product/lakebase.