惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Privacy & Cybersecurity Law Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
宝玉的分享
宝玉的分享
V
V2EX
爱范儿
爱范儿
Last Week in AI
Last Week in AI
美团技术团队
人人都是产品经理
人人都是产品经理
WordPress大学
WordPress大学
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 叶小钗
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Apple Machine Learning Research
Apple Machine Learning Research
Security Latest
Security Latest
C
Cybersecurity and Infrastructure Security Agency CISA
Know Your Adversary
Know Your Adversary
I
Intezer
K
Kaspersky official blog
阮一峰的网络日志
阮一峰的网络日志
大猫的无限游戏
大猫的无限游戏
T
Tenable Blog
AWS News Blog
AWS News Blog
小众软件
小众软件
博客园 - 司徒正美
Cyberwarzone
Cyberwarzone
NISL@THU
NISL@THU
博客园 - 三生石上(FineUI控件)
C
CERT Recently Published Vulnerability Notes
博客园 - 聂微东
量子位
有赞技术团队
有赞技术团队
S
Schneier on Security
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
S
Secure Thoughts
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
罗磊的独立博客
Hugging Face - Blog
Hugging Face - Blog
V
Visual Studio Blog
Google DeepMind News
Google DeepMind News
L
Lohrmann on Cybersecurity
P
Palo Alto Networks Blog
P
Privacy International News Feed
L
LINUX DO - 最新话题
博客园 - Franky
雷峰网
雷峰网
月光博客
月光博客
Hacker News: Ask HN
Hacker News: Ask HN
Forbes - Security
Forbes - Security
博客园 - 【当耐特】
C
Cyber Attacks, Cyber Crime and Cyber Security

Databricks

Why Talent Transformation Is the Missing Focus of Enterprise AI Public Health Intelligence Shouldn't Require a Data Scientist Mean Time to Detect Is a Data Access Problem First-party audience data is the ad sales relationship now Rethinking Distributed Systems for Serverless Performance and Reliability The AI Scaling Gap Hiding in Digital Native Companies 10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks AI success starts with clean data, not just better models How nOps Rebuilt Their Cloud Optimization Platform on Databricks Lakebase, and Why Other ISVs Should Too Peril Predicts: Precision Payouts for a Volatile World The foundation of AI scalability: one team, one platform, one operating model The Federal Data Paradox: Rich in Data, Poor in Access Driving Budapest Forward: How BKK Uses Databricks to Transform City Mobility LLM Vs AI: A Practical Guide to Differences, Use Cases, and Tools Model Risk Governance Is Not the Same as Risk Intelligence Generative AI for Business: A Complete Strategy and Implementation Guide Data Science vs Data Engineering: Choosing Analysis or Infrastructure AI Applications: Tools, Use Cases, and Platforms MLOps vs DevOps: A Practical Guide for Data Scientists and IT Teams Top Data Warehouse Tools For Modern Data Analytics Unlocking SAP Business Context in Databricks with Semantic Metadata Delta Sharing The marketing activation gap has a fix: Databricks and Stitch partner to turn data infrastructure into marketing performance Alert Fatigue Is a Business Risk Backstage with Lakebase Shipping Faster isn’t Learning Faster Why Your OEE Dashboard Is Lying to You The Turbine That Tried to Tell You It Was Failing Predicting Readmissions Isn't Enough. Acting in Time Is. Clinical Trials Run Longer Than They Have To. That's a Patient Problem Network Quality Is a Revenue Problem, Not a Technical One Shelf Availability Starts with Better Demand Visibility When Predicting the Next Hit Requires More Than Intuition Approximate Answers, Exact Decisions: New Sketch Functions for Analytics Companies Winning with AI Built the Data Layer First Rethinking SQL ETL for modern data platforms Stripe data now available on Databricks via Databricks Marketplace Databricks and Stripe Projects: Infrastructure Built for Agents Agents are ready but your architecture probably isn't Interoperability Between Unity Catalog and Google BigQuery via Catalog Federation Built In, Not Bolted On: What AI-Native Actually Means in Cybersecurity Operationalizing AI for public sector fraud prevention From months to minutes: Building real-time clinical data pipelines with natural language Agentic Data Engineering with Genie Code and Lakeflow Securely send first-party conversion signals with Snapchat Conversions API on Databricks Marketplace How leading tech companies are killing the builder’s tax with Lakebase Inside one of the first production deployments of Lakebase: LangGuard's agentic workflow governance engine The next generation of Databricks Genie Model Risk Management in 2026: A Banker’s Guide to the Revised Interagency Guidance OpenAI GPT-5.5 now available on Databricks, fully-governed through Unity AI Gateway Operational databases: How they work and when to use them Databricks partners with OpenAI on GPT-5.5 Announcing the Public Preview of Lakeflow Designer Are LLM agents good at join order optimization? How conversational analytics removes the BI bottleneck How to transform document activation workflows with Genie and Agent Bricks Beyond the spreadsheet: how Databricks is delivering the modern CFO in Financial Services AI App Development: Guide To Building AI-Powered Apps IoT in Manufacturing: Strategy, Components, Use Cases, and Challenges Stop Hand-Coding Change Data Capture Pipelines Multimodal Data Integration: Production Architectures for Healthcare AI Personalization Strategies for Media Companies A Modern AI Risk Management Framework Introducing the Databricks Excel Add-in for Business Users Real-Time Decisioning for AI Agents: Why you Need a Customer Context Layer First A Practical Guide to LLM Fine Tuning AI Data Transformation Guide for Data Engineers and Data Scientists Concurrency Control in DBMS: How Locking, MVCC and Optimistic Strategies Keep Data Consistent Bridging data science and marketing: Databricks unveils Delta Sharing integration for Adobe Experience Platform and agentic marketing workflows Take Control: Customer-Managed Keys for Lakebase Postgres Get hands on with agents, vibe coding and more at Data+ AI Summit Mercedes-Benz Builds a Cross-Cloud Data Mesh with Delta Sharing and Intelligent Replication, Cutting Costs by 66% What Is a Transactional Database? Introducing Genie Agent Mode Governing coding agent sprawl with Unity AI Gateway Governing Coding Agent Sprawl with Unity AI Gateway What is pgvector? Banks Don’t Have an AI Problem – They Have a Data Platform Problem Open Platform, Unified Pipelines: Why dbt on Databricks is Accelerating Why Your Agents Can’t Read Enterprise Documents — and How to Fix It Building with Databricks Document Intelligence and Lakeflow Databricks on Google Cloud: Innovate Faster. Smarter. Together. Introducing the Databricks Connector for Google Sheets: Real-Time, Governed Lakehouse Data in the Sheets Users Love Unity AI Gateway: How to connect agents to external MCPs securely Expanding agent governance with Unity AI Gateway Agentic reasoning in practice: Making sense of structured and unstructured data Agent Bricks: The Governed Enterprise Agent Platform 8 AI and data trends shaping financial services in 2026 Building real-time product search on Databricks Lovable + Databricks: Build Data-Driven Apps at the Speed of Thought Memory scaling for AI agents Powering clinical research innovation: How TriNetX uses Databricks to accelerate drug development Database Branching in Postgres: Git-Style Workflows with Databricks Lakebase How Zalando built a unified data foundation for AI and analytics on Databricks The next era of the open lakehouse: Apache Iceberg™ v3 in Public Preview on Databricks How FSIs eliminate silos between clients, operations, and finance How MakeMyTrip achieved millisecond personalization at scale with Databricks A multi-agent approach to audience intelligence AiChemy: Next-generation agent with MCP, skills and custom data for drug discovery Accelerate business insights with Lakeflow Connect, now with a Free Tier Unlocking Next-Gen Customer Experiences with Data Intelligence for Marketing
Backstage with Lakebase, part 2
2026-05-15 · via Databricks

In part 1 of this series, we explored how moving Backstage's underlying database to Databricks Lakebase turned risky schema migrations into 1-second branch-and-test operations. But a faster developer cycle only gets you so far if Security and Governance teams are still treating your operational database like a black box.  

In a traditional stack, your application database and your data lake live in two entirely different security paradigms. The ownership graph for your infrastructure lives in Backstage, backed by an isolated RDS instance and governed by complex IAM roles and Postgres native grants. Meanwhile, your warehouse data is governed by the data team using Unity Catalog. Unity Catalog is an Open Source framework created by Databricks that provides a unified governance layer for data, AI, and now operational databases – a single place to manage access controls, audit trails, lineage, and compliance across everything on the platform.

To audit a single table drop on RDS, you'd need to cross-reference CloudTrail for the IAM principal, pg_stat_activity or pgaudit logs for the SQL statement, and CloudWatch for the timestamp, three services, three query languages, three access policies. The operational database becomes a compliance side-channel.

Unity Catalog Absorbs the Operational DB

When we pointed Backstage at Lakebase, we didn't just change where the data lived; we changed where the access policy lived.

Because Lakebase is natively embedded inside Databricks, Unity Catalog extends directly over the operational Postgres database. In this POC, we used Lakehouse Federation to expose the Backstage catalog as a foreign catalog (lakebase_bs) in Unity Catalog. Once it's there, standard UC grants control who can see what, no Postgres-level role management required:

While we didn't build end-to-end Row-Level Security policies for Backstage in this POC, architecturally, the exact same RLS rules that protect sensitive billing tables can be applied directly to these operational tables. The wall between "operational" and "analytical" stops being a physical boundary, and simply becomes an access pattern.

A Unified Audit Trail Out of the Box

Remember the 1-second copy-on-write branching we executed in Part 1? In a traditional setup, proving to a security engineer that a developer only branched the database for an hour and then destroyed it is a manual exercise.

With Lakebase, every control-plane action against the operational database is automatically recorded in system.access.audit. To prove this, we queried the audit log for the exact branch operations from our Part 1 disaster-recovery experiment:

Result:

Every branch creation and deletion from our Part 1 experiments is logged. Each event is tied to a specific OAuth user identity and source IP, captured automatically, and governed by the exact same Row-Level Security controls as every other audit table in Unity Catalog. No CloudTrail cross-referencing. No RDS log parsing. One SQL query.

Automated Cost Attribution by Branch

A governance team doesn't just want to know who created a branch, they want to know what it cost.

In a traditional AWS environment, tracking the cost of an ephemeral RDS instance requires custom CloudWatch tagging strategies that often miss short-lived workloads. Because Lakebase integrates natively with Unity Catalog's system billing tables, compute costs break down automatically by project_idbranch_id, and endpoint_id.

In this POC, the production branch was billed at 31.6130 DBU, while the dropped test branch was independently attributed 0.0107 DBU. The audit trail and the cost trail are governed in the exact same place.

What This Means for Teams That Branch Every Day

Our governance story answers the compliance question: can we prove who did what, when, and what it cost? The answer is yes – one SQL query instead of three services. But there's a second governance question that matters just as much for development teams adopting the branching workflow from Part 1: what happens to governance when your team creates dozens of branches per sprint?

In Part 1, we described a workflow where every feature branch and every pull request gets its own isolated database copy. A team of six developers running two-week sprints might create and destroy 30-40 branches in a single sprint. That's 30-40 copies of production data, each one potentially containing sensitive fields – customer PII, financial records, health data.

This is where Unity Catalog's branch-level governance becomes load-bearing, not just convenient. When a Lakebase branch is created, Unity Catalog's attribute-level masking policies propagate automatically to the new branch. A developer working on their feature branch never sees unmasked production data – not because someone remembered to configure it, but because the governance layer enforces it at creation time. The CI branch that runs your PR tests is governed identically to production. The QA branch where a tester runs destructive scenarios is governed identically to production. There is no "non-production exception" where sensitive data leaks because someone forgot to apply the policy.

This matters more than it might seem. According to Perforce’s 2025 State of Data Compliance report, 60% of organizations have experienced breaches or theft in non-production environments where sensitive data was inadequately anonymized. The traditional approach – manually masking data when provisioning dev/test environments – doesn't scale when environments are created and destroyed in seconds. Governance has to be automatic, or it doesn't happen.

The DBA's New Opportunity

The audit trail and cost attribution data also signal a quieter shift: the DBA's role is evolving from reactive ticket work to strategic platform architecture.

Today, much of a DBA's time goes to operational requests – environment provisioning, schema reviews, data refreshes, access grants. A six-developer team can generate 30+ tickets per sprint, and the DBA's calendar becomes a queue. The expertise that makes DBAs valuable – understanding data integrity, performance, and governance at a deep level – gets buried under repetitive provisioning work.

When branching is self-service and governance is automatic, that repetitive work falls away. Developers provision their own environments in one second. Schema changes are reviewed asynchronously in pull requests – the DBA sees a formatted schema diff posted by CI, reviews it on their own schedule, and approves or requests changes through the normal PR workflow. With the time now available, those reviews go deeper: the DBA helps team members understand the existing data and structures in production, works with them to arrive at better solutions, and conducts thorough reviews that uphold data integrity and governance standards. Data masking is enforced by policy, not by manual intervention. Cost attribution is automatic, not a monthly reconciliation exercise.

What opens up is the work that actually leverages the DBA's expertise: defining branching policies, designing governance rules, architecting promotion workflows, tuning performance, and establishing the guardrails that make self-service safe. The DBA shifts from doing the work to designing how the work gets done – from 30+ operational tickets per sprint to fewer than 5 high-value policy reviews. The audit trail demonstrated above isn't just a compliance artifact – it's the DBA's new strategic dashboard, a real-time view of how the platform is being used and where to invest next.

From Role Shift to Tooling

The DBA's pivot from operational tickets to platform design only works if the tooling shifts with the role. The platform has to do the routine work on its own, and the DBA needs a place to design how that work gets done.

Two open-source tools, both deployed as Databricks Apps and both governed by the same Unity Catalog grants and audit trail described above, close that loop.

LakebaseOps is what the platform does on its own. Three agents – Provisioning, Performance, and Health – replace 51 of the tasks a DBA used to file tickets for. Seven of them run as scheduled Databricks Jobs and replace the pg_cron crontab a DBA would otherwise hand-maintain. A monitoring UI surfaces live pg_stat metrics, slow-query regressions, branch TTL enforcement, and a 9-KPI adoption dashboard. A migration wizard scores ten source engines (Aurora, RDS, Cloud SQL, AlloyDB, Cosmos DB, and more) against Lakebase, with live pricing from the AWS and Azure APIs.

Lakebase MCP is what the DBA does on top of the platform. A Model Context Protocol server exposing 46 tools to any MCP-capable AI agent (Claude, Copilot, GPT). The DBA stops opening pgAdmin and starts describing intent:

Two design choices keep this safe. First, dual-layer governance: a SQL-statement guard and a per-tool access guard, with four pre-built profiles (read_only, analyst, developer, admin) that map onto the same UC access patterns shown above. A coding assistant runs as read_only and physically cannot drop a table.

Second, every query is attributable – the server tags every statement with the originating tool:

Combined with the branch-level cost attribution shown earlier, you can answer "which agent on which branch generated the 4 AM CPU spike?" in one SQL query.

LakebaseOps runs for the team. Lakebase MCP runs with the team. Both inherit the governance posture you just saw.

In Part 3 of this series, we will look at the ultimate payoff: taking the infrastructure ownership data inside Backstage and joining it directly to cloud billing data in a single SQL query.