惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
C
Check Point Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
J
Java Code Geeks
博客园 - 【当耐特】
H
Hacker News: Front Page
S
Secure Thoughts
博客园_首页
Engineering at Meta
Engineering at Meta
N
News | PayPal Newsroom
美团技术团队
SecWiki News
SecWiki News
U
Unit 42
The Hacker News
The Hacker News
有赞技术团队
有赞技术团队
T
The Exploit Database - CXSecurity.com
M
MIT News - Artificial intelligence
T
Threat Research - Cisco Blogs
V
Vulnerabilities – Threatpost
TaoSecurity Blog
TaoSecurity Blog
The Last Watchdog
The Last Watchdog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
F
Fortinet All Blogs
T
Tor Project blog
T
Tailwind CSS Blog
Scott Helme
Scott Helme
Recorded Future
Recorded Future
Know Your Adversary
Know Your Adversary
The Register - Security
The Register - Security
Google DeepMind News
Google DeepMind News
Microsoft Security Blog
Microsoft Security Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Google DeepMind News
Google DeepMind News
S
Security @ Cisco Blogs
S
SegmentFault 最新的问题
Apple Machine Learning Research
Apple Machine Learning Research
月光博客
月光博客
阮一峰的网络日志
阮一峰的网络日志
H
Heimdal Security Blog
MongoDB | Blog
MongoDB | Blog
S
Securelist
C
CXSECURITY Database RSS Feed - CXSecurity.com
雷峰网
雷峰网
博客园 - 聂微东
S
Schneier on Security
T
Tenable Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
B
Blog

Databricks

Why Talent Transformation Is the Missing Focus of Enterprise AI Public Health Intelligence Shouldn't Require a Data Scientist Mean Time to Detect Is a Data Access Problem First-party audience data is the ad sales relationship now Rethinking Distributed Systems for Serverless Performance and Reliability The AI Scaling Gap Hiding in Digital Native Companies 10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks AI success starts with clean data, not just better models How nOps Rebuilt Their Cloud Optimization Platform on Databricks Lakebase, and Why Other ISVs Should Too Peril Predicts: Precision Payouts for a Volatile World The foundation of AI scalability: one team, one platform, one operating model The Federal Data Paradox: Rich in Data, Poor in Access Driving Budapest Forward: How BKK Uses Databricks to Transform City Mobility LLM Vs AI: A Practical Guide to Differences, Use Cases, and Tools Model Risk Governance Is Not the Same as Risk Intelligence Generative AI for Business: A Complete Strategy and Implementation Guide Data Science vs Data Engineering: Choosing Analysis or Infrastructure AI Applications: Tools, Use Cases, and Platforms MLOps vs DevOps: A Practical Guide for Data Scientists and IT Teams Top Data Warehouse Tools For Modern Data Analytics Unlocking SAP Business Context in Databricks with Semantic Metadata Delta Sharing The marketing activation gap has a fix: Databricks and Stitch partner to turn data infrastructure into marketing performance Alert Fatigue Is a Business Risk Backstage with Lakebase Shipping Faster isn’t Learning Faster Why Your OEE Dashboard Is Lying to You The Turbine That Tried to Tell You It Was Failing Predicting Readmissions Isn't Enough. Acting in Time Is. Clinical Trials Run Longer Than They Have To. That's a Patient Problem Network Quality Is a Revenue Problem, Not a Technical One Shelf Availability Starts with Better Demand Visibility When Predicting the Next Hit Requires More Than Intuition Approximate Answers, Exact Decisions: New Sketch Functions for Analytics Companies Winning with AI Built the Data Layer First Rethinking SQL ETL for modern data platforms Stripe data now available on Databricks via Databricks Marketplace Databricks and Stripe Projects: Infrastructure Built for Agents Agents are ready but your architecture probably isn't Interoperability Between Unity Catalog and Google BigQuery via Catalog Federation Built In, Not Bolted On: What AI-Native Actually Means in Cybersecurity Operationalizing AI for public sector fraud prevention From months to minutes: Building real-time clinical data pipelines with natural language Agentic Data Engineering with Genie Code and Lakeflow Securely send first-party conversion signals with Snapchat Conversions API on Databricks Marketplace How leading tech companies are killing the builder’s tax with Lakebase Inside one of the first production deployments of Lakebase: LangGuard's agentic workflow governance engine The next generation of Databricks Genie Model Risk Management in 2026: A Banker’s Guide to the Revised Interagency Guidance OpenAI GPT-5.5 now available on Databricks, fully-governed through Unity AI Gateway Operational databases: How they work and when to use them Databricks partners with OpenAI on GPT-5.5 Announcing the Public Preview of Lakeflow Designer Are LLM agents good at join order optimization? How conversational analytics removes the BI bottleneck How to transform document activation workflows with Genie and Agent Bricks Beyond the spreadsheet: how Databricks is delivering the modern CFO in Financial Services AI App Development: Guide To Building AI-Powered Apps IoT in Manufacturing: Strategy, Components, Use Cases, and Challenges Stop Hand-Coding Change Data Capture Pipelines Multimodal Data Integration: Production Architectures for Healthcare AI Personalization Strategies for Media Companies A Modern AI Risk Management Framework Introducing the Databricks Excel Add-in for Business Users Real-Time Decisioning for AI Agents: Why you Need a Customer Context Layer First A Practical Guide to LLM Fine Tuning AI Data Transformation Guide for Data Engineers and Data Scientists Concurrency Control in DBMS: How Locking, MVCC and Optimistic Strategies Keep Data Consistent Bridging data science and marketing: Databricks unveils Delta Sharing integration for Adobe Experience Platform and agentic marketing workflows Take Control: Customer-Managed Keys for Lakebase Postgres Get hands on with agents, vibe coding and more at Data+ AI Summit Mercedes-Benz Builds a Cross-Cloud Data Mesh with Delta Sharing and Intelligent Replication, Cutting Costs by 66% What Is a Transactional Database? Introducing Genie Agent Mode Governing coding agent sprawl with Unity AI Gateway Governing Coding Agent Sprawl with Unity AI Gateway What is pgvector? Banks Don’t Have an AI Problem – They Have a Data Platform Problem Open Platform, Unified Pipelines: Why dbt on Databricks is Accelerating Why Your Agents Can’t Read Enterprise Documents — and How to Fix It Building with Databricks Document Intelligence and Lakeflow Databricks on Google Cloud: Innovate Faster. Smarter. Together. Introducing the Databricks Connector for Google Sheets: Real-Time, Governed Lakehouse Data in the Sheets Users Love Unity AI Gateway: How to connect agents to external MCPs securely Expanding agent governance with Unity AI Gateway Agentic reasoning in practice: Making sense of structured and unstructured data Agent Bricks: The Governed Enterprise Agent Platform 8 AI and data trends shaping financial services in 2026 Building real-time product search on Databricks Lovable + Databricks: Build Data-Driven Apps at the Speed of Thought Memory scaling for AI agents Powering clinical research innovation: How TriNetX uses Databricks to accelerate drug development Database Branching in Postgres: Git-Style Workflows with Databricks Lakebase How Zalando built a unified data foundation for AI and analytics on Databricks The next era of the open lakehouse: Apache Iceberg™ v3 in Public Preview on Databricks How FSIs eliminate silos between clients, operations, and finance How MakeMyTrip achieved millisecond personalization at scale with Databricks A multi-agent approach to audience intelligence AiChemy: Next-generation agent with MCP, skills and custom data for drug discovery Accelerate business insights with Lakeflow Connect, now with a Free Tier Unlocking Next-Gen Customer Experiences with Data Intelligence for Marketing
Lakeflow: A new era of agentic data engineering
Bilal Aslam · 2026-06-16 · via Databricks

All analytics, AI, and applications start with data. Over the past few decades, data engineering tools have proliferated across a range of use cases and user personas. The result is that most enterprises end up with a very complex and fragmented data stack that is hard to integrate, maintain, or govern. With AI powering all data and AI practitioners, even more pressure will be placed on these brittle data stacks. 

This is why we set out to build Databricks Lakeflow, a unified platform for all of data engineering from ingestion, to transformation and orchestration. All Lakeflow capabilities are fully integrated and centrally governed by Unity Catalog. In the agentic era, this unified architecture offers significant advantages, enabling agents not only to build but also to operate your data pipelines. Today, at the Data + AI Summit, we’re announcing the next major evolution of Databricks Lakeflow.

Lakeflow: Connect, Spark Declarative Pipelines, Jobs, Designer

Genie Code and Lakeflow Designer: agentic pipeline development

Genie Code is now deeply integrated into every aspect of the Lakeflow user experience. You can use Genie Code to create ingestion connectors, build pipelines in Python and SQL and develop jobs with tasks, triggers and dependencies. All of this is made possible with the unified data engineering stack so that Genie Code has full and end-to-end context across your ingestion, transformation, and orchestration workloads.

Now generally available, Lakeflow Designer democratizes data engineering across the enterprise. This visual, AI-powered, no-code interface empowers teams to develop pipelines using a drag-and-drop canvas and natural language prompts. Business analysts and non-technical users can build production-ready ETL pipelines without writing code. Every visual Flow built in Designer natively runs on a production-ready Spark Declarative Pipeline, ensuring zero translation loss without complex handoffs. Data engineers can easily review and refine this code directly in place without switching context or rewriting logic.

Genie ZeroOps: Put data and AI operations on autopilot

Announced today, Genie ZeroOps, helps data teams operate data and AI assets in production. Genie ZeroOps is a purpose-built background AI agent that monitors and manages data and AI assets. ZeroOps detects failures and performs root-cause analysis to identify what went wrong using data quality metrics, error logs, and lineage data from Unity Catalog. Furthermore, it generates proposed fixes and validates them in a safe and isolated sandbox environment governed by Unity Catalog. Applying a fix is done with human-in-the-loop, so Genie ZeroOps does the heavy lifting, and you stay in control. Similar to agentic development, the functionality of Genie ZeroOps is only possible because of the full context awareness and end-to-end governance enabled by a unified data stack with Lakeflow.

Lakeflow Connect: Fast-growing ecosystem with 100+ built-in connectors

Automated pipelines are only as valuable as the data flowing through them. To build a complete "enterprise memory," and ground AI agents like Databricks Genie, you need seamless access to the latest governed context spanning every area of your business. Lakeflow Connect simplifies this process by incrementally ingesting fresh data from an ever-growing list of enterprise systems directly into Unity Catalog-governed Delta tables.

Today, we’re announcing that Lakeflow Connect is expanding to support more than 100 native, managed connectors across enterprise applications, databases, file sources and cloud storage. You can now eliminate brittle third-party tools and run optimized ingestion pipelines for the use cases that customers need most:  

  • Enterprise Knowledge Management: Unify business data from Jira (Beta), GitHub (Beta), and Confluence (GA) alongside unstructured documents, contracts, and PDFs from SharePoint (GA), Google Drive (Beta), and Outlook (Beta). Power context-aware AI applications, support agents, and intelligent document processing on a single foundation.
  • MarTech: Ingest campaign and customer data directly from Meta Ads (Beta), TikTok Ads (Beta), Google Ads (Beta), and HubSpot (GA) to drive real-time personalization.
  • IT & Security Operations: Centralize logs and telemetry for robust SIEM analysis. 
  • Query-based capture for all database connectors and Lakehouse Federation sources (GA):  Query the database directly for change capture without the need for log parsing.

For organizations with specialized or proprietary systems, Community Connectors (Beta) provide an open source solution built on Databricks. Deploy a pre-built connector from the community or build your own to share across your organization or the broader ecosystem. 

Panasonic used Lakeflow Connect to unify data from SAP, Workday, and SharePoint, replacing brittle, legacy ETL with a single platform for real-time, governed intelligence. 

“By moving from a rigid legacy ETL stack to the Databricks Platform, our BI teams can now easily discover and access critical data, cutting Power BI refresh times by 50%. We’re turning external, inconsistent data into trusted, production-grade assets that unlock new business insights and strengthen Panasonic’s competitive edge.”—Jerry Deng, BI Director, Panasonic

We’re also making it easier for organizations to permanently lower the TCO of high-volume ingestion with the Lakeflow Connect Free Tier. Customers automatically receive 100 free DBUs per day, supporting up to 100 million records daily across popular managed SaaS and database connectors.

Zerobus Ingest: Kafka-free ingestion for your data producers

Zerobus Ingest is changing how organizations handle high-volume event data, no message bus required. Near real-time writes in under 5 seconds and high throughput up to 100MB/s (over 10GB/s per table), Zerobus delivers data directly to your platform at scale.

However, performance only matters if your producers can connect without friction. A migration should be as simple as a config change. Since reaching General Availability earlier this year, Zerobus has expanded to meet your data producers where they already operate: 

  • Kafka-Compatible APIs (Beta): Your existing Kafka producers push data straight to Databricks—no code changes required.
  • gRPC & REST APIs (GA): Persistent gRPC streams for high-performance applications, or stateless REST APIs for webhooks and serverless functions.
  • SDK Ecosystem (GA): Production-ready SDKs for Python, Java, Rust, Go, and TypeScript make it easy to embed Zerobus directly into your custom applications.
  • OpenTelemetry (Public Preview): Send metrics, traces, and logs directly to the lakehouse with just a configuration change.

This multi-interface flexibility provides a direct, low-latency bridge to the cloud for global enterprises. For example, Meta has been using Zerobus Ingest to bridge its on-premises data centers to the cloud, enabling rapid development of data-driven solutions at scale.

“We cut our end-to-end pipeline latency to under a minute with Zerobus Ingest and Spark Declarative Pipelines, enabling faster time to value.”—Srikanth Sakhamuri, Data Engineering Leader, Meta

Once data lands in Unity Catalog-governed Delta tables, it’s instantly accessible to downstream AI and BI tools like Databricks Genie. As part of an end-to-end real-time analytic stack, Zerobus ingests the data and processes it using Real-Time Mode in Apache Spark™Declarative Pipelines (SDP) transforms it, and Lakehouse//RT, a new data warehouse type running on a fully native real-time engine, serves it at millisecond-scale performance. 

Spark Declarative Pipelines: Batch and streaming, SQL and Python, and now real-time

Achieving ultra-low-latency streaming has traditionally forced data teams to manage complex, fragmented architectures, often requiring the maintenance of a second specialized engine, such as Apache Flink, alongside Spark. Databricks initially solved this dual-engine complexity by introducing Real-Time Mode (RTM) for Spark Structured Streaming. By shifting from periodic microbatching to continuous stream processing, RTM currently powers pipelines for global brands including Coinbase, DraftKings, and MakeMyTrip.

Now, we are bringing that same power to our unified ETL product: Real-Time Mode (RTM) for Spark Declarative Pipelines is now in Public Preview. RTM for SDP achieves end-to-end latencies as low as 5 milliseconds without the complexity and cost of managing separate engines. Available on both classic and serverless compute, it delivers ultra-low-latency streaming alongside Spark Declarative Pipelines’ operational benefits: versionless execution, automated infrastructure upgrades, and low-to-zero-downtime maintenance.

Next, we are making the declarative APIs from Spark Declarative Pipelines— including Append, Auto CDC, incremental Replace Where, and Materialized View— available everywhere on the Databricks Platform. This means users can take advantage of incremental data processing directly from the product, compute type, and user interface they already know. All of these APIs are now available in Databricks SQL and will be available in serverless Notebooks and Lakeflow Designer in the next few weeks.

Lakeflow Jobs: Now with 50+ integrations

Orchestration shouldn’t be the most challenging part of managing your data pipeline. Whether you are running complex production DAGs, scheduling, or firing AI agents, Lakeflow Jobs is Databricks' native orchestration engine that handles all of these tasks. By bringing managed orchestration and end-to-end observability into every layer of the data lifecycle, data teams are consolidating legacy orchestrators, such as Apache Airflow, onto a single, unified platform.

Data and context-aware orchestration

Every cron schedule is a guess at when data is ready. Lakeflow Jobs lets you stop guessing and start triggering pipelines based on actual data readiness. By using plain English, you can ask Genie to write the SQL triggers that define what “ready” means in your data. Your job fires as soon as conditions are met, respecting your data contracts and ensuring you never process stale data.

“With Lakeflow Jobs, we were able to tap into data that legacy technologies could not access, empowering us to generate deeper, more reliable business insights."—Sachin Wadhwa, Director of Data Architecture and Platforms, The Rank Group

Universal orchestration for anything, anywhere

For customers with data workflows outside Databricks, Lakeflow Jobs provides External Orchestration to natively extend your reach to external systems without requiring you to rebuild integrations from scratch. By using an open operator framework, you can seamlessly trigger Snowflake jobs, fire custom REST APIs, or manage Slack and PagerDuty alerts. Compute is intelligently suspended while waiting for external conditions that may be hours away. We're publishing 40+ operator examples on GitHub and adding dozens of managed integrations in the coming quarters. In addition, every credential flows through Unity Catalog and has a full audit trail.

Getting started with Lakeflow

Lakeflow provides the unified data foundation you need to build reliable, agentic AI applications. To dive deeper into the technical configurations and see these new features in action, explore our hands-on tutorials or review our technical documentation to get started on your next project.

Ready to build? Try Databricks for free to experience Lakeflow today.