惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
D
DataBreaches.Net
Microsoft Azure Blog
Microsoft Azure Blog
大猫的无限游戏
大猫的无限游戏
雷峰网
雷峰网
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
P
Privacy & Cybersecurity Law Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
L
LINUX DO - 最新话题
L
LangChain Blog
量子位
P
Palo Alto Networks Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
L
Lohrmann on Cybersecurity
博客园 - 聂微东
人人都是产品经理
人人都是产品经理
Y
Y Combinator Blog
MongoDB | Blog
MongoDB | Blog
PCI Perspectives
PCI Perspectives
S
SegmentFault 最新的问题
O
OpenAI News
S
Securelist
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
C
Cybersecurity and Infrastructure Security Agency CISA
AWS News Blog
AWS News Blog
G
Google Developers Blog
博客园 - 叶小钗
C
Check Point Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Cisco Talos Blog
Cisco Talos Blog
罗磊的独立博客
V2EX - 技术
V2EX - 技术
小众软件
小众软件
IT之家
IT之家
Engineering at Meta
Engineering at Meta
Hacker News - Newest:
Hacker News - Newest: "LLM"
M
MIT News - Artificial intelligence
T
Threat Research - Cisco Blogs
Vercel News
Vercel News
酷 壳 – CoolShell
酷 壳 – CoolShell
A
About on SuperTechFans
Recorded Future
Recorded Future
N
News and Events Feed by Topic
Cloudbric
Cloudbric
W
WeLiveSecurity
T
The Exploit Database - CXSecurity.com
Martin Fowler
Martin Fowler
C
CXSECURITY Database RSS Feed - CXSecurity.com
The Cloudflare Blog
宝玉的分享
宝玉的分享

Databricks

Why Talent Transformation Is the Missing Focus of Enterprise AI Public Health Intelligence Shouldn't Require a Data Scientist Mean Time to Detect Is a Data Access Problem First-party audience data is the ad sales relationship now Rethinking Distributed Systems for Serverless Performance and Reliability The AI Scaling Gap Hiding in Digital Native Companies 10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks AI success starts with clean data, not just better models How nOps Rebuilt Their Cloud Optimization Platform on Databricks Lakebase, and Why Other ISVs Should Too Peril Predicts: Precision Payouts for a Volatile World The foundation of AI scalability: one team, one platform, one operating model The Federal Data Paradox: Rich in Data, Poor in Access Driving Budapest Forward: How BKK Uses Databricks to Transform City Mobility LLM Vs AI: A Practical Guide to Differences, Use Cases, and Tools Model Risk Governance Is Not the Same as Risk Intelligence Generative AI for Business: A Complete Strategy and Implementation Guide Data Science vs Data Engineering: Choosing Analysis or Infrastructure AI Applications: Tools, Use Cases, and Platforms MLOps vs DevOps: A Practical Guide for Data Scientists and IT Teams Top Data Warehouse Tools For Modern Data Analytics Unlocking SAP Business Context in Databricks with Semantic Metadata Delta Sharing The marketing activation gap has a fix: Databricks and Stitch partner to turn data infrastructure into marketing performance Alert Fatigue Is a Business Risk Backstage with Lakebase Shipping Faster isn’t Learning Faster Why Your OEE Dashboard Is Lying to You The Turbine That Tried to Tell You It Was Failing Predicting Readmissions Isn't Enough. Acting in Time Is. Clinical Trials Run Longer Than They Have To. That's a Patient Problem Network Quality Is a Revenue Problem, Not a Technical One Shelf Availability Starts with Better Demand Visibility When Predicting the Next Hit Requires More Than Intuition Approximate Answers, Exact Decisions: New Sketch Functions for Analytics Companies Winning with AI Built the Data Layer First Rethinking SQL ETL for modern data platforms Stripe data now available on Databricks via Databricks Marketplace Databricks and Stripe Projects: Infrastructure Built for Agents Agents are ready but your architecture probably isn't Interoperability Between Unity Catalog and Google BigQuery via Catalog Federation Built In, Not Bolted On: What AI-Native Actually Means in Cybersecurity Operationalizing AI for public sector fraud prevention From months to minutes: Building real-time clinical data pipelines with natural language Agentic Data Engineering with Genie Code and Lakeflow Securely send first-party conversion signals with Snapchat Conversions API on Databricks Marketplace How leading tech companies are killing the builder’s tax with Lakebase Inside one of the first production deployments of Lakebase: LangGuard's agentic workflow governance engine The next generation of Databricks Genie Model Risk Management in 2026: A Banker’s Guide to the Revised Interagency Guidance OpenAI GPT-5.5 now available on Databricks, fully-governed through Unity AI Gateway Operational databases: How they work and when to use them Databricks partners with OpenAI on GPT-5.5 Announcing the Public Preview of Lakeflow Designer Are LLM agents good at join order optimization? How conversational analytics removes the BI bottleneck How to transform document activation workflows with Genie and Agent Bricks Beyond the spreadsheet: how Databricks is delivering the modern CFO in Financial Services AI App Development: Guide To Building AI-Powered Apps IoT in Manufacturing: Strategy, Components, Use Cases, and Challenges Stop Hand-Coding Change Data Capture Pipelines Multimodal Data Integration: Production Architectures for Healthcare AI Personalization Strategies for Media Companies A Modern AI Risk Management Framework Introducing the Databricks Excel Add-in for Business Users Real-Time Decisioning for AI Agents: Why you Need a Customer Context Layer First A Practical Guide to LLM Fine Tuning AI Data Transformation Guide for Data Engineers and Data Scientists Concurrency Control in DBMS: How Locking, MVCC and Optimistic Strategies Keep Data Consistent Bridging data science and marketing: Databricks unveils Delta Sharing integration for Adobe Experience Platform and agentic marketing workflows Take Control: Customer-Managed Keys for Lakebase Postgres Get hands on with agents, vibe coding and more at Data+ AI Summit Mercedes-Benz Builds a Cross-Cloud Data Mesh with Delta Sharing and Intelligent Replication, Cutting Costs by 66% What Is a Transactional Database? Introducing Genie Agent Mode Governing coding agent sprawl with Unity AI Gateway Governing Coding Agent Sprawl with Unity AI Gateway What is pgvector? Banks Don’t Have an AI Problem – They Have a Data Platform Problem Open Platform, Unified Pipelines: Why dbt on Databricks is Accelerating Why Your Agents Can’t Read Enterprise Documents — and How to Fix It Building with Databricks Document Intelligence and Lakeflow Databricks on Google Cloud: Innovate Faster. Smarter. Together. Introducing the Databricks Connector for Google Sheets: Real-Time, Governed Lakehouse Data in the Sheets Users Love Unity AI Gateway: How to connect agents to external MCPs securely Expanding agent governance with Unity AI Gateway Agentic reasoning in practice: Making sense of structured and unstructured data Agent Bricks: The Governed Enterprise Agent Platform 8 AI and data trends shaping financial services in 2026 Building real-time product search on Databricks Lovable + Databricks: Build Data-Driven Apps at the Speed of Thought Memory scaling for AI agents Powering clinical research innovation: How TriNetX uses Databricks to accelerate drug development Database Branching in Postgres: Git-Style Workflows with Databricks Lakebase How Zalando built a unified data foundation for AI and analytics on Databricks The next era of the open lakehouse: Apache Iceberg™ v3 in Public Preview on Databricks How FSIs eliminate silos between clients, operations, and finance How MakeMyTrip achieved millisecond personalization at scale with Databricks A multi-agent approach to audience intelligence AiChemy: Next-generation agent with MCP, skills and custom data for drug discovery Accelerate business insights with Lakeflow Connect, now with a Free Tier Unlocking Next-Gen Customer Experiences with Data Intelligence for Marketing
From "What Happened?" to "What Will Happen?"
2026-05-21 · via Databricks

Business intelligence has always been about answering questions. For most organizations, those questions have been descriptive — what happened last quarter? — or diagnostic — why did churn spike in the Southeast? Databricks Genie has made these questions radically more accessible, enabling business users to get answers in natural language without writing SQL or waiting on an analyst.

But the questions that drive the most consequential decisions are predictive. Which customers are likely to churn next quarter? How will demand shift if we adjust pricing? How likely is this loan applicant to default? Answering these has historically required an entirely different set of tools, skills, and teams — a data scientist exploring the data, validating its fitness for prediction, engineering features, training a model, and maintaining that model as conditions change. The result: a hard boundary between the BI world, where business users operate with confidence, and the predictive analytics world, where only specialized teams can tread.

In a previous blog post, we showed how TabPFN — a foundation model for tabular data from Prior Labs — collapses much of that predictive workflow by delivering production-grade predictions in a single forward pass. But a key bottleneck remained: someone still needed to translate the business question into a well-formed dataset before TabPFN could make a prediction. The model may be instant, but the work that feeds it is not.

Genie as Feature Engineer, TabPFN as Universal Model

This is where Genie's role shifts from answering questions to enabling predictions. Genie already understands an organization's data — its schemas, relationships, and business semantics. By combining Genie with TabPFN within a multi-agent orchestrator, we create a closed loop: Genie dynamically translates a natural language question into the precise input data TabPFN needs, and TabPFN transforms that data into a prediction in a single forward pass. Every predictive question asked during the conversation received a tailored response on the fly. The space of questions you can answer becomes essentially unbounded — any question that can be framed as "given historical data with an outcome, predict an outcome for a new scenario" can be answered in seconds.

The result is a single, governed experience — grounded in Lakehouse data with full lineage and access control through Unity Catalog — where business users ask predictive questions in the same conversational interface they use for descriptive analytics.

In this post, we walk through the application architecture that makes this possible, introducing each technical component and showing how they come together to deliver predictive intelligence directly within conversational BI.

Video 1. Interacting with a multi-agent supervisor with Genie and TabPFN via a Databricks Apps interface

Architecture: A Multi-Agent Supervisor

The system is built as a multi-agent orchestrator deployed as a Databricks App, which connects the primary components using Agent Bricks, a platform for building and deploying enterprise agents on Databricks. Genie acts as a subagent for structured SQL analytics over governed Lakehouse data. TabPFN is connected to Unity Catalog as an external MCP server. The system also supports additional subagents and serving endpoints; other Databricks applications, or additional MCP servers, can be added as needed.

When a predictive question arrives, the orchestrator executes an agentic workflow. It interprets the user’s business intent. If answering the question requires predictive analysis, it queries Genie to extract the appropriate labeled data from the Lakehouse. After it has gathered all necessary data, it calls TabPFN, passing this data to the model in the right format. Finally, the supervisor interprets the predictions and delivers an actionable recommendation to the user (Figure 1).

Multi-agent supervisor architecture

Figure 1. Multi-agent supervisor architecture combining Databricks Genie and TabPFN via MCP to enable real-time predictive and descriptive analytics for business users

The Core Insight in Action

To make this concrete, consider what happens when a sales leader asks: "Which promotion type would most likely close the Horton-Cross deal?"

In a traditional workflow, answering this question requires a data scientist to understand the question and identify which tables and columns matter; extract the right training set from historical deals that include promotion types and win/loss outcomes; select an algorithm, tune hyperparameters, and validate performance; prepare inference data specific to the Horton-Cross deal; run the model; and translate the output into a business recommendation. Each of these steps takes time, expertise, and iteration. And the next question — "What is the optimal date to follow up to maximize win probability?" — requires an entirely different model built from scratch.

Now consider what happens with Genie and TabPFN under the same multi-agent supervisor. The supervisor interprets the natural language question and its semantic intent, then translates that intent into a specific request for Genie to generate a dataset. Genie recognizes that answering this question requires historical opportunities joined with promotions and accounts, using win or loss as the label, and generates precise SQL to extract this data instantly.

TabPFN receives that dataset and generates predictions in a single forward pass — no feature preprocessing, no model selection, no hyperparameter tuning. Finally, the supervisor returns a clear, data-driven recommendation. The entire pipeline — from question to prediction — assembles itself from natural language in a single conversation turn.

Assessing Quality and Limitations

The pattern has limitations: TabPFN is only as good as the data Genie produces. If Genie cannot construct a meaningful dataset with a clear label column for a given question, because the schema does not capture the right signal, the necessary joins do not exist, or the outcome is not represented in the data, then the prediction will not be reliable, regardless of how capable TabPFN is. See the best practices for building an effective Genie space here. On top of this, there is also a broader risk that an agent may hallucinate or omit key information during a multi-turn conversation.

That is exactly why systematic evaluation is essential. Unlike a static ML pipeline that must be validated once before deployment, this system dynamically constructs a distinct ML problem for each question. We need an evaluation framework to understand where the boundary lies: which classes of questions produce reliable predictions, and which ones exceed what Genie can express as a well-formed training set.

The solution accelerator ships with a comprehensive evaluation harness built on MLflow’s GenAI evaluation framework. It runs against the live agent and logs results to MLflow Experiment Tracking, giving teams a single pane of glass to evaluate and monitor quality over time. You can find the full details here.

Video 2. Evaluating a multi-agent supervisor with Genie and TabPFN via Databricks Experiments interface.

Without this evaluation loop, the system may confidently return predictions with no way to distinguish trustworthy from unreliable ones. This rigorous approach ensures coverage at every level: it catches conversational and behavioral regressions while also validating end-to-end correctness of the predictive pipeline. Together, these checks give teams the confidence to deploy this pattern in production, with a clear understanding of which question classes produce reliable predictions and where the system boundaries lie.

Get Started

The combination of Genie, TabPFN, and Agent Bricks reframes the relationship between descriptive and predictive analytics. Genie becomes the feature engineering layer. TabPFN removes the training and maintenance overhead. Agent Bricks provides the orchestration and governance backbone, while MLflow evaluates and monitors the quality of the responses. The result is that business users can ask predictive questions in the same conversational interface they already use for descriptive analytics.

The full Solution Accelerator is available here. The repository includes sample data generation, Genie Space configuration and the end-to-end evaluation harness described above. The pattern is domain-agnostic: while the accelerator demonstrates enterprise sales analytics, the same architecture applies to any domain where structured data with outcomes exists, including healthcare risk scoring, manufacturing quality prediction, financial fraud detection, customer churn analysis, and beyond.

Get started today and bring predictive intelligence to the conversations your teams are already having.