惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
L
LINUX DO - 最新话题
Stack Overflow Blog
Stack Overflow Blog
月光博客
月光博客
雷峰网
雷峰网
Apple Machine Learning Research
Apple Machine Learning Research
V
Visual Studio Blog
Attack and Defense Labs
Attack and Defense Labs
O
OpenAI News
The GitHub Blog
The GitHub Blog
A
About on SuperTechFans
B
Blog RSS Feed
H
Help Net Security
量子位
小众软件
小众软件
SecWiki News
SecWiki News
N
Netflix TechBlog - Medium
TaoSecurity Blog
TaoSecurity Blog
美团技术团队
博客园 - 司徒正美
Hacker News - Newest:
Hacker News - Newest: "LLM"
Recent Commits to openclaw:main
Recent Commits to openclaw:main
The Cloudflare Blog
N
News and Events Feed by Topic
C
Cybersecurity and Infrastructure Security Agency CISA
The Last Watchdog
The Last Watchdog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Scott Helme
Scott Helme
T
The Exploit Database - CXSecurity.com
K
Kaspersky official blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
Threat Research - Cisco Blogs
C
CERT Recently Published Vulnerability Notes
Application and Cybersecurity Blog
Application and Cybersecurity Blog
U
Unit 42
Google DeepMind News
Google DeepMind News
J
Java Code Geeks
Schneier on Security
Schneier on Security
G
Google Developers Blog
Forbes - Security
Forbes - Security
C
CXSECURITY Database RSS Feed - CXSecurity.com
Y
Y Combinator Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
P
Palo Alto Networks Blog
A
Arctic Wolf
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
The Hacker News
The Hacker News
B
Blog
D
DataBreaches.Net
Simon Willison's Weblog
Simon Willison's Weblog

Databricks

Why Talent Transformation Is the Missing Focus of Enterprise AI Public Health Intelligence Shouldn't Require a Data Scientist Mean Time to Detect Is a Data Access Problem First-party audience data is the ad sales relationship now Rethinking Distributed Systems for Serverless Performance and Reliability The AI Scaling Gap Hiding in Digital Native Companies 10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks AI success starts with clean data, not just better models How nOps Rebuilt Their Cloud Optimization Platform on Databricks Lakebase, and Why Other ISVs Should Too Peril Predicts: Precision Payouts for a Volatile World The foundation of AI scalability: one team, one platform, one operating model The Federal Data Paradox: Rich in Data, Poor in Access Driving Budapest Forward: How BKK Uses Databricks to Transform City Mobility LLM Vs AI: A Practical Guide to Differences, Use Cases, and Tools Model Risk Governance Is Not the Same as Risk Intelligence Generative AI for Business: A Complete Strategy and Implementation Guide Data Science vs Data Engineering: Choosing Analysis or Infrastructure AI Applications: Tools, Use Cases, and Platforms MLOps vs DevOps: A Practical Guide for Data Scientists and IT Teams Top Data Warehouse Tools For Modern Data Analytics Unlocking SAP Business Context in Databricks with Semantic Metadata Delta Sharing The marketing activation gap has a fix: Databricks and Stitch partner to turn data infrastructure into marketing performance Alert Fatigue Is a Business Risk Backstage with Lakebase Shipping Faster isn’t Learning Faster Why Your OEE Dashboard Is Lying to You The Turbine That Tried to Tell You It Was Failing Predicting Readmissions Isn't Enough. Acting in Time Is. Clinical Trials Run Longer Than They Have To. That's a Patient Problem Network Quality Is a Revenue Problem, Not a Technical One Shelf Availability Starts with Better Demand Visibility When Predicting the Next Hit Requires More Than Intuition Approximate Answers, Exact Decisions: New Sketch Functions for Analytics Companies Winning with AI Built the Data Layer First Rethinking SQL ETL for modern data platforms Stripe data now available on Databricks via Databricks Marketplace Databricks and Stripe Projects: Infrastructure Built for Agents Agents are ready but your architecture probably isn't Interoperability Between Unity Catalog and Google BigQuery via Catalog Federation Built In, Not Bolted On: What AI-Native Actually Means in Cybersecurity Operationalizing AI for public sector fraud prevention From months to minutes: Building real-time clinical data pipelines with natural language Agentic Data Engineering with Genie Code and Lakeflow Securely send first-party conversion signals with Snapchat Conversions API on Databricks Marketplace How leading tech companies are killing the builder’s tax with Lakebase Inside one of the first production deployments of Lakebase: LangGuard's agentic workflow governance engine The next generation of Databricks Genie Model Risk Management in 2026: A Banker’s Guide to the Revised Interagency Guidance OpenAI GPT-5.5 now available on Databricks, fully-governed through Unity AI Gateway Operational databases: How they work and when to use them Databricks partners with OpenAI on GPT-5.5 Announcing the Public Preview of Lakeflow Designer Are LLM agents good at join order optimization? How conversational analytics removes the BI bottleneck How to transform document activation workflows with Genie and Agent Bricks Beyond the spreadsheet: how Databricks is delivering the modern CFO in Financial Services AI App Development: Guide To Building AI-Powered Apps IoT in Manufacturing: Strategy, Components, Use Cases, and Challenges Stop Hand-Coding Change Data Capture Pipelines Multimodal Data Integration: Production Architectures for Healthcare AI Personalization Strategies for Media Companies A Modern AI Risk Management Framework Introducing the Databricks Excel Add-in for Business Users Real-Time Decisioning for AI Agents: Why you Need a Customer Context Layer First A Practical Guide to LLM Fine Tuning AI Data Transformation Guide for Data Engineers and Data Scientists Concurrency Control in DBMS: How Locking, MVCC and Optimistic Strategies Keep Data Consistent Bridging data science and marketing: Databricks unveils Delta Sharing integration for Adobe Experience Platform and agentic marketing workflows Take Control: Customer-Managed Keys for Lakebase Postgres Get hands on with agents, vibe coding and more at Data+ AI Summit Mercedes-Benz Builds a Cross-Cloud Data Mesh with Delta Sharing and Intelligent Replication, Cutting Costs by 66% What Is a Transactional Database? Introducing Genie Agent Mode Governing coding agent sprawl with Unity AI Gateway Governing Coding Agent Sprawl with Unity AI Gateway What is pgvector? Banks Don’t Have an AI Problem – They Have a Data Platform Problem Open Platform, Unified Pipelines: Why dbt on Databricks is Accelerating Why Your Agents Can’t Read Enterprise Documents — and How to Fix It Building with Databricks Document Intelligence and Lakeflow Databricks on Google Cloud: Innovate Faster. Smarter. Together. Introducing the Databricks Connector for Google Sheets: Real-Time, Governed Lakehouse Data in the Sheets Users Love Unity AI Gateway: How to connect agents to external MCPs securely Expanding agent governance with Unity AI Gateway Agentic reasoning in practice: Making sense of structured and unstructured data Agent Bricks: The Governed Enterprise Agent Platform 8 AI and data trends shaping financial services in 2026 Building real-time product search on Databricks Lovable + Databricks: Build Data-Driven Apps at the Speed of Thought Memory scaling for AI agents Powering clinical research innovation: How TriNetX uses Databricks to accelerate drug development Database Branching in Postgres: Git-Style Workflows with Databricks Lakebase How Zalando built a unified data foundation for AI and analytics on Databricks The next era of the open lakehouse: Apache Iceberg™ v3 in Public Preview on Databricks How FSIs eliminate silos between clients, operations, and finance How MakeMyTrip achieved millisecond personalization at scale with Databricks A multi-agent approach to audience intelligence AiChemy: Next-generation agent with MCP, skills and custom data for drug discovery Accelerate business insights with Lakeflow Connect, now with a Free Tier Unlocking Next-Gen Customer Experiences with Data Intelligence for Marketing
How Daikin Applied Americas builds consistent data pipelines at scale with Genie Code
Trent Lezer · 2026-06-25 · via Databricks

Agentic data engineering is changing how pipelines are built

Daikin Applied Americas (DAA) manufactures and services commercial HVAC systems across North America. That means managing large volumes of operational, manufacturing and service data across systems, from equipment telemetry and supply chain data to field service records.

The data team supports analytics and AI use cases across engineering, operations and customer service, all of which depend on reliable, well-structured pipelines.

As those demands grew, so did the pressure on the data team, including more pipelines, more use cases and more coordination across teams. To address this, the team defined a more structured operating model for how pipelines are designed, built and governed, and used Databricks Genie Code to accelerate execution within that model.

The team leveraged Genie Code as an AI-assisted approach to data engineering. Working directly against governed data in Unity Catalog can help plan and generate multi-step pipelines across the workflow. This allows engineers to move from an idea to a working pipeline much faster, without switching tools or manually stitching components together.

That speed fundamentally changed how the team worked. Pipelines that previously took days to prototype could be generated in minutes. Iteration cycles were shortened, and engineers spent less time writing boilerplate and more time refining logic and outcomes.

At the same time, operating in a large, shared data environment requires consistency. Pipelines must follow common architectural patterns, use shared definitions and behave predictably across teams.

Large language models introduce a structural challenge in this context. When teams rely on varied prompts or loosely defined instructions, the same request can yield inconsistent outputs and lead to architectural drift over time.

To address this, the DAA team focused on defining how AI should operate within a governed enterprise environment, rather than relying solely on prompt engineering.

As Trent Lezer, Sr. Director, Data & Analytics at Daikin Applied Americas, puts it: “Genie Code works best when treated like a junior engineer who works fast but must respect the same architectural constraints as everyone else, no special exemptions ‘because it’s AI.’”

Scaling data engineering through reusable skills

Early usage of Genie Code followed a familiar pattern: long prompts that attempted to encode architecture rules, naming standards, transformation logic and documentation requirements in a single block of text.

This approach did not scale. Instructions varied across teams, prompts became difficult to maintain and similar tasks produced inconsistent outputs.

To address this, the team introduced a MECE (Mutually Exclusive, Collectively Exhaustive) skill framework. As Trent explains: “We implemented a MECE skill framework, each skill defines one coherent competency, skills are non-overlapping and the full set covers the entire lifecycle of data engineering work.”

Each skill defines a specific capability in the data engineering lifecycle. Together, the skills are non-overlapping and cover the full workflow. These skills include medallion architecture design, source readiness and grain definition, transformation patterns, canonical alignment and governance standards.

Instead of embedding rules inside prompts, the team structured the environment so Genie Code loads the appropriate skills at runtime and applies them during planning and execution. This shifts behavior from interpreting ad hoc instructions to operating within a defined execution model.

From a governance perspective, this also changes how standards are enforced. As James VanGordon, Solutions Architect at Databricks, notes: “The pattern I keep seeing with Genie Code is pretty simple: prompts get you started, but they are a bad place to enforce team standards. If the same rule matters more than once, it should live in the workspace as a skill, where Genie Code can actually use it.”

He also emphasizes embedding standards directly into the execution environment: “That is what makes this real instead of wishful thinking. The skills, Unity Catalog context and Genie Code are working in the same place. The guidance sits where the work is being created, not off to the side in a review process someone has to remember later.”

Using the medallion architecture to guide pipeline development

The team also strengthened the role of the medallion architecture as both a governance and reasoning framework. Bronze, Silver and Gold layers already existed, but the shift was making them explicit decision boundaries during pipeline generation, not just storage tiers.

Bronze represents raw source truth. Silver represents cleaned and conformed data. Gold represents business-ready analytics.

To operationalize this structure, the team introduced checkpoints between layers. Before data advances, requirements such as source grain definition, join validation and data stability checks must be satisfied.

These checkpoints are enforced within the development workflow itself, not as downstream review steps. Genie Code operates within these constraints as pipelines are generated and modified.

This ensures consistency across teams while reducing the risk of architectural shortcuts during rapid development.

Connecting pipelines to business concepts

A recurring challenge in enterprise data engineering is aligning technical models with business language.

At DAA, stakeholders think in terms of equipment, customers, service events and contracts, not tables, joins or transformations.

To address this, the team anchored pipeline design in stable business entities. Rather than starting with technical structures, engineers begin by identifying what the data represents and how it behaves over time.

This shift improves downstream efforts and reduces ambiguity when datasets are reused across domains.

Over time, Silver-layer models and Gold datasets become more consistent because they are grounded in shared business concepts rather than isolated technical decisions.

What changed for the team

With this operating model in place and AI embedded, the team saw a clear shift in how work was executed.

Pipeline development accelerated, particularly during early exploration and iteration. Engineers spent less time writing boilerplate code and more time refining business logic.

Outputs also became more consistent across teams. Similar use cases followed similar structural patterns, improving maintainability and reuse.

Importantly, trust in generated outputs increased. Engineers spent less time validating structural correctness and could iterate more quickly.

Standardizing decision-making within the development workflow

To make these gains repeatable, the team standardized key decisions within the development process.

Rather than relying on implicit knowledge, definitions were made explicit, including what qualifies as Bronze, Silver and Gold data, how source grain is defined, which transformation patterns are reusable and how business entities are represented. This structure was critical for scale. It ensures AI operates within a consistent framework across teams, even as use cases evolve.

The payoff: what this unlocked at scale

The result of this operating model was not just faster pipelines. It was the ability to scale data engineering in a governed enterprise environment.

Faster delivery with fewer corrections

Engineers spend less time fixing structurally incorrect pipelines and more time refining logic and business outcomes.

Reduced architectural drift across teams

Consistent application of skills and governance checkpoints prevents divergence across teams working on similar data challenges.

Stronger alignment between engineering and business

Grounding pipelines in business concepts improves clarity and reduces downstream rework.

Scalable governance without manual overhead

Guardrails are embedded directly into the system, reducing reliance on manual enforcement.

Increased trust in AI-generated outputs

Because defined skills and checkpoints constrain outputs, AI operates reliably within production workflows.

As Trent summarizes: “The goal isn’t to make AI follow more rules. It’s to make the right rules impossible to ignore.”

Conclusion

At Daikin Applied Americas, combining a structured operating model with AI-assisted development allowed the data team to scale faster while maintaining consistency, clarity and control.

By defining how pipelines should be built and embedding those rules directly into the development environment, the team created a system in which speed and governance reinforce each other rather than compete.

Learn more about Genie Code.