惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
S
Securelist
GbyAI
GbyAI
The Register - Security
The Register - Security
B
Blog
Recorded Future
Recorded Future
D
DataBreaches.Net
C
Cybersecurity and Infrastructure Security Agency CISA
A
About on SuperTechFans
C
CERT Recently Published Vulnerability Notes
T
The Blog of Author Tim Ferriss
Vercel News
Vercel News
Google DeepMind News
Google DeepMind News
S
Schneier on Security
S
SegmentFault 最新的问题
Martin Fowler
Martin Fowler
T
Tenable Blog
T
The Exploit Database - CXSecurity.com
阮一峰的网络日志
阮一峰的网络日志
宝玉的分享
宝玉的分享
AWS News Blog
AWS News Blog
L
Lohrmann on Cybersecurity
Spread Privacy
Spread Privacy
N
News | PayPal Newsroom
Engineering at Meta
Engineering at Meta
T
Tor Project blog
The Hacker News
The Hacker News
量子位
酷 壳 – CoolShell
酷 壳 – CoolShell
MongoDB | Blog
MongoDB | Blog
Cyberwarzone
Cyberwarzone
Security Archives - TechRepublic
Security Archives - TechRepublic
爱范儿
爱范儿
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
C
Cyber Attacks, Cyber Crime and Cyber Security
T
Threatpost
WordPress大学
WordPress大学
Google Online Security Blog
Google Online Security Blog
G
GRAHAM CLULEY
Google DeepMind News
Google DeepMind News
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Attack and Defense Labs
Attack and Defense Labs
N
Netflix TechBlog - Medium
SecWiki News
SecWiki News
Hacker News: Ask HN
Hacker News: Ask HN
M
MIT News - Artificial intelligence
Scott Helme
Scott Helme
Microsoft Security Blog
Microsoft Security Blog
H
Help Net Security
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Using Microsoft Fabric Shortcuts to Avoid Duplicate Data Copies
Ravi Kiran Pagidi · 2026-06-05 · via DEV Community

Enterprise data platforms are really good at one thing: creating copies of the same data everywhere. Different teams copy the same curated folders into their own lakehouses, then copy again into another workspace "for reporting," then again for a data science sandbox. Storage grows, pipelines multiply, and nobody is sure which copy is the source of truth anymore.

Microsoft Fabric Shortcuts give us a way out of that pattern by letting a Fabric Lakehouse reference data where it already lives instead of copying it again. You still get a first-class experience in the Lakehouse, SQL endpoint, and Power BI, but the bytes stay in one place.


What Are Microsoft Fabric Shortcuts?

In plain terms, a shortcut in Microsoft Fabric is a logical link that points from your Lakehouse (or other Fabric item) to some other storage location. In the Lakehouse explorer it looks like a regular folder or table, but the data is actually being read from the target location.

Supported shortcut targets today include:

  • Other OneLake locations (files or tables in different workspaces or lakehouses)
  • Azure Data Lake Storage Gen2 (ADLS Gen2) accounts and containers
  • Amazon S3 buckets
  • Dataverse and other external sources via Fabric connectors

Think of a shortcut as a "symbolic link" in OneLake: it has a shortcut path where it appears in your Lakehouse and a target path that points to where the data really lives. No data is physically moved or duplicated when you create one.


Why Duplicate Data Copies Become a Problem

Most enterprises end up with multiple copies of the same core datasets scattered across environments and workspaces. Typical failure modes:

  • Every team builds its own "copy pipeline" from the same source into its own lakehouse or workspace
  • Dev, UAT, and prod end up being fed by slightly different pipelines or schedules, so schema and data drift over time
  • Pipelines exist solely to move data from one lake to another (ADLS to Fabric, or workspace-to-workspace) with no real transformation
  • Storage costs grow linearly with the number of teams and environments, and nobody feels responsible because "storage is cheap"
  • Freshness SLAs become hard to manage, because each copy has its own schedule and failure modes
  • Governance teams now have to manage access and data protection policies across several physical copies of the same sensitive data
  • Debugging becomes painful when teams disagree about which version of a dataset is correct

A concrete example: A customer transactions table is copied from ADLS into a Fabric Lakehouse. Then it is copied again into another lakehouse for reporting. Then copied again for data science experiments. Each copy adds storage cost, a new pipeline to monitor, another access policy to manage, and another potential source of stale or inconsistent data. By the time something breaks, you are not sure which copy is authoritative.


How Fabric Shortcuts Solve This

Shortcuts let teams:

  • Reference data without physically moving it
  • Build logical lakehouse views over existing data sources
  • Reduce redundant ETL/ELT pipelines
  • Keep the original data as the single source of truth
  • Enable multiple teams to consume the same dataset consistently
  • Simplify medallion or domain-based architectures

Here is how the data flow looks when shortcuts are used correctly:

Source Data in ADLS / OneLake / S3
          |
    Fabric Shortcut
          |
    Fabric Lakehouse
          |
  SQL Endpoint / Semantic Model / Power BI / Data Science

Enter fullscreen mode Exit fullscreen mode

And a broader architecture view:

[Enterprise Data Lake (ADLS Gen2 / OneLake)]
          |
          |  Shortcut (no physical copy)
          v
[Fabric Lakehouse: Curated Zone]
          |
          +---> [Power BI Semantic Model]
          |
          +---> [Data Science Notebook]
          |
          +---> [SQL Analytics Endpoint]
          |
          +---> [Downstream Data Product]

Enter fullscreen mode Exit fullscreen mode

The source remains authoritative. Consumers get clean, governed access without owning the underlying data.


Real-World Use Case

Scenario: A large enterprise already has curated Delta tables in Azure Data Lake Storage Gen2. Multiple teams want to use those datasets in Microsoft Fabric for reporting, analytics, and AI use cases. Instead of building new copy pipelines into Fabric, the data engineering team creates shortcuts from a Fabric Lakehouse to the existing curated folders in ADLS.

Implementation steps:

Step 1: Identify trusted curated data in ADLS
Before creating any shortcuts, confirm the source folders contain governed, validated data. Raw or unvalidated folders are not good shortcut candidates.

Step 2: Create a Fabric Lakehouse for the analytics domain
Set up a dedicated lakehouse in the appropriate Fabric workspace. Apply workspace roles and permissions aligned with the consuming team.

Step 3: Add shortcuts to the curated folders
Navigate to the Lakehouse, select New Shortcut, choose ADLS Gen2, provide the connection and folder path, and create the shortcut. It appears immediately as a folder or table reference in the Lakehouse.

Step 4: Validate table structure and permissions
Confirm the shortcut resolves correctly, the data schema is as expected, and that end users have the right access through Fabric's permission model and OneLake security.

Step 5: Build SQL views or semantic models on top
Use the SQL Analytics Endpoint to create views or expose tables. Build a Power BI semantic model on top for reporting teams. Keep the raw shortcut path abstracted from end users.

Step 6: Let reporting and analytics teams consume without extra copies
Reporting, data science, and analytics teams now access the same data through Fabric. No additional pipelines. No additional storage. One source of truth.

Business impact:

  • Less storage duplication
  • Fewer pipelines to build and maintain
  • Faster onboarding of new data products
  • Reduced data freshness issues
  • Better alignment with data governance policies
  • Simpler architecture to explain and audit

Architecture Pattern: Shortcut-Based Lakehouse Consumption

Pattern name: Shortcut-Based Lakehouse Consumption Pattern

Layers:

Layer What It Contains
Source Layer ADLS Gen2, OneLake, S3, Dataverse, existing lakehouses
Shortcut Layer Logical references inside Fabric Lakehouse
Consumption Layer Lakehouse tables, SQL endpoint, notebooks, semantic models
Governance Layer Microsoft Purview, Fabric permissions, workspace roles, OneLake security
Monitoring Layer Pipeline monitoring, usage tracking, access auditing

Architecture diagram:

[Source Systems]
       |
       v
[Raw / Curated Data in ADLS or OneLake]
       |
       |  Fabric Shortcut
       v
[Fabric Lakehouse]
       |
       +---> [SQL Endpoint]
       +---> [Power BI Semantic Model]
       +---> [Data Science / ML]
       +---> [Business Data Product]
       |
       v
[Governance, Security, Monitoring]

Enter fullscreen mode Exit fullscreen mode


Example Implementation

Creating a Shortcut from a Fabric Lakehouse to ADLS Gen2

Steps in the Fabric UI:

  1. Open your Fabric workspace
  2. Create or open an existing Lakehouse
  3. In the Lakehouse explorer, go to Files or Tables
  4. Click New Shortcut
  5. Choose Azure Data Lake Storage Gen2 as the source
  6. Provide the connection details (storage account, container, credential)
  7. Select the target folder
  8. Name the shortcut and create it
  9. Validate the data in Lakehouse Explorer
  10. Use the data from notebooks, SQL endpoint, or Power BI

Reading shortcut data from a PySpark notebook

df = spark.read.format("delta").load("Files/shortcuts/customer_transactions")
display(df.limit(10))

Enter fullscreen mode Exit fullscreen mode

If the shortcut points to a Delta-formatted folder, Spark reads it directly. If the data is in Parquet or CSV, adjust the format accordingly.

Querying via the SQL Analytics Endpoint

SELECT TOP 100 *
FROM lakehouse.customer_transactions;

Enter fullscreen mode Exit fullscreen mode

Note: whether a shortcut appears under Files or Tables in the Lakehouse explorer depends on how it was created and whether the target folder is a recognized Delta table. If it appears under Files only, you can register it as a table using CREATE TABLE in a notebook or via the Lakehouse UI.


When to Use Shortcuts

Good scenarios:

  • You already have trusted, governed data in ADLS or OneLake
  • Multiple teams need access to the same dataset without owning it
  • You want to avoid building copy pipelines just to move data between workspaces
  • You are building domain-oriented or product-oriented lakehouses
  • You want faster analytics access without waiting for a pipeline to run
  • You are connecting Fabric to an existing cloud storage investment

When Not to Use Shortcuts

Shortcuts are not always the right answer. Avoid or be careful when:

  • The source data is not governed, validated, or trusted
  • Permissions are unclear or inconsistently applied at the source
  • Performance requirements need optimized physical layout inside Fabric (compaction, partitioning, Z-ordering)
  • The source folder structure is messy or changes frequently
  • The consuming team expects full ownership and control of the data
  • Cross-cloud latency or egress cost is a concern (for example, S3 shortcuts)
  • The source is controlled by an external team with no SLA alignment

Best Practices

  • Use shortcuts primarily for trusted curated datasets, not raw ingestion zones
  • Keep naming conventions clean and consistent with your lakehouse standards
  • Document shortcut ownership: who created it, what it points to, and who the source owner is
  • Avoid creating shortcuts to random raw folders just because it is convenient
  • Validate access controls before promoting shortcuts to production
  • Use semantic models or SQL views to abstract the shortcut path from end consumers
  • Monitor usage and performance, especially for cross-cloud shortcuts
  • Align shortcuts with data product boundaries, not just individual tables
  • Do not treat shortcuts as a substitute for proper data governance
  • Maintain a clear source-of-truth policy so teams know which shortcut is authoritative

Common Pitfalls

  • Users assume the data is physically stored in Fabric. It is not. If the source is unavailable or deleted, the shortcut breaks.
  • Teams delete or reorganize source folders without knowing shortcuts depend on them. This silently breaks downstream consumption.
  • Permissions work for engineers but fail for business users. Always test access with a non-admin account before go-live.
  • Too many shortcuts create a confusing lakehouse structure. Organize them with clear folder hierarchies and naming.
  • No documentation for where shortcut data comes from. Future team members have no idea what the shortcut points to or why.
  • Shortcuts are used to bypass proper data modeling. A shortcut to a raw table is not a curated data product.

Conclusion

Microsoft Fabric Shortcuts are not just a convenience feature. They are an important architectural pattern for reducing duplicate data copies, simplifying enterprise lakehouse design, and accelerating analytics adoption. Used correctly, they help teams build cleaner, cheaper, and more governable data platforms.

But like any architecture pattern, they need ownership, naming standards, security design, and monitoring. A shortcut without governance is just technical debt with a different shape.

"The best data architecture is not always the one that moves data faster. Sometimes, it is the one that avoids moving data unnecessarily."