惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Threatpost
Forbes - Security
Forbes - Security
P
Palo Alto Networks Blog
Scott Helme
Scott Helme
S
Securelist
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
Visual Studio Blog
P
Proofpoint News Feed
大猫的无限游戏
大猫的无限游戏
博客园 - 三生石上(FineUI控件)
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
AWS News Blog
AWS News Blog
P
Privacy & Cybersecurity Law Blog
L
LINUX DO - 热门话题
L
Lohrmann on Cybersecurity
Last Week in AI
Last Week in AI
Engineering at Meta
Engineering at Meta
Spread Privacy
Spread Privacy
博客园_首页
T
Tor Project blog
The Cloudflare Blog
博客园 - 聂微东
罗磊的独立博客
Cyberwarzone
Cyberwarzone
腾讯CDC
T
The Exploit Database - CXSecurity.com
Google DeepMind News
Google DeepMind News
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
P
Privacy International News Feed
D
Darknet – Hacking Tools, Hacker News & Cyber Security
U
Unit 42
A
Arctic Wolf
C
Cybersecurity and Infrastructure Security Agency CISA
A
About on SuperTechFans
G
GRAHAM CLULEY
K
Kaspersky official blog
月光博客
月光博客
Microsoft Security Blog
Microsoft Security Blog
T
Tenable Blog
L
LINUX DO - 最新话题
酷 壳 – CoolShell
酷 壳 – CoolShell
IT之家
IT之家
TaoSecurity Blog
TaoSecurity Blog
Security Archives - TechRepublic
Security Archives - TechRepublic
MongoDB | Blog
MongoDB | Blog
T
The Blog of Author Tim Ferriss
C
Cyber Attacks, Cyber Crime and Cyber Security
博客园 - 司徒正美
O
OpenAI News
Recent Announcements
Recent Announcements

WhatIs

Strategic IT outlook: Tech conferences and events calendar | TechTarget 8 AI use cases in manufacturing Enterprises are making an AI native transformation Zero trust in the IT ops stack: Securing hybrid workloads How algorithmic value sets enhance clinical decision-making Top methods for collecting customer feedback Build a data governance team that delivers results How to calculate the total cost of ownership of ERP software Communities call for transparency in AI data center deals Scalable IT infrastructure: Balancing speed with stability How health systems are tackling 'Kill the Clipboard' obstacles Understanding the science behind AI-based hiring assessments Tape's strategic role in modern data protection How to choose an HR software system in 2026: A complete guide The UC stack gets the policy job Top zero-trust use cases in the enterprise 13 top IT infrastructure conferences in 2026 SNMP vs. CMIP: What's the difference? 3 essential network analytics use cases AI Security Risks Force CIOs to Rethink Strategy Red Hat Summit 2026 news and conference guide | TechTarget What is HR technology (human resources tech)? Understand, optimize and track customer journey touchpoints Should IT use Apple Business Manager without MDM? Build and organize an effective machine learning team The storage modernization imperative in a fast-changing IT landscape Procurement automation use cases for CSCOs to consider 3 steps for health system leaders to drive patient safety culture What is DevOps? Meaning, methodology and guide Enterprises Face New Storage Bottlenecks as AI Grows A guide to Intune Suite licensing for endpoint management Epic controls 42% of the US EHR market. Does that help or hurt interoperability? SAP Sapphire 2026 news, trends and analysis | TechTarget How to develop a data governance strategy: 7 key steps 12 generative AI tools for marketing and sales teams Top 9 smart contract platforms to consider in 2026 Top 8 e-signature software providers for 2026 Rise with SAP vs. S/4HANA Cloud: What are the differences? How businesses use KPIs to measure AI's performance 5 clues your network has shadow AI How do digital signatures work? Collaboration security and governance must be proactive Compare SAP greenfield vs. brownfield approach for S/4HANA Merck, Home Depot tap Gemini Enterprise for AI agent development Rural challenges may dampen digital healthcare's potential Build an ethical AI framework: 12 top resources The great workload reshuffle: Choices for AI and analytics How to remove a device from Intune enrollment Cisco unveils quantum network advancements 3 BYOD security risks and how to prevent them 10 of the top carbon accounting software 8 trends powering machine learning's dynamic new roles Network engineers must take the lead to push DDI to the cloud How does Microsoft 365 Copilot pricing and licensing work? ONC highlights behavioral health EHR adoption trends, data exchange barriers LLMs struggle with clinical reasoning, study finds Democratizing AI in business: The good, bad and ugly What can organizations do to address BYOD privacy concerns? Fix the service path before you optimize it with AI How AI reshapes upselling in customer experience platforms When collaboration starts becoming operational drag Balancing health AI management with growing vendor sprawl Career cure for AI phobia: Be a beekeeper, not a worker bee 16 top applicant tracking systems for 2026 How a rural community hospital deploys AI to detect heart disease 8 examples of document version control Guide to 30+ sustainability certifications for professionals AI agents are only as smart as the data that feeds them AI could earn trust in transactional work first How to fix keyboard connection issues on a remote desktop How to add and enroll devices to Microsoft Intune 11 DevSecOps best practices to prioritize in 2026 6 key components of a successful data strategy How to enable Copilot in Microsoft 365: A step-by-step guide What CIOs need to know about Meta's proposed CEO AI agent Top AI recruiting tools and software of 2026 How contact centers detect and prevent fraud 10 essential skills for modern contact center agents Beyond the chatbot: Engineering the agentic enterprise AI in business intelligence: How to manage it effectively Why legacy networks are a growing liability Failure is an option as an IT leadership tool How HR can create a successful change management strategy HR AI is becoming a change management story Digital transformation: Balancing speed and governance RSAC 2026 Conference: Key news and industry analysis | TechTarget 8 best practices for a bulletproof IAM strategy 5 customer journey phases businesses should understand 12 top HR software and tool options to consider in 2025 6 contact center trends shaping the future of customer service Contact center monitoring best practices for CX leaders Cloud vs. local backup: Which is right for your organization? 6 steps for when remote desktop credentials are not working How governance maturity affects M&A integration outcomes Inside the push to turn AI agents into suite functionality How should contact centers use AI today? Accenture global health lead on scaling AI in healthcare with governance and intent 10 best free DevOps certifications and training courses in 2026 What is compensation management? What CIOs must know about bossware strategy
8 Data Integration Challenges and How to Overcome Them | TechTarget
Donald Farmer · 2026-06-15 · via WhatIs

Every so often, I meet a CIO who tells me their team has consolidated numerous legacy data sources into a single new platform, and its worst data integration challenges will now be a thing of the past. Unfortunately, the reality is quite different.

More often than not, much of the organization's data resides in disparate systems -- on desktops, in file shares, streaming from various devices and external data collected from the web. The need for complicated data integration never went away.

In fact, almost every organization, regardless of size, needs to integrate data from multiple sources to support business processes and get a consistent view of their current state, customer behavior and ongoing operations. It's not an easy task.

The first step in tackling data integration is to understand how the process fits into your overall data management strategy. For example, an ERP system might be governed and administered by a set of data management policies focused on financial integrity. Meanwhile, the CRM system is governed by the need to comply with customer data privacy regulations. What policies will govern a process that integrates both systems and creates a new, consolidated data set for BI and analytics uses?

Data integration is a vital part of data management, but it comes with plenty of technical challenges. The eight described here are common, especially in modern data architectures that are increasingly diverse and dynamic.

1. Managing data volumes while preserving lineage

Most businesses today recognize data as a valuable asset, but many struggle with the sheer scale of what is available to them. Data storage is relatively cheap, and analytics tools are capable of handling large amounts of data, so where does the problem lie?

Data integration and related disciplines, such as managing data quality, can be a real challenge when data volumes become massive. To add to the complication, regulators increasingly demand audit trails for all data used in decision support or for training AI systems. This is the problem of data lineage, where the path data takes through systems and transformations must be traced. Lineage answers questions like: where did this data come from, what happened to it along the way, and where did it go?

Many data integration tasks that are routinely performed with modest data volumes become more taxing as workloads grow very large. But most of the techniques for handling complex integration processes will also help with large volumes of data. In addition, running smaller, more efficient batch integration jobs and optimizing the integration workflow prevent the pipeline from being held up. But the problem of maintaining lineage through these transformations remains.

One approach is to embed lineage into the architecture from the start by running smaller batch jobs and optimizing workflows to maintain metadata. Where lineage becomes too complex to maintain, it may be acceptable to track data provenance instead. Provenance is closely related to lineage but focused more on where this data originated, who created it and whether we can trust it.

2. Integrating diverse data sources

At one time, extract, transform and load (ETL) and other data integration processes mostly involved a mix of text files and database extracts that contained familiar structured data types. Now, you might also need to integrate streaming data from device logs and online services, including:

  • Social media data, including text and images.
  • Public data, such as weather data, commodity prices from government websites, or specialized information providers.
  • In some sectors, increasing amounts of ecosystem data from customers, partners, suppliers or shippers.

This data diversity complicates integration work, but you can manage it with a careful choice of data platforms and tool sets. A traditional data warehouse built on a relational database handles structured diversity reasonably well, but struggles with data streams and unstructured content. A data lake processes streams and unstructured data effectively, but historically offered weaker guarantees around integrity, transaction support and availability.

Recently, hybrid data lakehouse architecture has become a critical component in the data ecosystem to support analytics, real-time and AI scenarios. A lakehouse adds transactional capabilities and schemas over data lake storage, most popularly built on Apache Iceberg. In this way, a single platform can now serve BI workloads that require strict consistency, data science workloads that require access to raw, fine-grained data and AI workloads that may include very diverse data types.

This convergence simplifies governance because rather than maintaining separate policies for warehouse and lake environments, organizations can enforce access controls, lineage tracking and quality metrics through a unified layer.

3. Hybrid cloud and on-premises environments

Not all large-scale IT workloads have moved to the cloud, and some that did have been repatriated to on-prem storage. Several factors drive this, including the following:

  • Cloud systems are flexible and easy to scale up or down, but they aren't cheap. Many enterprises have been surprised by the cost of cloud computing -- they find that the elastic nature of cloud services and the simplicity of scaling lead to inefficient practices that result in usage costs being far more than expected.
  • There are increasing regulatory demands for data sovereignty, which refers not only to the physical storage of data, but also to AI models trained over it and who (or what systems) can access those assets. This trend is driven by privacy concerns, but also by worries about the transfer of AI technologies across national boundaries.
  • Organizations training their own AI models are often concerned about data exposure, especially when using cloud-hosted services. They may prefer to run smaller models on premises to keep proprietary data and trade secrets under their control.

4. Poor or inconsistent data quality

I often say there's only one measurement of data quality: Is it fit for purpose? But the same data can be used for very different purposes, which makes assessing its quality in that way tricky. Also, key dimensions of data quality, such as accuracy, timeliness, completeness and especially consistency, become more difficult to maintain as data volumes grow and data comes from diverse sources.

That could mean business decisions are based on bad, incomplete or duplicate data. To prevent that, organizations must identify and address data quality issues during data integration, through steps such as data profiling and data cleansing. In a data lakehouse, however, it's advised not to run destructive data integration processes that overwrite or discard the original data, which may be of analytical value to data scientists and other users as is. Rather, ensure the raw data is still available in a separate architectural zone and document quality metrics for each zone so that downstream consumers understand what they are working with.

5. Serving multiple use cases without sacrificing compliance

It's not only data quality that is affected by multiple use cases -- data integration is too. Data scientists often want to build analytical models with fine-grained raw data so they don't miss potential insights, but business analysts more often want to work with aggregated data that's aligned with their best practices. Similarly, data may need to be in different formats to be consumed by different tools. For example, data science tools might optimally use Apache Parquet files, but BI dashboards need to be built through a database connection.

AI workloads demand specific data types and access patterns, including large volumes of labeled data for training, low-latency feature data for inference, and vector embeddings. For these AI workloads, tracking data lineage is crucial to address regulatory constraints on model learning and to ensure data quality. These demands introduce new integration points, in addition to traditional analytics and BI needs, often requiring distinct pipelines for data transformation and delivery.

The secret is not only to understand these different use cases, but also to accommodate them. Don't try to force users to change data platforms or analytics tools or to use a suboptimal connection just to make data integration easier. Instead, look for a data integration platform that supports a wide variety of targets and don't be reluctant to build separate data pipelines for specific use cases.

6. Monitoring and observability

In a data integration context, it's easy to mistake data observability for a more traditional approach that involves careful logging and reporting of integration processes. However, there's more to it than that.

Modern data integration processes -- diverse, distributed and dynamic as they are -- also have numerous points of failure. Very often, data pipelines and the dependencies between them are designed so that no single failure breaks the whole process. This robustness reduces downtime, which used to be the bane of overnight ETL processes that broke too easily and were difficult to restart.

However, if the process as a whole doesn't break but one part of it fails, what is the current state of your data? Monitoring can be a challenge.

This is where data observability comes in. It measures data delivery -- including when the data was last fully processed-- logs all processes and traces both the source and impact of any errors. When you run a report, you can see not only when the data was refreshed, but also which components -- even which rows or cells -- may not be fully up to date or of the best quality.

Observability typically focuses on the following five key attributes of data:

  • Is it timely?
  • Is it structured as we expect?
  • Is it within the expected data quality limits?
  • Is the data complete?
  • What is the lineage of the data -- where was it sourced from?

In a perfect world, all of our data would be right all the time. In the real world, pay attention to data observability and, if needed, invest in specific tools for this work.

7. Integrating streaming data with event-driven architectures

A stream of data is continuous and unbound -- there's no beginning or end, just a sequence of recorded events. And therein lies the challenge: How do you integrate this open-ended data with more traditional architectures that expect data sets, not streams?

With many devices or streaming services, it's possible to request data in a batch. You could get all the data for the last five minutes, for example. What you receive is a file, structured much like an XML record or a database table. You could then get another one five minutes later, building up a continuous flow of data. This isn't really streaming -- in fact, many call this micro-batching because you're receiving numerous small, discrete batches or data sets.

Likewise, there are integration tools, such as change data capture software, that enable you to query a stream as if it were a batch, often using SQL. This makes data integration pretty straightforward, and I recommend you start here. Handling streams that are interrupted and then restarted remains a challenge. But your chosen tool may offer best practices or specific features that make this catch-up effort, sometimes called rehydrating the stream, easier.

Nevertheless, because you're using a workaround to integrate a stream, data observability is again essential to identify and respond to any issues that arise.

8. Mixed tool sets and architecture

To combat the previous challenges, I've recommended various tools and platforms. In doing so, I've now created another challenge: handling all these tools and the complex architecture that can result from using them.

The good news is that, except in the most complex scenarios, you won't be trying to solve every challenge listed. Also, if your systems are running on a cloud platform from one of the mega-vendors, versions of many of these tools will likely be available in some form. They might not always be best-in-class, but they're a good place to start -- the integration and consistency between them is helpful when starting out.

On the other hand, your choice of a data integration platform might be driven by your need for a very specific set of tools. For example, some data engineers -- particularly in manufacturing companies -- find that only one streaming tool really meets their needs. The platform that supports it best wins the day.

Whatever tools you select, governance should operate at a layer above them. One path forward may be metadata-driven automation: define governance policies as metadata that travels with the data, enforced automatically by whatever tool processes it. In this way, access controls, quality thresholds, retention policies, and lineage requirements are not tied to which tool happens to touch the data; they follow the data itself.

Best practices for managing data integration processes

There's a lot to consider with these challenges, and you'll likely come across more. The following are some best practices that can help make data integration strategies and efforts successful:

  • Process data as close to the source as possible, both to minimize data movement and to remove or select out unneeded data for efficiency as soon as possible.
  • Document integration processes and catalog the integrated data carefully, so business users and data scientists alike can find what they're looking for. As data integration grows in complexity, your documentation should grow in precision -- and volume, unfortunately. But look at that as an investment in future ease of use, data recovery and high availability, rather than as a burden.
  • Treat lineage as a first-class requirement, especially when your data feeds into AI systems. Retrofitting lineage onto pipelines built without it is costly, so build it in from the start.
  • Similarly, if your data is used to train or fine-tune AI models, maintain a strict versioning discipline. For both debugging models and regulatory audits, you should be able to reconstruct the exact data set used to train a model at a particular point in time.
  • Keep your data integration processing nondestructive in analytical data sets. You never know when users will need to get back to the original data -- also, use cases vary greatly and change over time.
  • With that in mind, you may find a data lakehouse architecture most appropriate, but still include a landing zone for raw data, a staging area for temporary data, and zones where you can save integrated, conformed and cleansed data to meet the requirements of data science, BI and AI use cases.
  • In general, focus on data integration techniques rather than tools, especially at the design or prototyping stage. Integration tools are helpful, but they can constrain what you build around their own capabilities. In practice, many data engineers or integration developers address challenges such as those involving a lot of code and a limited set of tools. That's not for everyone, of course, but a lot of example code and architectural advice is available from people who have solved similar integration problems already.
  • Don't just rely on logging to track and monitor integration processes. Use data observability techniques so the availability and quality of data are documented and well understood by users and data administrators alike.
  • Recognize that AI pipelines may fail differently from traditional ETL. A batch job either completes or throws an error, but stale or outdated data, or data with emerging biases, may continue running silently and degrade. Monitoring AI integration points requires attention not only to technical execution, but to data freshness and the quality of inferences or predictions.

Donald Farmer is a data strategist with 30+ years of experience, including as a product team leader at Microsoft and Qlik. He advises global clients on data, analytics, AI and innovation strategy, with expertise spanning from tech giants to startups. He lives in an experimental woodland home near Seattle.

Next Steps

Emerging data integration trends to assess

Why culture is critical for data integration

Establish big data integration techniques and best practices

Dig Deeper on Data integration