惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
罗磊的独立博客
人人都是产品经理
人人都是产品经理
博客园_首页
Hugging Face - Blog
Hugging Face - Blog
美团技术团队
L
Lohrmann on Cybersecurity
博客园 - 【当耐特】
量子位
Last Week in AI
Last Week in AI
D
Darknet – Hacking Tools, Hacker News & Cyber Security
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
C
Cyber Attacks, Cyber Crime and Cyber Security
腾讯CDC
有赞技术团队
有赞技术团队
Cyberwarzone
Cyberwarzone
T
Tor Project blog
V
V2EX
L
LINUX DO - 热门话题
Security Latest
Security Latest
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
NISL@THU
NISL@THU
C
Cisco Blogs
T
Tailwind CSS Blog
G
GRAHAM CLULEY
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - Franky
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
小众软件
小众软件
K
Kaspersky official blog
博客园 - 司徒正美
IT之家
IT之家
大猫的无限游戏
大猫的无限游戏
Jina AI
Jina AI
S
Schneier on Security
月光博客
月光博客
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
The Exploit Database - CXSecurity.com
Scott Helme
Scott Helme
J
Java Code Geeks
博客园 - 聂微东
Martin Fowler
Martin Fowler
MongoDB | Blog
MongoDB | Blog
AWS News Blog
AWS News Blog
Know Your Adversary
Know Your Adversary
C
Cybersecurity and Infrastructure Security Agency CISA
F
Fortinet All Blogs
T
Threat Research - Cisco Blogs
C
CXSECURITY Database RSS Feed - CXSecurity.com
雷峰网
雷峰网

Catchpoint Blog

SRE Report: AI optimism and the economics of effort SRE Report: Why fast is what users trust SRE Report 2026: What surprised us, what didn't, and why the gaps matter most The SRE Report 2026: Defensible Ns Why Synthetic Tracing Delivers Better Data, Not Just More Data A New Chapter: LogicMonitor + Catchpoint – A Personal Note from Mehdi Mezmo + Catchpoint deliver observability SREs can rely on The four pillars holding up your digital business, and what happens when they crumble When payments pause: lessons from a global payments outage Observability 2025 Decoded: What the DZone Report Means for SLO-Driven Ops The next evolution of WebPageTest has arrived, and it’s a game-changer The Monitoring Blind Spot That Could Cost You Black Friday Powering Mexico’s Digital Future: Expanded Internet Observability with Catchpoint The Next Chapter of WebPageTest: Your New Experience Starts Soon SRE Report Retrospectives — Have AIOps Predictions Held Up? When BGP becomes UX: The inside story of a SaaS routing decision gone wrong (or right) Session Replay explained: A guide to seeing digital experience through your user’s eyes Making the invisible visible: Are your cloud firewalls and DDoS protection really working? Why it’s time to move beyond APM: Monitoring from the user’s perspective When metrics mislead: Inside the 2025 Retail Web Performance Benchmark The vendor trap: why your next outage won’t be your fault—but will be your problem LLMs don’t stand still: How to monitor and trust the models powering your AI Semantic Caching: What We Measured, Why It Matters The Annual SRE Survey Is Open—We Want to Hear from You Observability isn’t about the tool. It’s about the truth Invisible dependencies, visible impact: Lessons from the Google Cloud outage Real-time detection of BGP blackholing and prefix hijacks Leading analyst firm reveals the real cost of internet disruptions The Power of Over 3000 Intelligent Observability Agents Monitoring in the Age of Complexity: 5 Assumptions CIOs Need to Rethink Why Intelligent Traffic Steering is Critical for Performance and Cost Optimization Retail digital performance event recap: Key insights from IBM & Catchpoint Zendesk outage: A case for proactive monitoring and faster incident response Silence during chaos: Why the X outage is a call to arms for proactive monitoring The $1 Million Lesson: Building a Culture of Quality Through SLAs When AI tools fail: How to map your AI dependencies for proactive visibility Why Super Bowl 2025 was a triumph for Internet Resilience Why Internet Performance Monitoring is the new health check for IT organizations Why use Playwright in Catchpoint for synthetic monitoring Introducing WebPageTest Expert Plan: Real-Time Insights, Synthetic + RUM together in One Platform The shift to digital: How businesses are reshaping their priorities for 2025 The SRE Report 2025's Call to Action Monitoring in the Age of the Internet: DEM, IPM, and APM—What You Need to Know SSL Monitoring, Trust, and McLOVIN Performing for the holidays: Look beyond uptime for season sales success Lessons from Microsoft’s office 365 Outage: The Importance of third-party monitoring Web Performance Experts Look into the Future of Web Performance The hidden challenges of Internet Resilience: Key insights from 2024 report When SSL Issues aren’t just about SSL: A deep dive into the TIBCO Mashery outage The curious case of Marriott and the untold impact of web performance on revenue Preparing for the unexpected: Lessons from the AJIO and Jio Outage It’s time to stop neglecting the elephant in the room: Performance Matters! The Need for Speed: Highlights from IBM and Catchpoint’s Global DNS Performance Study Learnings from ServiceNow’s Proactive Response to a Network Breakdown Webinar Recap: Taking Web Performance to the Next Level Use the Catchpoint Terraform Provider in your CI/CD workflows Is the Internet ready for L4S? Takeaways from the CrowdStrike outage: third-parties can pose risk July 19th global IT outage reminds us of digital complexity 5 Actions you can take to improve digital performance 2024: A banner year for Internet Resilience APM vs Observability: Both-and, not either-or AppAssure: Ensuring the resilience of your Tier-1 applications just became easier APM vs observability: why your definitions are broken APM vs Observability: What comes next? APM vs Observability: Observing beyond APM AWS Outage: How do you prepare for the failure of your own safety net? Agentic AI: Powerful But Fragile—What You Need to Know Catch frustration before it costs you: New tools for a better user experience Catchpoint Expands Observability Network to Barcelona: A Growing Internet Hub Catchpoint Peak Performance Summit 2025: Redefining Observability for the Outcome Economy Catchpoint named a leader in the 2024 Gartner® Magic Quadrant™ for Digital Experience Monitoring Consolidation and Modernization in Enterprise Observability Connected Devices: Unlocking the next frontier of Internet Performance Monitoring Cloud Monitoring's Blind Spot: The User Perspective Cloudflare’s Resolver Outage: More Than Just DNS Cloudflare outage: another wake-up call for resilience planning Demystifying API Monitoring and Testing with IPM Creating the IPM Category: Catchpoint’s Journey to Leadership and the LogicMonitor Era Critical Requirements for Modern API Monitoring Customer Survey 2024: Unveiling insights and impact Did Delta's slow web performance signal trouble before CrowdStrike? Diagnosing Wi-Fi failures that traditional tools miss: a case study DNS misconfiguration can happen to anyone - the question is how fast can you detect it? ECN explained: Navigate congestion for faster, smoother data delivery Don’t get caught in the dark: Lessons from a Lumen & AWS micro-outage Escalating risk, shrinking margins: The 2025 Internet Resilience Report From refresh to results: the metrics that shaped Election Day 2024 coverage Fast and furious: The importance of performance in the digital age Getting Started with Traceroute From the source to the edge: the six agent types you can’t ignore From SEO to AEO: Why Web Performance Is the Key to AI Search Success Going for gold: Testing the resilience of Olympic websites Here’s the proof: What the fastest sites on the web have in common Google’s Agent-to-Agent (A2A) Protocol is here—Now Let’s Make it Observable How IPM helped a top tech brand catch an OpenAI outage before it became a crisis How AI Turns Monitoring From “What Now?” Into “What’s Next?” How SAP achieved world-class uptime through modern observability How to Monitor AI Agents in Commerce Systems
Achieving stability with agility in your CI/CD pipeline
2026-05-31 · via Catchpoint Blog

in this blog post

After reading this blog and watching our recent webinar on the same topic, it’s our hope you will be in a position to build better software, collaborate more effectively across the software development lifecycle, and understand how monitoring your entire DevOps lifecycle using Internet Performance Monitoring (IPM) can aid you in doing so.  

Who does this apply to? DevOps and AppDev practitioners, in particular those at more mature companies.  

Let’s face it. At an early startup, everyone does everything. You must be agile to succeed. As a result, you don’t necessarily need to be stable. You’re trying to acquire customers and that requires experimentation. But those companies grow up. You have customers and you can’t fail them. You can’t go down, you can’t afford to have major features break or have bad data, but nor can you stop innovating. That’s who we’re aiming this at.

When we talk about how to achieve stability with agility, essentially that’s the crux: how to keep innovating and acquiring new customers while at the same time keeping your existing customers happy.

Something we all strive for, but it isn’t always easy.

Using IPM to achieve stability with agility in CI/CD

Admittedly, my role as VP of Engineering (Sergey here) involves being responsible for Catchpoint’s IPM platform, so of course I am biased, but at the same time, Catchpoint very much falls into the mature category I have just described. We have a large customer base whom we can’t let down. We have a large infrastructure that has thousands of components and many, many dependencies (we’ll get more into that momentarily). Nonetheless, one of my jobs is to keep innovating and delivering new functionality and features. Here at Catchpoint, we have to be both agile and stable.  

5 IPM-centered CI/CD tips

I’d like to share a few tips from my own hard-won experience on how we strive to achieve this delicate balance - with IPM as an anchor tool:

#1 - Set up the observability you need to achieve stability

There have been times when I made the wrong decision. Times I chose agility over stability. It can be a tough choice.  

As engineering leaders, we always want to move as fast as possible. We want to minimize toil for our teams, and we want them to work on the most interesting projects for the highest reward. But moving too fast without the right systems set up to observe what’s going on inside our applications and platforms can cause massive issues.

The way to reduce toil is not just to do CI/CD, but to do it properly, you need to understand – and have a way to accurately measure - your whole system.  

#2 - Analyze ALL your dependencies (internal and external)

In our recent Internet Resilience Report, 77% of respondents said that third-party dependencies were extremely or highly critical to their Internet Resilience success.

The different teams involved in maintaining or developing your product are likely to be analyzing your internal dependencies, such as your databases, servers, storage, etc., but they’re less likely to be assessing the external dependencies – DNS, CDN, email services, etc.  

Additionally, almost no one thinks about the customer side. What if all your customers are in a single location and there’s a major ISP outage in that region, meaning none of your customers can reach your site.  

The question to repeatedly ask yourself: is your stability affected when any of these dependencies are impacted?

If so, you need to find a way to incorporate your third-party dependencies into your CI/CD playbooks and runbooks. IPM covers the entire Internet Stack. IPM will help you implement resilience where it’s most important. It will help you understand your dependencies for both your internal customers (your employees) and your external employees (your customers) so they are properly evaluated, considered, prioritized and monitored.  

A diagram of different types of softwareDescription automatically generated

#3 - Shift wide (i.e. left and right)

What does shift wide mean? Well, to back up for a moment… Shift left is the concept that by moving things earlier in the development process, you find problems faster and therefore it costs less.  

However, that’s not always the case; for example, for trivial changes such as changing the color of a button, it’s usually OK to test during production i.e. to shift right (more on that here).  

To shift wide (and encompass both), you need a ground truth i.e. a shared set of standardized measurements to be able to react quickly and fix issues if and where there’s a problem… which is where IPM comes in.

#4 - Use IPM as a common language at the center of your CI/CD pipeline

Now, let’s get down to the CI/CD pipeline in more detail.  

A diagram of a software development processDescription automatically generated

You can look across all its different phases and notice that ‘Monitor’ is right in the middle.  

Indeed, IPM specifically is a crucial stage of the CI/CD pipeline because it creates a ground truth or a common language that all the different teams responsible for different stages of your software lifecycle can use to communicate.

Your product management team, for example, can use your IPM data to specify what the performance of the system needs to be. They can then work with your DevOps or Ops team to create a monitoring strategy – before your software exists. Your developers can then use the same data to make sure their code works properly. The testers can use it for their system testing, and by the time it gets to operations, everyone is sure they’re on the same page and the entire system is working as intended.  

I call this the observable software development lifecycle (OSDLC).  

#5 - “As code” automation and integrations to reduce toil and save OpEx

Lastly, let’s talk about Observability as Code.  

After all, this is CI/CD, everything has to be automated into your workflow in order to truly be of benefit. Not only do we need automation, but it needs to be seamless and simple. In order to achieve that, you have to treat the observability you are implementing as part of your OSDLC as code.  

This way, you reduce toil for your teams, particularly when working at scale and potentially save on OpEx, removing the need for time and resources to be spent on manual, repetitive processes.

With Catchpoint, we have of course a number of ways to incorporate IPM into your CI/CD pipeline. Perhaps the most mature way I’ve seen customers run Observability as Code is to create new test configurations as new features are released using the REST API.  

This way, you tag deployments, run ad hoc testing to ensure things are working, trigger actions or other automations using Webhooks or emails, then send all your collected data to a data warehouse to combine it with other data sources to ultimately further your designated outcomes, such as business KPI reporting.  

It’s not if, but when

It all comes down to a simple question. It’s not a question of if an incident occurs, but when. What are you prepared for the impact to be?  

In the Internet Resilience Report we mentioned earlier, we found that 43% of companies estimate a total economic impact or loss of more than $1 million monthly due to Internet outages or degradations. Get your entire teams speaking the same language with IPM across the OSDLC and begin to get ahead of those outages before they impact your business.  

A graph of percentages and numbersDescription automatically generated with medium confidence

Internet Resilience Report 2024 (Catchpoint)

To find out more about how Catchpoint’s IPM platform can help you, talk to a Solution Engineer.

Further resources

Watch the related webinar, which goes into these concepts in more depth (including a sample CI/CD workflow): https://www.catchpoint.com/webinar/how-to-achieve-agility-with-stability

Find out how Catchpoint can help you across the DevOps lifecycle: https://www.catchpoint.com/application-experience/devops-lifecycle

Summary

After reading this blog and watching our recent webinar on the same topic, it’s our hope you will be in a position to build better software, collaborate more effectively across the software development lifecycle, and understand how monitoring your entire DevOps lifecycle using Internet Performance Monitoring (IPM) can aid you in doing so.  

Who does this apply to? DevOps and AppDev practitioners, in particular those at more mature companies.  

Let’s face it. At an early startup, everyone does everything. You must be agile to succeed. As a result, you don’t necessarily need to be stable. You’re trying to acquire customers and that requires experimentation. But those companies grow up. You have customers and you can’t fail them. You can’t go down, you can’t afford to have major features break or have bad data, but nor can you stop innovating. That’s who we’re aiming this at.

When we talk about how to achieve stability with agility, essentially that’s the crux: how to keep innovating and acquiring new customers while at the same time keeping your existing customers happy.

Something we all strive for, but it isn’t always easy.

Using IPM to achieve stability with agility in CI/CD

Admittedly, my role as VP of Engineering (Sergey here) involves being responsible for Catchpoint’s IPM platform, so of course I am biased, but at the same time, Catchpoint very much falls into the mature category I have just described. We have a large customer base whom we can’t let down. We have a large infrastructure that has thousands of components and many, many dependencies (we’ll get more into that momentarily). Nonetheless, one of my jobs is to keep innovating and delivering new functionality and features. Here at Catchpoint, we have to be both agile and stable.  

5 IPM-centered CI/CD tips

I’d like to share a few tips from my own hard-won experience on how we strive to achieve this delicate balance - with IPM as an anchor tool:

#1 - Set up the observability you need to achieve stability

There have been times when I made the wrong decision. Times I chose agility over stability. It can be a tough choice.  

As engineering leaders, we always want to move as fast as possible. We want to minimize toil for our teams, and we want them to work on the most interesting projects for the highest reward. But moving too fast without the right systems set up to observe what’s going on inside our applications and platforms can cause massive issues.

The way to reduce toil is not just to do CI/CD, but to do it properly, you need to understand – and have a way to accurately measure - your whole system.  

#2 - Analyze ALL your dependencies (internal and external)

In our recent Internet Resilience Report, 77% of respondents said that third-party dependencies were extremely or highly critical to their Internet Resilience success.

The different teams involved in maintaining or developing your product are likely to be analyzing your internal dependencies, such as your databases, servers, storage, etc., but they’re less likely to be assessing the external dependencies – DNS, CDN, email services, etc.  

Additionally, almost no one thinks about the customer side. What if all your customers are in a single location and there’s a major ISP outage in that region, meaning none of your customers can reach your site.  

The question to repeatedly ask yourself: is your stability affected when any of these dependencies are impacted?

If so, you need to find a way to incorporate your third-party dependencies into your CI/CD playbooks and runbooks. IPM covers the entire Internet Stack. IPM will help you implement resilience where it’s most important. It will help you understand your dependencies for both your internal customers (your employees) and your external employees (your customers) so they are properly evaluated, considered, prioritized and monitored.  

A diagram of different types of softwareDescription automatically generated

#3 - Shift wide (i.e. left and right)

What does shift wide mean? Well, to back up for a moment… Shift left is the concept that by moving things earlier in the development process, you find problems faster and therefore it costs less.  

However, that’s not always the case; for example, for trivial changes such as changing the color of a button, it’s usually OK to test during production i.e. to shift right (more on that here).  

To shift wide (and encompass both), you need a ground truth i.e. a shared set of standardized measurements to be able to react quickly and fix issues if and where there’s a problem… which is where IPM comes in.

#4 - Use IPM as a common language at the center of your CI/CD pipeline

Now, let’s get down to the CI/CD pipeline in more detail.  

A diagram of a software development processDescription automatically generated

You can look across all its different phases and notice that ‘Monitor’ is right in the middle.  

Indeed, IPM specifically is a crucial stage of the CI/CD pipeline because it creates a ground truth or a common language that all the different teams responsible for different stages of your software lifecycle can use to communicate.

Your product management team, for example, can use your IPM data to specify what the performance of the system needs to be. They can then work with your DevOps or Ops team to create a monitoring strategy – before your software exists. Your developers can then use the same data to make sure their code works properly. The testers can use it for their system testing, and by the time it gets to operations, everyone is sure they’re on the same page and the entire system is working as intended.  

I call this the observable software development lifecycle (OSDLC).  

#5 - “As code” automation and integrations to reduce toil and save OpEx

Lastly, let’s talk about Observability as Code.  

After all, this is CI/CD, everything has to be automated into your workflow in order to truly be of benefit. Not only do we need automation, but it needs to be seamless and simple. In order to achieve that, you have to treat the observability you are implementing as part of your OSDLC as code.  

This way, you reduce toil for your teams, particularly when working at scale and potentially save on OpEx, removing the need for time and resources to be spent on manual, repetitive processes.

With Catchpoint, we have of course a number of ways to incorporate IPM into your CI/CD pipeline. Perhaps the most mature way I’ve seen customers run Observability as Code is to create new test configurations as new features are released using the REST API.  

This way, you tag deployments, run ad hoc testing to ensure things are working, trigger actions or other automations using Webhooks or emails, then send all your collected data to a data warehouse to combine it with other data sources to ultimately further your designated outcomes, such as business KPI reporting.  

It’s not if, but when

It all comes down to a simple question. It’s not a question of if an incident occurs, but when. What are you prepared for the impact to be?  

In the Internet Resilience Report we mentioned earlier, we found that 43% of companies estimate a total economic impact or loss of more than $1 million monthly due to Internet outages or degradations. Get your entire teams speaking the same language with IPM across the OSDLC and begin to get ahead of those outages before they impact your business.  

A graph of percentages and numbersDescription automatically generated with medium confidence

Internet Resilience Report 2024 (Catchpoint)

To find out more about how Catchpoint’s IPM platform can help you, talk to a Solution Engineer.

Further resources

Watch the related webinar, which goes into these concepts in more depth (including a sample CI/CD workflow): https://www.catchpoint.com/webinar/how-to-achieve-agility-with-stability

Find out how Catchpoint can help you across the DevOps lifecycle: https://www.catchpoint.com/application-experience/devops-lifecycle

This is some text inside of a div block.