惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
阮一峰的网络日志
阮一峰的网络日志
有赞技术团队
有赞技术团队
小众软件
小众软件
P
Proofpoint News Feed
P
Proofpoint News Feed
Apple Machine Learning Research
Apple Machine Learning Research
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
The Last Watchdog
The Last Watchdog
O
OpenAI News
Security Latest
Security Latest
博客园 - Franky
Forbes - Security
Forbes - Security
N
Netflix TechBlog - Medium
H
Hacker News: Front Page
Cloudbric
Cloudbric
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Hugging Face - Blog
Hugging Face - Blog
Microsoft Security Blog
Microsoft Security Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
S
Security Affairs
Recent Announcements
Recent Announcements
The GitHub Blog
The GitHub Blog
S
Schneier on Security
MongoDB | Blog
MongoDB | Blog
WordPress大学
WordPress大学
Last Week in AI
Last Week in AI
博客园 - 【当耐特】
Attack and Defense Labs
Attack and Defense Labs
C
Cyber Attacks, Cyber Crime and Cyber Security
F
Fortinet All Blogs
Webroot Blog
Webroot Blog
S
Secure Thoughts
Spread Privacy
Spread Privacy
Blog — PlanetScale
Blog — PlanetScale
T
Troy Hunt's Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
V
V2EX
Security Archives - TechRepublic
Security Archives - TechRepublic
P
Privacy & Cybersecurity Law Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Simon Willison's Weblog
Simon Willison's Weblog
C
Check Point Blog
L
LINUX DO - 最新话题
NISL@THU
NISL@THU
博客园_首页
罗磊的独立博客
A
Arctic Wolf
U
Unit 42

VictoriaMetrics: Simple & Reliable Monitoring for Everyone on VictoriaMetrics

Operator now has Long-Term Support (LTS) version Multi-tiered Observability: A Practical Way to Handle Diverse Workloads VictoriaMetrics April 2026 Ecosystem Updates Not All Telemetry Requires Premium Pricing VictoriaMetrics at KubeCon Amsterdam: Community Highlights What's new in VictoriaMetrics Anomaly Detection (Q1 2026) What's New in VictoriaMetrics Cloud Q1 2026? Logs, MCP Server, Better Alerting, and... a Secret Project VictoriaMetrics at KubeCon: Optimizing Tail Sampling in OpenTelemetry with Retroactive Sampling VictoriaMetrics March 2026 Ecosystem Updates Observability Lessons From OpenAI Benchmarking Kubernetes Log Collectors: vlagent, Vector, Fluent Bit, OpenTelemetry Collector, and more VictoriaMetrics February 2026 Ecosystem Updates VictoriaMetrics at FOSDEM, Cloud Native Days France, and CfgMgmtCamp Ghent VictoriaLogs in VictoriaMetrics Cloud: Fast, Cost-Effective Log Management is Here What’s new in VictoriaMetrics Anomaly Detection (2025) VictoriaMetrics January 2026 Ecosystem Updates VictoriaLogs Basics: What You Need to Know, with Examples & Visuals What's New in VictoriaMetrics Cloud Q4 2025? New tiers, more deployment options, IaC and alerting rules. Vibe coding tools observability with VictoriaMetrics Stack and OpenTelemetry How a US Software Provider Improved Traffic Alerting with VictoriaMetrics Anomaly Detection VictoriaMetrics 2025 Developer Experience: A Year in Review Spotify’s performance & control across large monitoring environments with VictoriaMetrics VictoriaMetrics Achieves Red Hat OpenShift Operator Certification Our latest updates across the VictoriaMetrics Observability ecosystem New Capacity Tiers in VictoriaMetrics Cloud Announcing 1B+ Downloads & Product Development With Logs, Traces, Metrics AI Agents Observability with OpenTelemetry and the VictoriaMetrics Stack Discarding gRPC-Go: The Story Behind OTLP/gRPC Support in VictoriaTraces What's New in VictoriaMetrics Cloud Q3 2025? From new region in Asia to proactive alerts How DreamHost Slashed Memory Usage by 80% and Scaled to 76 Million Time Series Upcoming Conferences & Meetups: Where to Meet Our Team VictoriaMetrics Long-Term Support (LTS): H2 2025 Update Creating a Sustainable Open Source Business Model - Introduction Full-Stack Observability with VictoriaMetrics in the OTel Demo Alerting Best Practices vmanomaly Deep Dive: Smarter Alerting with AI (Tech Talk Companion) VictoriaLogs Practical Ingestion Guide for Message, Time and Streams Monotonic and Wall Clock Time in the Go time package Hello Singapore! VictoriaMetrics Cloud Expands to Asia Pacific MCP Server Integration & Much More: What's New in VictoriaMetrics Cloud Q2 2025 FIPS 140-3 Compatible Builds for VictoriaMetrics Enterprise Components VictoriaLogs Unleashed: Cluster Version Now Available for Exceptional, Linear Scaling Integrations made easy with VictoriaMetrics Cloud Developer's Note: Research on Distributed Tracing, Comparing With Tempo and ClickHouse vmagent: Key Features Explained in Under 15 Minutes Go synctest: Solving Flaky Tests vmalert: Maximize Your Monitoring (Tech Talk Companion) Celebrating 14K Stars on GitHub: Spring Update vmalert: Maximize Your Monitoring VictoriaMetrics Connects with the Open Source Community at LinuxFest Northwest 2025 Graceful Shutdown in Go: Practical Patterns VictoriaLogs: Gaps, Gains & Growth Prometheus Monitoring: Functions, Subqueries, Operators, and Modifiers VictoriaMetrics Cloud: What's New in Q1 2025? Don’t default to microservices: You’ll thank us later! Container CPU Requests & Limits Explained with GOMAXPROCS Tuning gRPC in Go: Streaming RPCs, Interceptors, and Metadata From Chaos to Clarity with VictoriaLogs Prometheus Alerting 101: Rules, Recording Rules, and Alertmanager Heading to London: Meet Our Team at KubeCon Europe 2025 Inside vmselect: The Query Processing Engine of VictoriaMetrics Meet Our Team at Scale 22x Practical Protobuf - From Basic to Best Practices VictoriaLogs Status Update: Heading Towards the Cluster Version 24th of February 2025 Statement: VictoriaMetrics Stands with Ukraine! Prometheus Metrics Explained: Counters, Gauges, Histograms & Summaries Prometheus Monitoring: Instant Queries and Range Queries Explained 300%+ Growth in 2024: Join Our Team in 2025! FOSDEM 2025 recap How Protobuf Works—The Art of Data Encoding OpenTelemetry, Prometheus, and More: Which Is Better for Metrics Collection and Propagation? How vmstorage Handles Query Requests From vmselect How vmstorage's IndexDB Works VictoriaMetrics Tech Talk Stream: A Deep Dive into Blackbox Monitoring How HTTP/2 Works and How to Enable It in Go VictoriaMetrics Cloud: What's New in Q4 2024? How vmstorage Processes Data: Retention, Merging, Deduplication,... How vmstorage Handles Data Ingestion From vminsert When Metrics Meet vminsert: A Data-Delivery Story From net/rpc to gRPC in Go Applications Piros | VictoriaMetrics Partner Allenta | VictoriaMetrics Partner CloudRaft | VictoriaMetrics Partner Sensedia & VictoriaMetrics: API-compatible Efficient Storage Scalable Prometheus: Why DSV Chose VictoriaMetrics Sensor Factory | VictoriaMetrics Partner Erythix | VictoriaMetrics Partner Groove X & VictoriaMetrics: Faster Device Health Monitoring Scaled & Performant Monitoring at Spotify with VictoriaMetrics Grammarly & VictoriaMetrics: 10× Lower Costs & Direct Access Zelarsoft | VictoriaMetrics Partner DFKI & VictoriaMetrics: Efficient Long-Term Metric Storage Niubits | VictoriaMetrics Partner Megazone Cloud | VictoriaMetrics Partner Cogito Software | VictoriaMetrics Partner Bajau | VictoriaMetrics Partner Find Out Why Dig Security Chose VictoriaMetrics! Ness | VictoriaMetrics Partner Alpha Data | VictoriaMetrics Partner SIOS Technology | VictoriaMetrics Partner
Rules backfilling via vmalert
Roman Khavronenko · 2023-01-31 · via VictoriaMetrics: Simple & Reliable Monitoring for Everyone on VictoriaMetrics

Recording rules is a clever concept introduced by Prometheus for storing results of query expressions in a form of a new time series. It is similar to materialized view and helps to speed up queries by using data pre-computed in advance instead of doing all the hard work on query time.

Like materialized views, recording rules are extremely useful when user knows exactly what needs to be pre-computed. For example, a complex panel on Grafana dashboard or SLO objective. Both have queries that rarely change, so recording rules could significantly simplify and speed up the execution.

But recording rules do not have a retroactive effect. Pre-computed results start to appear only after the moment recording rule was configured. Data before that time will be missing. And results of the recording rule can’t be changed in the past, only deleted or replaced.

Starting from v1.61.0, VictoriaMetrics gained a feature named replay as a part vmalert component. It allows running vmalert in a special mode to retroactively evaluate recording or alerting rules and backfill their results back to the database. Let’s see how this feature can be used in practice.

SLI/SLO calculation

#

SLI/SLO calculation is one of the best examples of using recording and alerting rules. The common practice is to measure SLO on big time windows, of at least 30d. For example, the service shouldn’t return more than 1% of errors on a 30d interval. And since SLI metrics are usually generated via recording rules, users need to wait for 30d until rules evaluate and SLO becomes meaningful. Let’s see how we can make it better.

I have 6 months of metrics collected from one of the VictoriaMetrics sandbox clusters. As an SLO I’d like to define 99.9% of successful requests served by VictoriaMetrics cluster on 30 days interval. To generate rules for this objective, I decided to use one of the most popular SLI/SLO frameworks slok/sloth with the following config:

version: "prometheus/v1"
service: "sandbox-vmcluster"
slos:
 # We allow failing 1 request every 1000 requests (99.9%).
 - name: "requests-availability"
   objective: 99.9
   description: "SLO based on availability for HTTP request responses."
   sli:
     events:
       error_query: sum(rate(vm_http_request_errors_total{job="vmselect-benchmark-vm-cluster"}[{{.window}}]))
       total_query: sum(rate(vm_http_requests_total{job="vmselect-benchmark-vm-cluster"}[{{.window}}]))
   alerting:
     name: VMHighErrorRate

To generate recording and alerting rules from this config, I ran the following command:

./sloth -i slo_config.yml > slo_rules.yml

The result of the command execution is a bunch of recording and alerting rules of the following form:

groups:
- name: sloth-slo-sli-recordings-sandbox-vmcluster-requests-availability
  rules:
  - record: slo:sli_error:ratio_rate5m
    expr: |
      (sum(rate(vm_http_request_errors_total{job="vmselect-benchmark-vm-cluster"}[5m])))
      /
      (sum(rate(vm_http_requests_total{job="vmselect-benchmark-vm-cluster"}[5m])))
...

The full list of generated rules is available here.

Now we can feed this configuration to the vmalert, and it will start evaluating the rules. But to retroactively evaluate these rules for all the data I already have, I’m going to run vmalert in replay mode using the same generated configuration file:

./vmalert -rule=slo_rules.yml \               # path to the configuration file
    -datasource.url=http://localhost:8428 \   # where to read metrics from
    -remoteWrite.url=http://localhost:8428 \  # where to persist results to
    -replay.timeFrom=2022-07-21T00:00:00Z \   # when to start the evaluation 
    -replay.timeTo=2023-01-21T00:00:00Z       # when to end the evaluation

As a source of data and destination for persisting results, I’m using a local installation of single-node VictoriaMetrics. The process of the replay mode looks like the following:

The vmalert's replay mode in-progress. The vmalert's replay mode in-progress.

In replay mode, rule groups and rules within groups are executed sequentially one by one. This is especially important for rules chaining - approach when the rule depends on the results of the previous rule. Just like in the SLI rules we generated above.

During evaluation, vmalert executes the rule’s expression via /query_range API to minimize the number of API calls. The configured time range is split into smaller ranges, so the API calls remain efficient and resilient to time series churn. The cache on the VictoriaMetrics side is automatically disabled by vmalert in order to prevent cache pollution. But VictoriaMetrics could have already cached responses for previously made requests, so it is recommended to follow general recommendations after data backfilling.

Once the replay is done, we should be able to see the recording rules backfilled and available for the whole time range:

Screenshot from the Grafana dashboard for SLO/SLI metrics <a href='https://sloth.dev/introduction/dashboards/' target='_blank'>by slok/sloth</a>. Screenshot from the Grafana dashboard for SLO/SLI metrics by slok/sloth.

And if we zoom-in and compare generated recording rule with the actual query used for its generation, we’ll see how they match:

Comparing recording rule results to the actual query in vmui. Comparing recording rule results to the actual query in vmui.

Retroactive alerts

#

Another cool thing about replay mode is that it supports retroactive evaluation of alerting rules. Actually, this mode was introduced specifically to support alerts evaluation in the past for satellite operator. So after we “replayed” SLO rules, we should be able to see whether any alerts triggered in the past:

vmui screenshot of the triggered alerting rule after `replay`. vmui screenshot of the triggered alerting rule after `replay`.

From the screenshot it is clear that alert actually triggered and its state correlates with errors rate spike.

Summary

#

Replay mode is a great feature for retroactive rules evaluation, both recording and alerting. The advantages of the replay mode are the following:

  • supports both recordings and alerting rules;
  • support rules chaining within the group;
  • uses Prometheus HTTP API and Remote Write protocol which makes it compatible with many other TSDBs;
  • can be used for evaluating rules against single and clustered installations;
  • can be configured with different endpoints for reading and writing, which makes it possible to migrate data from one installation to another via recording rules.

With all said, replay mode overcomes limitations of Prometheus backfilling for recording rules. But the replay mode has its own limitations: