惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LINUX DO - 最新话题
A
Arctic Wolf
I
Intezer
V
Vulnerabilities – Threatpost
C
Cisco Blogs
MyScale Blog
MyScale Blog
NISL@THU
NISL@THU
Y
Y Combinator Blog
C
CERT Recently Published Vulnerability Notes
P
Privacy International News Feed
H
Hackread – Cybersecurity News, Data Breaches, AI and More
酷 壳 – CoolShell
酷 壳 – CoolShell
Recorded Future
Recorded Future
云风的 BLOG
云风的 BLOG
S
SegmentFault 最新的问题
Microsoft Security Blog
Microsoft Security Blog
L
LangChain Blog
博客园 - 聂微东
博客园 - 叶小钗
F
Fortinet All Blogs
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Recent Announcements
Recent Announcements
C
Cyber Attacks, Cyber Crime and Cyber Security
Latest news
Latest news
Simon Willison's Weblog
Simon Willison's Weblog
P
Palo Alto Networks Blog
S
Schneier on Security
C
Cybersecurity and Infrastructure Security Agency CISA
V
V2EX
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
The Hacker News
The Hacker News
博客园 - 司徒正美
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
LINUX DO - 热门话题
罗磊的独立博客
K
Kaspersky official blog
Last Week in AI
Last Week in AI
Know Your Adversary
Know Your Adversary
小众软件
小众软件
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
D
DataBreaches.Net
Scott Helme
Scott Helme
P
Proofpoint News Feed
P
Privacy & Cybersecurity Law Blog
P
Proofpoint News Feed
博客园 - 三生石上(FineUI控件)
Hugging Face - Blog
Hugging Face - Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
F
Full Disclosure

Redis

Real-Time Fraud Detection: Latency, Features & Scale Context window in AI: why every token is a budget decision Connecting to Redis Cloud with AWS PrivateLink vs. VPC peering | Redis Redis Data Integration in Redis Cloud is now GA in AWS | Redis Why AI Misses Business Context & How Teams Fix It AI Reasoning Explained: Why Context Matters Semantic Layer vs Context Layer: Key Differences Redis array data type: How it works and when to use it Context Graphs vs. Vector Search: When RAG Falls Short What’s new in two – May 2026 edition Redis 8.8 performance improvements: Faster string, hash, streams, SCAN & more Redis 8.8: New array data structure & open source features How Conflict-free Replicated Data Types power active-active database replication Context Orchestration: What It Is & How It Works Context Compaction for AI Agents: A Complete Guide Prompt Bloat: Causes, Costs & Fixes for LLM Apps Agentic Retrieval Techniques: A Complete Guide Single-shot reliable consumers with XREADGROUP CLAIM in Redis 8.4 | Redis Long-Horizon AI Agents: Memory & State Infrastructure What is a context engine? What Is a Context Layer? AI Agent Infrastructure Context Retrieval for AI Agents: What It Is & Why It Matters Context Poisoning: How Bad Data Breaks Agent Reasoning Context is all you need: Introducing Redis Iris | Redis Context Engineering for AI: What It Is & How to Build It Dynamic endpoints: Migrate databases without changing your endpoint | Redis AI Shopping Assistants: How They Work & What to Build Endless Aisle Retail: Infrastructure & Real-Time Data LLM Speed Benchmarks: Metrics & Infrastructure Guide Context Pruning: Cut LLM Tokens Without Losing Quality What’s new in two – April 2026 edition Agentic AI Architecture: 5 Patterns Explained AI Agent vs Chatbot: Key Differences Explained Advantages of Building a Vector Search Solution API Latency in LLM Apps: Causes & How to Fix It Security advisory: [CVE‑2026‑23479] [CVE‑2026‑25243] [CVE-2026-25588] [CVE‑2026‑25589] [CVE-2026-23631] | Redis Edge Computing Latency: Causes & How to Reduce It AI Agents vs Workflows: When to Use Each Streaming LLM Responses: Make Your AI App Feel Fast Active-Active vs Active-Passive Database Architecture Prefill vs Decode: LLM Inference Phases Explained Long-Term Memory Architectures for AI Agents Time to First Byte Test: Tools, Causes & Fixes Speculative decoding: how it works & when to use it P95 Latency: What It Is & Why It Matters Why Multi-Agent LLM Systems Fail & How to Fix Them AI Human in the Loop: Production Oversight Patterns Native OpenTelemetry metrics for Redis client libraries | Redis Client-side geographic failover for Redis Active-Active | Redis Use Redis with SQL | Redis Introducing Redis Feature Form Build Google ADK Agents with persistent, real-time memory on Redis | Redis Startup Spotlight: Neuron Systems API Throttling: Algorithms, Patterns & Mistakes Agentic AI Examples Across 6 Industries Best Chunking Strategies for RAG Pipelines Agentic AI Guardrails: Controls That Work Redis joins AWS at GDC to support the next generation of gaming | Redis Designing a semantic routing system: From static rules to dynamic intelligence with Redis and Java | Redis Real-Time Dispatch System: A Complete Guide P99 Latency: What It Means & How to Fix It Tokenization in LLMs: What AI App Devs Need to Know TTFT Meaning: What is Time to First Token? Atomic slot migration with Redis 8.4 Hybrid search benefits: Why your RAG system needs both keyword & vector search What’s new in two: March 2026 edition Vector embedding generators: How they work & how to use them Throughput-optimizing Redis for L2 KV Cache Reuse What is a data pipeline? Building AI agent pipelines that don't forget, fail, or fall apart Redis achieves Google Cloud Ready, Distributed Cloud status ahead of Google Cloud Next ‘26 | Redis Real-time network monitoring: what your data platform needs to keep up AI agent API: How agents connect to the real world What is multicloud infrastructure? A guide for 2026 What is a transaction monitoring system & how does it work? Why your AI agent fails in production & how tracing helps AI agent benchmarks: Where they fall short & why your infrastructure matters What is a JSON database (and when should you use one)? Introducing the Redis Partner Network: A new foundation for real-time innovation How real-time customer segmentation works in retail Payment orchestration & vault architecture in retail Agentic systems vs. GenAI: when generation isn't enough What is fuzzy matching? Semantic caching & routing: two powerful patterns for vector classification Redis alternatives: Why there are no exact substitutes Connect to Azure Managed Redis with Redis Insight 3.2.0 How to tame the thundering herd problem Redis to Manage Storage Replication | Redis How hierarchical navigable small world (HNSW) algorithms can improve search | Redis How leading financial institutions use Redis to drive growth | Redis What’s new in two: May 2025 | Redis Introducing Model Context Protocol (MCP) for Redis | Redis Redis vs. Elasticsearch: What’s faster for GenAI & vector search? | Redis Build fast, production-worthy AI apps with Spring AI and Redis | Redis Azure Managed Redis is GA today | Redis Redis then & now: Adapting with developers through every era | Redis What’s new in two: April 2025 | Redis Redis 8 is now GA, loaded with new features and more than 30 performance improvements | Redis What is a data strategy? 6 key components explained Data replication explained: types, examples & use cases
Supercharge Your AI with OpenShift AI and Redis: Unleash speed and scalability | Redis
2025-05-03 · via Redis

Since the birth of large language models (LLMs) and the release of ChatGPT, artificial intelligence (AI) has gone from being an out-of-reach concept to showing real promise in the business landscape for every industry and business. From personalized customer experiences to streamlined operations and increased security, the possibilities are endless.

The reality we encounter in many companies, however, shows a lack of robust and flexible infrastructure technologies that will allow the business to harness the full potential of these AI promises. Leaders see the potential of AI and want to adopt it, but the teams do not have access to the basic tooling that will allow them to explore it.

This is where Red Hat OpenShift AI and Redis come into play, offering a powerful environment for data scientists and machine learning (ML) engineers.

OpenShift AI provides organizations with an efficient way to deploy and manage a comprehensive set of AI/ML tools. Its ability to create custom environments ensures that data scientists and machine learning engineers always have the right resources at their fingertips.

Redis is the world’s fastest in-memory database. It’s a versatile solution that has evolved beyond a simple key-value data store to support a wide range of use cases, including:

  • Vector database: Store and query vector data for similarity searches
  • Retrieval augmented generation (RAG): Enhance LLM accuracy and relevance by grounding searches in real-time data
  • LLM memory: Manage the context window of LLMs for more coherent and context-aware conversations
  • Semantic cache: Optimize performance and reduce LLM costs by caching semantically similar prompts and responses

Better together: OpenShift AI and Redis

AI applications, especially those involving generative AI (gen AI), demand high performance and low latency. Users expect real-time responses and personalized experiences. The combination of OpenShift AI and Redis addresses these challenges head-on.

Supercharge Your AI

OpenShift AI provides the environment where data scientists can use different tools, including embedding models, third-party frameworks like LangChain or LlamaIndex and multiple LLMs to implement their gen AI use cases at scale. Redis delivers the sub-second latency that gen AI use cases need.

Let’s dive into specific use cases:

1. Retrieval augmented generation (RAG)

RAG enhances the knowledge of LLMs by integrating external data sources. Instead of solely relying on their pretrained knowledge, LLMs can fetch relevant information from a database in real time to generate more accurate and contextually appropriate responses. Fine-tuning an LLM with business-specific data is traditionally a costly and time-consuming process, and it may not be viable, depending on how often the knowledge base changes (with new or updated records). Keeping the knowledge base external to the model provides more flexibility and makes it easier to ensure that the LLM always has the latest information to serve the users.

Redis Technical Diagram RedHat Blog RAG

The business benefit: Improved accuracy, reduced hallucinations and access to up-to-date information for chatbots, content generation tools and more.

In this use case, Redis plays the role of the vector database, while Openshift AI provides the compute resources and notebook environment with pipelines and model serving tools for the data scientists to prepare the vector data and test the quality and performance of semantic searches. OpenShift AI also provides advanced tooling like guardrails to improve LLM accuracy and monitor and better safeguard both user input interactions and model outputs.

By using Redis as the vector database, administrators can configure role-based access control (RBAC) and access control lists (ACLs) to separate vector data between different users, departments, etc. This allows vectors that contain sensitive information that can only be shared to certain users. Redis implements this control at the highest level, not only as a query parameter.

2. Semantic cache

LLMs can be expensive to run, especially for repetitive queries. A semantic cache stores LLM responses based on the meaning of the query, not just the exact text. When the user submits a new prompt, the system will look for similar prompts, and if it finds a match, it will retrieve the LLM response directly from the cache, saving a trip to the LLM server and the associated token cost (when using a hosted LLM service) or compute capacity (when self-hosting a model).

Redis Technical Diagram RedHat Blog Semantic Cache

The business benefit: Semantic caching can greatly reduce LLM costs, especially for use cases where users are expected to ask basic or generic questions (FAQs, etc). Other benefits include faster response times and improved scalability for AI-powered applications.

In this use case, Redis will be used as the vector database and semantic cache. Some frameworks, like LangChain, are configured immediately to save prompts to the semantic cache and check automatically for each new prompt. This allows developers to quickly take advantage of this capability without having to write a lot of code, or control reads and writes to the cache. Data scientists can define the distance threshold for the cache based on the requirements and characteristics of the use case.

OpenShift AI provides the environment and tooling such as Jupyter and data science pipelines to create and run the code that will generate the vector data and load it. This can help data scientists not only generate the vector data, but also preload the semantic cache with thousands of questions and answers that can be generated with the help of a LLM. That way, when the first user asks a question, there is a good probability that a similar question is already in the cache, reducing the response time to milliseconds, potentially 15x faster than the traditional RAG pipeline.

If we can ensure a user experience where most of the prompts take only milliseconds to respond, how else can we enrich this user experience? What other data or services can we add to it, if we know that users won’t be waiting for seconds every time?

3. LLM memory

LLMs are, by definition, stateless. Meaning, they keep no record of any previous interaction with the user. As far as the model is aware, every prompt is an entirely new prompt, with no past or history to consider.

To get around this limitation, client applications (like chatbots) keep the conversation history between the user and the model (plus some additional information) and serve this data to the model every time the user submits a new prompt. The capacity to use this data is called the “context window,”’ and it makes it possible for the model to keep the answers within the context of the conversation that is happening.

Redis Technical Diagram RedHat Blog LLM Memory

The business benefit: More engaging and context-aware chatbots, personalized customer experiences and improved ability to handle complex conversations.

Redis not only stores the context window data, but again provides immediate integrations with frameworks like LangChain, which allows developers to enable LLM memory using only 2 lines of code. Additionally, using Redis to store LLM memory has other impactful benefits, such as:

  • It enables multichannel user experiences. Users can close the browser-based chatbot and open the voice assistant in their mobile application and continue the conversation exactly where they left off, because both clients are pulling the conversation history from Redis;
  • In case a call center or some other internal team needs to review the conversation history between user and bot, it can very easily be retrieved from Redis. This allows internal staff to understand exactly what happened in that conversation and whether or not the information provided by the model was correct.

Additionally, OpenShift AI provides data scientists with an environment that can be customized to grant access to the conversation history within Redis to fine-tune the conversation model. Having access to a dataset that includes real questions and answers can be critical to ensure the continuous improvement of the LLM responses. With OpenShift AI, data scientists have the resources and tooling to analyse and prepare a dataset to use for fine-tuning the embedding model or the LLM (or to prepare new prompts to ‘pre-warm’ the semantic cache).

Now that we’ve covered the main use cases, let’s see how we can put these ideas into action.

Getting started

Deploying Redis to Red Hat OpenShift can be greatly simplified using the OperatorHub. In the OpenShift web console, go to the OperatorHub page (in the Operators section of the left-side navigation panel).

From there, you can browse to the Database tab and look for Redis, or simply type Redis in the search bar.

Supercharge Your AI

Then you can open the Details page and click on the Install button to deploy the Redis operator.

Once the Operator is deployed, Redis resources can be easily and quickly created through the OpenShift UI:

Supercharge Your AI Blog

There are 2 main resources for Redis: the cluster and the database (along with their active-active counterparts, which is outside the scope of this article).

The Redis cluster manages multiple databases, ensuring high availability and scalability. This is the first resource that needs to be created. Once the cluster is created, a new database can be created to serve the use cases discussed above. Make sure to enable the Search and JSON support capabilities, as they are necessary for vector search.

Once the database is created, users can access the Redis console to retrieve the connection information, check metrics and track the overall health of the database.

Supercharge Your AI

Next, the OpenShift AI environment can be configured. Here, users can create a notebook environment, provision inference services for local LLMs, create and configure data science pipelines and much more.

Supercharge Your AI

Jupyter notebooks provide a very simple and convenient way to experiment with vector searches. Users can connect to the Redis database with only a few lines of code, and from there, they can take advantage of popular frameworks like LangChain and LlamaIndex, or they can use the redis-vl package, which allows them to use vector capabilities without requiring a specific framework.

Supercharge Your AI Blog

Conclusion

The increasing popularity of AI tools has unquestionably enabled a broad set of opportunities for companies that are looking to modernize, innovate and stay ahead of the competition.

Providing an environment with the resources and capabilities needed to take advantage of these AI tools is one of the greatest challenges companies face, as they try to understand the value AI could bring to their business.

OpenShift AI and Redis offer a powerful combination for organizations looking to use AI effectively. By providing a flexible gen AI development environment and a high-performance data platform, these technologies empower data scientists and machine learning engineers to build innovative AI solutions that create real business value. Whether it’s RAG, semantic caching or LLM memory, OpenShift AI and Redis provide the foundation for building intelligent applications that are fast, scalable and context-aware.