惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Privacy & Cybersecurity Law Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
D
Docker
V
V2EX
GbyAI
GbyAI
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - Franky
Jina AI
Jina AI
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
I
InfoQ
博客园 - 司徒正美
雷峰网
雷峰网
F
Full Disclosure
S
SegmentFault 最新的问题
大猫的无限游戏
大猫的无限游戏
博客园 - 叶小钗
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
D
DataBreaches.Net
M
MIT News - Artificial intelligence
V
Visual Studio Blog
H
Help Net Security
月光博客
月光博客
博客园 - 三生石上(FineUI控件)
博客园_首页
O
OpenAI News
人人都是产品经理
人人都是产品经理
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Attack and Defense Labs
Attack and Defense Labs
Blog — PlanetScale
Blog — PlanetScale
爱范儿
爱范儿
罗磊的独立博客
P
Palo Alto Networks Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
博客园 - 聂微东
Last Week in AI
Last Week in AI
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
Lohrmann on Cybersecurity
N
News and Events Feed by Topic
有赞技术团队
有赞技术团队
The Register - Security
The Register - Security
S
Security @ Cisco Blogs
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
酷 壳 – CoolShell
酷 壳 – CoolShell
AWS News Blog
AWS News Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
L
LINUX DO - 最新话题
Hacker News - Newest:
Hacker News - Newest: "LLM"
T
Threat Research - Cisco Blogs

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Logging Your AI Events (from Ollama) in Bronto
Patrick Lond · 2026-05-20 · via DEV Community

Authored by David Tracey

Many software companies are investigating the use of Large Language Models (LLMs) in their products. At Bronto we've announced our Bronto Labs initiative, with AI features including auto-parsing, AI dashboard creation, and Bronto Scope for error investigation.

This post explores a different angle: using logs in the development of AI applications. We'll focus on Ollama — an open source tool for running LLMs locally — and show how to pipe its logs into Bronto for search and analysis.

LLMs are complex, non-deterministic systems. Beyond traditional logging use cases (performance monitoring, API usage), their unpredictable nature increases the need for logging — particularly to record and track responses to prompts. Individual log events can be large when they include a full prompt or response. Meta found this problem significant enough at their scale to build a dedicated Meta AI Logging Engine.

The fundamental requirements for logging AI applications are:

  • Ability to handle large log events
  • Ability to handle high volumes at low cost
  • Ability to search across high volumes quickly

These are exactly the requirements Bronto was designed to meet.


Setting Up Ollama

Recommended specs:

  • 16GB RAM (8GB works for smaller models)
  • 12GB disk space for Ollama and basic models
  • Modern CPU with at least 4 cores (8 preferred)
  • Optional: GPU for improved performance

Install and Run the Server

Install from ollama.com/download for your OS, then start the server:

ollama serve

Enter fullscreen mode Exit fullscreen mode

You'll see output including the default port it's listening on (11434).

Download and Run a Model

# Pull a model from the registry
ollama pull gemma:2b

# List downloaded models
ollama list

# Run a model interactively
ollama run gemma:2b

Enter fullscreen mode Exit fullscreen mode

The run command gives you a >>> prompt where you can enter prompts or /help for commands.


Sending Ollama Logs to Bronto

Step 1: Configure Ollama Logging to File

Stop the server and restart it writing logs to a file:

ollama serve > /your_log_path/.ollama/logs/server.log 2>&1

Enter fullscreen mode Exit fullscreen mode

For more detailed debug logs, add to your shell profile (.zprofile etc.):

export OLLAMA_LOG_LEVEL=DEBUG
export OLLAMA_DEBUG=true

Enter fullscreen mode Exit fullscreen mode

To redirect model client logs:

# stderr only (keeps console interactive)
ollama run gemma:2b 2>>/your_log_path/.ollama/logs/gemma.log

# both stdout and stderr (API use only — disables console input)
ollama run gemma:2b > /your_log_path/.ollama/logs/gemma.log 2>&1

Enter fullscreen mode Exit fullscreen mode

Verify logs are flowing:

tail -f /your_log_path/.ollama/logs/server.log

Enter fullscreen mode Exit fullscreen mode

Step 2: Install OpenTelemetry Collector

Download for your platform from opentelemetry.io. Example for Mac ARM64:

curl --proto '=https' --tlsv1.2 -fOL \
  https://github.com/open-telemetry/opentelemetry-collector-releases/releases/download/v0.114.0/otelcol-contrib_0.114.0_darwin_arm64.tar.gz

chmod +x otelcol-contrib
mv otelcol-contrib /usr/local/bin/otelcol

# Verify
otelcol --version

Enter fullscreen mode Exit fullscreen mode

Step 3: Configure OpenTelemetry to Forward to Bronto

Create /etc/otelcol/config.yaml:

receivers:
  filelog/Ollama_Server:
    include:
      - /your_log_path/.ollama/logs/server.log
    resource:
      service.name: LaptopServer
      service.namespace: Ollama

  filelog/Ollama_Gemma:
    include:
      - /your_log_path/.ollama/logs/gemma.log
    resource:
      service.name: LaptopGemma
      service.namespace: Ollama

processors:
  batch:

exporters:
  otlphttp/brontobytes:
    logs_endpoint: "https://ingestion.us.bronto.io/v1/logs"
    compression: none
    headers:
      x-bronto-api-key: replace_this_with_your_bronto_apikey

service:
  pipelines:
    logs:
      receivers: [filelog/Ollama_Server, filelog/Ollama_Gemma]
      processors: [batch]
      exporters: [otlphttp/brontobytes]
  # Useful for debugging:
  # telemetry:
  #   logs:
  #     level: "debug"
  #     output_paths: [/your_log_path/otelcol/debug.log]

Enter fullscreen mode Exit fullscreen mode

Validate and run:

otelcol validate --config=/etc/otelcol/config.yaml
otelcol --config=/etc/otelcol/config.yaml

Enter fullscreen mode Exit fullscreen mode


A Simple Ollama API Program

The Python script below (ollama-log-demo.py) uses the Ollama API to send prompts against a log file and print the response. Example usage:

# Summarize 100 lines of CDN logs
python3 ollama-log-demo.py 100lines-CDN-log.csv \
  --model "gemma:2b" \
  --prompt "You have been given 100 lines from a CDN log in CSV format. Summarise the logs provided."

# Find errors and suggest fixes
python3 ollama-log-demo.py 100lines-search-log.csv \
  --model "gemma:2b" \
  --prompt "Find errors in this log and suggest how to fix them"

Enter fullscreen mode Exit fullscreen mode

The final line of each Ollama response includes useful performance metadata:

Field Description
total_duration Total time spent generating the response
load_duration Time spent loading the model (nanoseconds)
prompt_eval_count Number of tokens in the prompt
prompt_eval_duration Time spent evaluating the prompt (nanoseconds)
eval_count Number of tokens in the response
eval_duration Time spent generating the response (nanoseconds)
context Conversation encoding — pass in next request to maintain memory
response Empty if streamed; full response if not streamed

Model notes from testing: gemma:2b is good for summarizing but tends to give high-level summaries even when asked for specifics. mistral takes longer but produces more detailed, data-specific responses. Defining the right prompt for your use case is key.


Searching Ollama Logs in Bronto

Ollama server logs include a mix of structured and unstructured entries:

Standard log levels:

INFO [main] HTTP server listening | hostname="127.0.0.1" port="11434"
level=INFO source=sched.go:714 msg="new model will fit in available VRAM"
level=DEBUG source=memory.go:103 msg=evaluating library=metal gpu_count=1

Enter fullscreen mode Exit fullscreen mode

Model and resource logs:

llm_load_print_meta: max token length = 93
llama_model_loader: - kv 0: general.architecture str = gemma
level=INFO source=server.go:105 msg="system memory" total="8.0 GiB" free="1.2 GiB"

Enter fullscreen mode Exit fullscreen mode

Even a small test with short prompts generates surprisingly large log volumes — 244 events totaling ~2MB in our test. Bronto handles these unstructured and semi-structured formats natively, and you can add a custom parser to make them more convenient to search and view.

Example searches in Bronto:
Fig.1 — Searching for log events containing "tokens"
Searching for log events containing

Fig.2 — Searching for log events containing "prompt"
Searching for log events containing
Fig.3 — Grouping by prompt evaluation time per task_id

Grouping by prompt evaluation time per task_id


Conclusion

This post introduced Ollama as an example of an LLM system and explained why AI applications create unique logging challenges — large events, high volumes, non-deterministic outputs, and distributed agents. We walked through setting up Ollama locally, configuring OpenTelemetry to forward logs to Bronto, and writing a simple Python API program to experiment with prompts against log data.

Future posts will develop the theme further with other AI systems including AWS Bedrock.


Appendix: ollama-log-demo.py

import argparse
import json
import requests


def print_ollama_stats(json_response):
    load_duration = json_response.get("load_duration")
    if load_duration:
        print("\n--- load_duration = ", load_duration)

    total_duration = json_response.get("total_duration")
    if total_duration:
        print("\n--- total_duration = ", total_duration)

    eval_duration = json_response.get("eval_duration")
    if eval_duration:
        print("\n--- eval_duration = ", eval_duration)

    prompt_eval_duration = json_response.get("prompt_eval_duration")
    if prompt_eval_duration:
        print("\n--- prompt_eval_duration = ", prompt_eval_duration)

    prompt_eval_count = json_response.get("prompt_eval_count")
    if prompt_eval_count:
        print("\n--- prompt_eval_count = ", prompt_eval_count)

    eval_count = json_response.get("eval_count")
    if eval_count:
        print("\n--- eval_count = ", eval_count)


def examine_log_with_prompt(file_path, input_prompt, input_model):
    with open(file_path, 'r') as file:
        log_data = file.read()

    req_params = {
        "model": input_model,
        "prompt": f"{input_prompt}\n\n{log_data}"
    }

    try:
        # Update localhost URL to match your Ollama API endpoint
        response = requests.post(
            "http://localhost:11434/api/generate",
            headers={"Content-Type": "application/json"},
            data=json.dumps(req_params),
            stream=True
        )
        if response.status_code == 200:
            print("\n--- Processing Successful Ollama Response ---")
            line_count = 0
            for line in response.iter_lines():
                if line:
                    try:
                        json_line = line.decode('utf-8')
                        line_count += 1
                        json_response = json.loads(json_line)
                        print(json_response["response"], end='', flush=True)
                    except json.JSONDecodeError as e:
                        print(f"Error decoding JSON on line {line_count + 1}: {e}")
                    except UnicodeDecodeError as e:
                        print(f"Error decoding line to UTF-8 on line {line_count + 1}: {e}")
            if line_count == 0:
                print("No JSON lines found or response was empty.")
            print("\n--------------------------------------------------")
            print_ollama_stats(json_response)
            print("\n--------------------------------------------------")
        else:
            print(f"\nError - Response Status code: {response.status_code}")
            print(response.text)
    except Exception as e:
        print(e)


def main():
    parser = argparse.ArgumentParser(description="Ollama API Demo for Logs")
    parser.add_argument('file', type=str, help='Path to the log file to be examined')
    parser.add_argument('--model', type=str, help='Model to use in analysis', default=None)
    parser.add_argument('--prompt', type=str, help='Prompt to send to model', default=None)
    args = parser.parse_args()
    examine_log_with_prompt(args.file, args.prompt, args.model)


if __name__ == "__main__":
    main()

Enter fullscreen mode Exit fullscreen mode

Explore Bronto's AI Features