惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
S
Securelist
GbyAI
GbyAI
The Register - Security
The Register - Security
B
Blog
Recorded Future
Recorded Future
D
DataBreaches.Net
C
Cybersecurity and Infrastructure Security Agency CISA
A
About on SuperTechFans
C
CERT Recently Published Vulnerability Notes
T
The Blog of Author Tim Ferriss
Vercel News
Vercel News
Google DeepMind News
Google DeepMind News
S
Schneier on Security
S
SegmentFault 最新的问题
Martin Fowler
Martin Fowler
T
Tenable Blog
T
The Exploit Database - CXSecurity.com
阮一峰的网络日志
阮一峰的网络日志
宝玉的分享
宝玉的分享
AWS News Blog
AWS News Blog
L
Lohrmann on Cybersecurity
Spread Privacy
Spread Privacy
N
News | PayPal Newsroom
Engineering at Meta
Engineering at Meta
T
Tor Project blog
The Hacker News
The Hacker News
量子位
酷 壳 – CoolShell
酷 壳 – CoolShell
MongoDB | Blog
MongoDB | Blog
Cyberwarzone
Cyberwarzone
Security Archives - TechRepublic
Security Archives - TechRepublic
爱范儿
爱范儿
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
C
Cyber Attacks, Cyber Crime and Cyber Security
T
Threatpost
WordPress大学
WordPress大学
Google Online Security Blog
Google Online Security Blog
G
GRAHAM CLULEY
Google DeepMind News
Google DeepMind News
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Attack and Defense Labs
Attack and Defense Labs
N
Netflix TechBlog - Medium
SecWiki News
SecWiki News
Hacker News: Ask HN
Hacker News: Ask HN
M
MIT News - Artificial intelligence
Scott Helme
Scott Helme
Microsoft Security Blog
Microsoft Security Blog
H
Help Net Security
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Using Ollama with the Laravel AI SDK: Run Local LLMs for Free
Hafiz · 2026-05-18 · via DEV Community

Originally published at hafiz.dev


API costs add up fast during AI development. You prompt an agent 50 times debugging a tool, that's 50 API calls. You run your test suite, that's another batch. Multiply that across a team and you're spending real money before shipping anything.

Ollama solves this cleanly. It runs open-source models locally on your machine (Llama 3, Qwen, Mistral, and dozens more) and the Laravel AI SDK treats it as a first-party provider, exactly like OpenAI or Anthropic. Switch between them with a single environment variable. No code changes, no new packages, no API keys.

This post covers the full setup: installing Ollama, configuring it in the Laravel AI SDK, building agents that run locally, and the dev/production workflow that lets you use Ollama locally while shipping with a cloud provider.

What Ollama Does

Ollama is a lightweight tool that downloads and serves open-source language models locally. Once it's running, it exposes an HTTP API on localhost:11434 that the Laravel AI SDK connects to directly.

There's no internet connection required after the initial model download. No rate limits. No costs per token. If you've built your app with the Laravel AI SDK smart assistant tutorial, your existing agents work with Ollama with a one-line change.

The tradeoff is hardware. Larger models need more RAM and a capable GPU to run at acceptable speeds. But for development, smaller models like Llama 3.2:3B run well on any modern developer machine.

Installing Ollama

macOS:

brew install ollama

Enter fullscreen mode Exit fullscreen mode

Or download the macOS app from ollama.com which installs as a menu bar app and starts automatically.

Linux:

curl -fsSL https://ollama.com/install.sh | sh

Enter fullscreen mode Exit fullscreen mode

Windows:
Download the installer from ollama.com. Ollama runs as a background service after installation.

Once installed, verify it's running:

curl http://localhost:11434
# Should return: Ollama is running

Enter fullscreen mode Exit fullscreen mode

Pulling Models

Download a model with ollama pull:

# General purpose, runs on any machine with 4GB+ RAM
ollama pull llama3.2

# Smaller version, 2GB, good for constrained machines
ollama pull llama3.2:1b

# Strong at code-related tasks, good for Laravel AI agents
ollama pull qwen2.5-coder:7b

# Mistral, fast and capable general model
ollama pull mistral

Enter fullscreen mode Exit fullscreen mode

You can list all downloaded models:

ollama list

Enter fullscreen mode Exit fullscreen mode

And test a model from the terminal before wiring it into Laravel:

ollama run llama3.2 "Explain Laravel service containers in one sentence"

Enter fullscreen mode Exit fullscreen mode

For the artisan commands used in this guide, having at least one model pulled before starting saves debugging time.

Configuring the Laravel AI SDK

The SDK ships with Ollama support out of the box. The only .env addition is:

OLLAMA_API_KEY=

Enter fullscreen mode Exit fullscreen mode

Leave the value blank. Ollama doesn't require authentication for local use, but the SDK expects the variable to exist. Add it to your .env and .env.example.

If Ollama is running on the default port, that's all you need. If you've changed the port or are running Ollama on a remote machine, configure the URL in config/ai.php:

'providers' => [
    // ... other providers

    'ollama' => [
        'driver' => 'ollama',
        'key'    => env('OLLAMA_API_KEY', ''),
        'url'    => env('OLLAMA_URL', 'http://localhost:11434/api'),
    ],
],

Enter fullscreen mode Exit fullscreen mode

And in .env:

OLLAMA_URL=http://localhost:11434/api

Enter fullscreen mode Exit fullscreen mode

The default URL is http://localhost:11434/api so for standard setups you don't need to add this. It works without it.

Using Ollama in Your Agents

Two ways to route an agent to Ollama: set it as the default provider for a specific agent class, or override it per-prompt at runtime.

Per-Agent with PHP Attributes

Add #[Provider] and #[Model] attributes to your agent class:

<?php

namespace App\Ai\Agents;

use Laravel\Ai\Attributes\Model;
use Laravel\Ai\Attributes\Provider;
use Laravel\Ai\Contracts\Agent;
use Laravel\Ai\Enums\Lab;
use Laravel\Ai\Promptable;

#[Provider(Lab::Ollama)]
#[Model('llama3.2')]
class SupportAgent implements Agent
{
    use Promptable;

    public function instructions(): string
    {
        return 'You are a helpful support agent. Answer questions about our product concisely.';
    }
}

Enter fullscreen mode Exit fullscreen mode

Now every time you prompt this agent, it uses Ollama locally:

$response = SupportAgent::make()->prompt('How do I reset my password?');

return (string) $response;

Enter fullscreen mode Exit fullscreen mode

This is the cleanest pattern for development. You write your agent once with Ollama attributes, build and test locally with no API costs, then change the attributes (or override them via .env) when deploying to production.

Overriding Per-Prompt

For one-off local testing without modifying the agent class:

use Laravel\Ai\Enums\Lab;

$response = SupportAgent::make()
    ->prompt('How do I reset my password?', provider: Lab::Ollama, model: 'llama3.2');

Enter fullscreen mode Exit fullscreen mode

This is useful when you want to quickly compare responses between Ollama and a cloud provider without changing the agent configuration.

The Dev/Production Workflow

The cleanest approach is to set a default provider at the application level in config/ai.php, driven by environment variables:

'default' => [
    'text' => [
        'provider' => env('AI_PROVIDER', 'openai'),
        'model'    => env('AI_MODEL', 'gpt-4o'),
    ],
],

Enter fullscreen mode Exit fullscreen mode

Then in your local .env:

AI_PROVIDER=ollama
AI_MODEL=llama3.2

Enter fullscreen mode Exit fullscreen mode

And in production .env (or your Forge/Vapor environment):

AI_PROVIDER=anthropic
AI_MODEL=claude-sonnet-4-5

Enter fullscreen mode Exit fullscreen mode

Zero code changes between environments. Your agents, tools, and structured output stay identical. Only the provider changes. This works well for any agent that doesn't use PHP attribute overrides; those take precedence over the default config.

For agents with explicit #[Provider] attributes, you'd need to either remove the attributes or use a different approach for environment-based switching. The attribute approach is better for agents that should always use a specific provider (a code review agent that truly needs a smart model in all environments). The default config approach is better for general-purpose agents where Ollama in dev and a cloud model in prod makes sense.

Which Models to Use

Not all models are equal, and the right choice depends on what your agent is doing. Here's a practical guide based on common Laravel AI SDK use cases.

Llama 3.2 (3B or 8B) is the safe default for most use cases. The 3B version runs comfortably on any developer machine with 4GB RAM. The 8B version is noticeably better at following complex instructions but needs 8GB. Good for support agents, document summarisation, and general Q&A. Start here if you're not sure.

Qwen 2.5 Coder (7B) is the right choice for agents that work with code. It outperforms Llama on code generation and review tasks despite similar size. If you're building an agent that analyzes PHP files, generates migrations, or reviews code quality, use this one instead.

Mistral (7B) is fast and reliable for instruction-following tasks. If you need quick responses and the task isn't code-heavy, Mistral is worth trying. It tends to be faster than Llama 3.2 at the same quality level.

Avoid very large models (30B+) for development. They're slow on typical developer machines and the speed penalty makes iteration painful. The quality gap between 7B and 30B matters less in development where you're primarily testing tool calls and output format, not production response quality. Save the big models for your production cloud provider.

A practical setup for a Laravel SaaS would be: use llama3.2:8b for general agents and qwen2.5-coder:7b for any agent touching code. Both run on a 16GB machine without issues. If you're on a 8GB machine, use llama3.2:3b for everything and accept slightly weaker instruction following in exchange for speed.

If you've already built a multi-agent system with the SDK, you can route different sub-agents to different Ollama models the same way you'd assign different cloud models, and the AI SDK overview covers the broader SDK capabilities worth knowing before diving into local model optimization.

Embeddings with Ollama

Ollama also works for local embeddings, which means you can do RAG development with zero API costs:

use Laravel\Ai\Facades\Ai;
use Laravel\Ai\Enums\Lab;

$embedding = Ai::embed(
    'How do I cancel my subscription?',
    provider: Lab::Ollama,
    model: 'nomic-embed-text'
);

Enter fullscreen mode Exit fullscreen mode

Pull the embedding model first:

ollama pull nomic-embed-text

Enter fullscreen mode Exit fullscreen mode

nomic-embed-text is a solid local embedding model that produces 768-dimension vectors. For production RAG you'd swap to OpenAI's text-embedding-3-small or a similar cloud model, but for building and testing your vector search logic, Ollama keeps costs at zero.

What Ollama Doesn't Support

The Laravel AI SDK's Ollama integration covers text generation and embeddings. It does not support image generation, text-to-speech, speech-to-text, or file uploads. If your agents use those capabilities, you'll need a cloud provider for those specific features.

This is usually fine for a dev/production split. Most agent logic (tools, structured output, conversation flow) doesn't depend on images or audio. You can run the core agent logic against Ollama locally, and the multimedia features only come into play in staging or production against cloud providers.

Testing Agents That Use Ollama

One thing to be aware of: when running your test suite, you probably don't want tests making real Ollama calls any more than you'd want real OpenAI calls. The SDK's fake testing utilities work regardless of which provider is configured:

it('responds to password reset questions', function () {
    SupportAgent::fake([
        'To reset your password, visit the login page and click "Forgot password".',
    ]);

    $response = SupportAgent::make()->prompt('How do I reset my password?');

    expect((string) $response)->toContain('password');
});

Enter fullscreen mode Exit fullscreen mode

Faking the agent response means your tests are fast, deterministic, and don't depend on Ollama being installed or running. The agent safety post covers more on keeping agent behavior predictable in tests.

The development workflow then becomes: build and iterate against real Ollama locally, run the test suite with faked responses, deploy with cloud providers in production.

Ollama on a Shared Dev Server

If your team uses a shared development server, you can run Ollama there and point everyone's local Laravel instances at it. Just update OLLAMA_URL in each developer's .env:

OLLAMA_URL=http://your-dev-server:11434/api

Enter fullscreen mode Exit fullscreen mode

Make sure Ollama is configured to accept connections from outside localhost on the server:

OLLAMA_HOST=0.0.0.0 ollama serve

Enter fullscreen mode Exit fullscreen mode

This means one machine does the model serving and your team shares it, without everyone needing to pull and run models locally. Useful if some team members are on constrained hardware.

FAQ

Does Ollama work with Laravel AI SDK agents that use tools?

Yes, but model quality matters more for tool use. Some smaller models handle tool calls inconsistently. Llama 3.2 8B is reliable for tool use. If you're seeing missed or malformed tool calls, try a larger or more capable model.

Can I use Ollama in production?

You can if you have dedicated server hardware with enough RAM and ideally a GPU. Most teams use Ollama for local development and testing, then cloud providers in production. The cost and maintenance overhead of running Ollama in production usually outweighs the savings unless you have high volume and a specific privacy requirement.

What's the difference between Ollama and running models via API?

With Ollama, the model runs on your machine. No data leaves your network. With cloud APIs (OpenAI, Anthropic), your prompts are sent to the provider's servers. For development involving sensitive or proprietary data, Ollama is the better choice.

Do I need a GPU?

No. Most models run on CPU, just more slowly. For development iteration a CPU is fine. Responses take 5-15 seconds depending on model size and your hardware. A GPU drops that to under 2 seconds for 7B models.

Can I use Ollama with the sub-agents pattern?

Yes. Each sub-agent can have its own #[Provider(Lab::Ollama)] and #[Model] attributes. The sub-agents guide covers the full pattern; the Ollama attributes drop in without any other changes.

Start Locally

The setup comes down to four steps: install Ollama, pull a model, add OLLAMA_API_KEY= to your .env, and add #[Provider(Lab::Ollama)] to your agent class. After that, you're running AI locally with no API costs and no rate limits while you build.

In production, switch back to OpenAI or Anthropic by changing the provider attribute or your default config. The rest of your code stays exactly the same.

If you're setting this up for a team or have questions about the dev/production split, get in touch.