惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
腾讯CDC
阮一峰的网络日志
阮一峰的网络日志
GbyAI
GbyAI
B
Blog RSS Feed
Engineering at Meta
Engineering at Meta
Google DeepMind News
Google DeepMind News
MyScale Blog
MyScale Blog
Last Week in AI
Last Week in AI
F
Fortinet All Blogs
云风的 BLOG
云风的 BLOG
N
Netflix TechBlog - Medium
G
Google Developers Blog
博客园_首页
有赞技术团队
有赞技术团队
V
V2EX
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
MongoDB | Blog
MongoDB | Blog
H
Help Net Security
aimingoo的专栏
aimingoo的专栏
月光博客
月光博客
Hugging Face - Blog
Hugging Face - Blog
The GitHub Blog
The GitHub Blog
S
SegmentFault 最新的问题

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
Running Local GGUF Models with Ollama (GPU Enabled)
KALPESH · 2026-05-16 · via DEV Community
Cover image for Running Local GGUF Models with Ollama (GPU Enabled)

KALPESH

1. Install & Start Ollama

curl -fsSL https://ollama.com/install.sh | sh
systemctl start ollama
ollama --version

Enter fullscreen mode Exit fullscreen mode


2. Verify GPU Detection

NVIDIA

nvidia-smi

Enter fullscreen mode Exit fullscreen mode

AMD

rocm-smi

Enter fullscreen mode Exit fullscreen mode


3. Set Up Model Directory

mkdir -p ~/Documents/LLM
cd ~/Documents/LLM
# Copy your .gguf file here

Enter fullscreen mode Exit fullscreen mode


4. Create a Modelfile

vim Modelfile

Enter fullscreen mode Exit fullscreen mode

Vim quick reference:

  • i — enter insert mode (start typing)
  • Esc — exit insert mode
  • :wq — save and quit
  • :q! — quit without saving
FROM ./Phi-4-mini-instruct-Q4_K_M.gguf

SYSTEM """
You are a helpful AI assistant.
"""

TEMPLATE """<|user|>
{{ .Prompt }}<|end|>
<|assistant|>
"""

PARAMETER stop "<|user|>"
PARAMETER stop "<|assistant|>"
PARAMETER stop "<|end|>"
PARAMETER temperature 0.7
PARAMETER num_ctx 8192

Enter fullscreen mode Exit fullscreen mode

Note: Always include TEMPLATE for custom GGUFs. Use instruct/chat variants, not base models.


5. Create & Run the Model

ollama create mymodel -f Modelfile
ollama run mymodel

Enter fullscreen mode Exit fullscreen mode


6. Verify GPU Usage

Open a second terminal and monitor VRAM — an increase confirms GPU acceleration.

# NVIDIA
watch -n 1 nvidia-smi

# AMD
watch -n 1 rocm-smi

Enter fullscreen mode Exit fullscreen mode

To confirm via logs:

journalctl -u ollama -f
# Look for: "using CUDA" or "offloading layers to GPU"

Enter fullscreen mode Exit fullscreen mode


7. Ollama Command Reference

Model Management

Task Command
Pull a model ollama pull <model>
Create from Modelfile ollama create <name> -f Modelfile
List installed models ollama list
Show model details ollama show <model>
Copy a model ollama cp <source> <dest>
Remove a model ollama rm <model>
Push model to registry ollama push <model>

Running Models

Task Command
Run model (interactive) ollama run <model>
Run with single prompt ollama run <model> "your prompt"
Run with stdin input `echo "prompt" \
Show running models {% raw %}ollama ps
Stop a running model ollama stop <model>

In-Chat Commands

Command Action
/clear Clear chat history
/bye Exit chat
/set parameter <key> <val> Change param on the fly
/show info Show model info
/show modelfile Show current Modelfile
/show parameters Show active parameters
/help List all in-chat commands

API (REST)

Ollama runs a local server at http://localhost:11434.

# Generate (single turn)
curl http://localhost:11434/api/generate -d '{
  "model": "mymodel",
  "prompt": "Explain Docker in simple terms",
  "stream": false
}'

# Chat (multi-turn)
curl http://localhost:11434/api/chat -d '{
  "model": "mymodel",
  "messages": [
    { "role": "user", "content": "Hello!" }
  ]
}'

# List models via API
curl http://localhost:11434/api/tags

# Check running models
curl http://localhost:11434/api/ps

Enter fullscreen mode Exit fullscreen mode


8. Manage Ollama Service (systemctl)

Start / Stop / Restart

# Start Ollama service
systemctl start ollama

# Stop Ollama service
systemctl stop ollama

# Restart Ollama service
systemctl restart ollama

Enter fullscreen mode Exit fullscreen mode

Status & Logs

# Check service status
systemctl status ollama

# View live logs
journalctl -u ollama -f

# View last 50 log lines
journalctl -u ollama -n 50

Enter fullscreen mode Exit fullscreen mode

Enable / Disable on Boot

# Enable Ollama to start on boot
systemctl enable ollama

# Disable autostart
systemctl disable ollama

# Check if enabled
systemctl is-enabled ollama

Enter fullscreen mode Exit fullscreen mode


9. Gollama — Chat TUI for Ollama

Gollama is a terminal chat interface for Ollama with conversation history saved via SQLite.

Install Go (Fedora)

sudo dnf install golang -y
go version

Enter fullscreen mode Exit fullscreen mode

Install Gollama

go install github.com/gaurav-gosain/gollama@latest

# Add Go binaries to PATH
echo 'export PATH=$PATH:~/go/bin' >> ~/.bashrc
source ~/.bashrc

Enter fullscreen mode Exit fullscreen mode

Launch

gollama

Enter fullscreen mode Exit fullscreen mode

Keyboard Shortcuts

Key Action
/ k Navigate up
/ j Navigate down
Ctrl+N New chat
/ Fuzzy search chats
d Delete chat
Ctrl+C Quit