惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
Docker
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - Franky
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
A
About on SuperTechFans
博客园 - 【当耐特】
Microsoft Security Blog
Microsoft Security Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
The GitHub Blog
The GitHub Blog
雷峰网
雷峰网
博客园_首页
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
IT之家
IT之家
博客园 - 叶小钗
Google DeepMind News
Google DeepMind News
aimingoo的专栏
aimingoo的专栏
博客园 - 聂微东
B
Blog RSS Feed
H
Help Net Security
Recent Announcements
Recent Announcements
阮一峰的网络日志
阮一峰的网络日志
D
DataBreaches.Net
L
LangChain Blog
Vercel News
Vercel News

Recent Commits to openclaw:main

test: merge chat side-result checks · openclaw/openclaw@ddd2c2a test: merge cron history checks · openclaw/openclaw@f7eb746 test: merge responsive navigation shell checks · openclaw/openclaw@c2e4b47 docs(changelog): add codex oauth fixes · openclaw/openclaw@628e6cd test: merge navigation routing cases · openclaw/openclaw@5d8cecb Tests: mock channel registry bundled fallback · openclaw/openclaw@2b08233 Secrets: avoid broad web search discovery for single plugin config · openclaw/openclaw@a464f59 test: merge config view browser checks · openclaw/openclaw@20cf511 fix(status): align oauth health with runtime · openclaw/openclaw@eed7116 feat: add macOS screen snapshots for monitor preview (#67954) thanks … · openclaw/openclaw@f377db1 fix: report shared auth scopes in hello-ok (#67810) thanks @BunsDev · openclaw/openclaw@0b6c39b Auto-reply: avoid eager bundled route fallback · openclaw/openclaw@3ea1bf4 Tests: narrow session binding contract setup · openclaw/openclaw@54e4e16 fix(macOS): enable undo/redo in webchat composer text input (#34962) · openclaw/openclaw@00951dc Tests: speed up channel setup promotion · openclaw/openclaw@82b529a Docs: refresh agent instructions · openclaw/openclaw@5775fe2 fix(auth): serialize OAuth refresh across agents to fix #26322 (#67876) · openclaw/openclaw@8e79080 test: allow ollama public surface boundary test · openclaw/openclaw@7d4f1a6 Docs: add test performance guardrails · openclaw/openclaw@89706d3 Tests: restore context-engine usage proof · openclaw/openclaw@e4c4f95 Tests: slim context engine runtime coverage · openclaw/openclaw@74c198f ci: retry failed custom checkouts · openclaw/openclaw@0ee5baf test: trim duplicate provider auth onboarding cases · openclaw/openclaw@1ffc02e matrix: fix sessions_spawn --thread subagent session spawning (#67643) · openclaw/openclaw@1ce2596 test: reduce auth choice fixture churn · openclaw/openclaw@857b9cd test: mock health status config boundaries · openclaw/openclaw@9d5ab4a test: mock onboard config io boundary · openclaw/openclaw@299694d test: mock legacy state plugin boundaries · openclaw/openclaw@2713089 test: mock channel install boundaries · openclaw/openclaw@b945248 test: mock doctor preview channel boundaries · openclaw/openclaw@b1a3ad4
fix: stabilize qa lab mock suite · openclaw/openclaw@903308d
steipete · 2026-04-24 · via Recent Commits to openclaw:main

@@ -207,15 +207,15 @@ refs and write a judged Markdown report:

207207208208

```bash

209209

pnpm openclaw qa character-eval \

210-

--model openai-codex/gpt-5.5,thinking=xhigh \

210+

--model openai/gpt-5.4,thinking=medium,fast \

211211

--model openai/gpt-5.2,thinking=xhigh \

212212

--model openai/gpt-5,thinking=xhigh \

213213

--model anthropic/claude-opus-4-6,thinking=high \

214214

--model anthropic/claude-sonnet-4-6,thinking=high \

215215

--model zai/glm-5.1,thinking=high \

216216

--model moonshot/kimi-k2.5,thinking=high \

217217

--model google/gemini-3.1-pro-preview,thinking=high \

218-

--judge-model openai-codex/gpt-5.5,thinking=xhigh,fast \

218+

--judge-model openai/gpt-5.4,thinking=xhigh,fast \

219219

--judge-model anthropic/claude-opus-4-6,thinking=high \

220220

--blind-judge-models \

221221

--concurrency 16 \

@@ -227,13 +227,13 @@ scenarios should set the persona through `SOUL.md`, then run ordinary user turns

227227

such as chat, workspace help, and small file tasks. The candidate model should

228228

not be told that it is being evaluated. The command preserves each full

229229

transcript, records basic run stats, then asks the judge models in fast mode with

230-

`xhigh` reasoning to rank the runs by naturalness, vibe, and humor.

230+

`xhigh` reasoning where supported to rank the runs by naturalness, vibe, and humor.

231231

Use `--blind-judge-models` when comparing providers: the judge prompt still gets

232232

every transcript and run status, but candidate refs are replaced with neutral

233233

labels such as `candidate-01`; the report maps rankings back to real refs after

234234

parsing.

235-

Candidate runs default to `high` thinking, with `xhigh` for OpenAI models that

236-

support it. Override a specific candidate inline with

235+

Candidate runs default to `high` thinking, with `medium` for GPT-5.4 and `xhigh`

236+

for older OpenAI eval refs that support it. Override a specific candidate inline with

237237

`--model provider/model,thinking=<level>`. `--thinking <level>` still sets a

238238

global fallback, and the older `--model-thinking <provider/model=level>` form is

239239

kept for compatibility.

@@ -247,12 +247,12 @@ Candidate and judge model runs both default to concurrency 16. Lower

247247

`--concurrency` or `--judge-concurrency` when provider limits or local gateway

248248

pressure make a run too noisy.

249249

When no candidate `--model` is passed, the character eval defaults to

250-

`openai-codex/gpt-5.5`, `openai/gpt-5.4`, `openai/gpt-5.2`, `anthropic/claude-opus-4-6`,

250+

`openai/gpt-5.4`, `openai/gpt-5.2`, `openai/gpt-5`, `anthropic/claude-opus-4-6`,

251251

`anthropic/claude-sonnet-4-6`, `zai/glm-5.1`,

252252

`moonshot/kimi-k2.5`, and

253253

`google/gemini-3.1-pro-preview` when no `--model` is passed.

254254

When no `--judge-model` is passed, the judges default to

255-

`openai-codex/gpt-5.5,thinking=xhigh,fast` and

255+

`openai/gpt-5.4,thinking=xhigh,fast` and

256256

`anthropic/claude-opus-4-6,thinking=high`.

257257258258

## Related docs