惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

酷 壳 – CoolShell
酷 壳 – CoolShell
Schneier on Security
Schneier on Security
H
Help Net Security
PCI Perspectives
PCI Perspectives
博客园 - 司徒正美
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Google Online Security Blog
Google Online Security Blog
V
Visual Studio Blog
Engineering at Meta
Engineering at Meta
Last Week in AI
Last Week in AI
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
L
LINUX DO - 最新话题
GbyAI
GbyAI
IT之家
IT之家
TaoSecurity Blog
TaoSecurity Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
J
Java Code Geeks
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
N
News and Events Feed by Topic
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
美团技术团队
T
Troy Hunt's Blog
Security Archives - TechRepublic
Security Archives - TechRepublic
Cloudbric
Cloudbric
A
About on SuperTechFans
Recorded Future
Recorded Future
Microsoft Security Blog
Microsoft Security Blog
阮一峰的网络日志
阮一峰的网络日志
H
Hacker News: Front Page
Forbes - Security
Forbes - Security
Webroot Blog
Webroot Blog
D
DataBreaches.Net
L
LangChain Blog
S
Schneier on Security
博客园_首页
S
SegmentFault 最新的问题
Apple Machine Learning Research
Apple Machine Learning Research
N
News | PayPal Newsroom
Hacker News - Newest:
Hacker News - Newest: "LLM"
爱范儿
爱范儿
量子位
T
The Exploit Database - CXSecurity.com
博客园 - 【当耐特】
T
Threatpost
The Hacker News
The Hacker News
N
News and Events Feed by Topic
罗磊的独立博客
Spread Privacy
Spread Privacy
Hacker News: Ask HN
Hacker News: Ask HN

Recent Commits to openclaw:main

test: merge chat side-result checks · openclaw/openclaw@ddd2c2a test: merge cron history checks · openclaw/openclaw@f7eb746 test: merge responsive navigation shell checks · openclaw/openclaw@c2e4b47 docs(changelog): add codex oauth fixes · openclaw/openclaw@628e6cd test: merge navigation routing cases · openclaw/openclaw@5d8cecb Tests: mock channel registry bundled fallback · openclaw/openclaw@2b08233 Secrets: avoid broad web search discovery for single plugin config · openclaw/openclaw@a464f59 test: merge config view browser checks · openclaw/openclaw@20cf511 fix(status): align oauth health with runtime · openclaw/openclaw@eed7116 test: merge chat context notice checks · openclaw/openclaw@5c2f4af feat: add macOS screen snapshots for monitor preview (#67954) thanks … · openclaw/openclaw@f377db1 fix: report shared auth scopes in hello-ok (#67810) thanks @BunsDev · openclaw/openclaw@0b6c39b Auto-reply: avoid eager bundled route fallback · openclaw/openclaw@3ea1bf4 Tests: narrow session binding contract setup · openclaw/openclaw@54e4e16 fix(macOS): enable undo/redo in webchat composer text input (#34962) · openclaw/openclaw@00951dc Tests: speed up channel setup promotion · openclaw/openclaw@82b529a Docs: refresh agent instructions · openclaw/openclaw@5775fe2 fix(auth): serialize OAuth refresh across agents to fix #26322 (#67876) · openclaw/openclaw@8e79080 test: allow ollama public surface boundary test · openclaw/openclaw@7d4f1a6 Docs: add test performance guardrails · openclaw/openclaw@89706d3 Tests: restore context-engine usage proof · openclaw/openclaw@e4c4f95 Tests: slim context engine runtime coverage · openclaw/openclaw@74c198f ci: retry failed custom checkouts · openclaw/openclaw@0ee5baf test: trim duplicate provider auth onboarding cases · openclaw/openclaw@1ffc02e matrix: fix sessions_spawn --thread subagent session spawning (#67643) · openclaw/openclaw@1ce2596 test: reduce auth choice fixture churn · openclaw/openclaw@857b9cd test: mock health status config boundaries · openclaw/openclaw@9d5ab4a test: mock onboard config io boundary · openclaw/openclaw@299694d test: mock legacy state plugin boundaries · openclaw/openclaw@2713089 test: mock channel install boundaries · openclaw/openclaw@b945248 test: mock doctor preview channel boundaries · openclaw/openclaw@b1a3ad4 test: trim doctor command hotspots · openclaw/openclaw@c66f16a test: isolate agent auth and spawn hotspots · openclaw/openclaw@9285935 test: stabilize MCP startup disposal race · openclaw/openclaw@dd9d2eb test: merge browser contract server suites · openclaw/openclaw@5817a76 test: narrow ollama provider discovery setup · openclaw/openclaw@a0d9598 build: declare qa-lab aimock runtime dependency · openclaw/openclaw@24431e5 test: speed up safe-bins exec harness · openclaw/openclaw@ee856ab test: preserve tool helpers in embedded runner mocks · openclaw/openclaw@acd86a0 refactor: move memory embeddings into provider plugins · openclaw/openclaw@77e6e4c test: reuse system-run temp fixtures · openclaw/openclaw@7e9ff0f test: trim hotspot wait overhead · openclaw/openclaw@12a59b0 Check: avoid duplicate boundary prep · openclaw/openclaw@baf11b8 test: reduce hotspot fixture overhead · openclaw/openclaw@3a59edd feat(ui): overhaul settings and slash command UX (#67819) thanks @Bun… · openclaw/openclaw@2cfb660 QA Matrix: exit cleanly on failure · openclaw/openclaw@42805d2 QA Matrix: isolate scenario coverage · openclaw/openclaw@7e659e1 Matrix: refresh crypto bootstrap state · openclaw/openclaw@94081d8 QA Lab: add provider registry · openclaw/openclaw@bb7e982 Matrix: add plugin changelog · openclaw/openclaw@4acab55 test: trim more hotspot overhead · openclaw/openclaw@f485311 test: trim remaining hotspot tests · openclaw/openclaw@6ba8626 test: narrow hotspot mocks · openclaw/openclaw@dbc8179 test: isolate gemini embedding request helpers · openclaw/openclaw@cd330f5 test: trim memory and mcp hotspots · openclaw/openclaw@fd48dfa test: slim provider registry mocks · openclaw/openclaw@2e08c77 test: harden Parallels update smoke · openclaw/openclaw@1a98090 feat: default Anthropic to Opus 4.7 · openclaw/openclaw@628b454 fix: harden node-host shell payload mutability checks · openclaw/openclaw@75c551e fix: land node-host approval binding for native binaries (#66731) (th… · openclaw/openclaw@29919bb CI: add daily schedule to CodeQL workflow (#67645) · openclaw/openclaw@69d25f5 fix(gateway): capture config hash after plugin auto-enable to prevent… · openclaw/openclaw@8c11210 fix: repair sanitized replay tool results before send (#67620) (thank… · openclaw/openclaw@c3c7a99 fix: restrict HTML timeout short-circuit to transient statuses · openclaw/openclaw@de129a6 fix: keep TUI watchdog bound to active run (#67401) (thanks @xantorres) · openclaw/openclaw@3525273 Gateway/skills: dedupe skills prefix-match + drop dead fallback on log · openclaw/openclaw@d7f489f TUI/streaming: add watchdog that resets the activity indicator after … · openclaw/openclaw@f44ab20 Agents/tool-loop: enable unknown-tool stream guard by default · openclaw/openclaw@36ed367 Gateway/skills: invalidate session skills snapshot on config write · openclaw/openclaw@b23d59a fix: classify HTML provider error pages correctly (#67642) (thanks @s… · openclaw/openclaw@e588e90 fix(skills): remove unused model-usage import (#67641) · openclaw/openclaw@55f05df docs(changelog): credit codex fix superseded PRs · openclaw/openclaw@e485f24 fix(openai-codex): normalize stale transport metadata in resolution a… · openclaw/openclaw@90801ba CI: pin Docker-related GitHub Actions (#67632) · openclaw/openclaw@f697b01 Android: modernize WebView and discovery API usage (#67627) · openclaw/openclaw@44a6e50 fix(deps): bump hono to 4.12.14 and @hono/node-server to 1.19.14 (GHS… · openclaw/openclaw@fbccc18 fix(deps): bump dompurify to 3.4.0 (#67614) · openclaw/openclaw@2c2dc00 CI: add explicit permissions to all workflow jobs (fixes code-scannin… · openclaw/openclaw@01b7516 fix: register bundled TTS providers and route overrides correctly (#6… · openclaw/openclaw@6ea3cdd fix: align host tilde paths with OS home (#62804) (thanks @stainlu) · openclaw/openclaw@ecfaf64 fix: flush creds queue before reconnect socket open (#67464) (thanks … · openclaw/openclaw@405c63f fix: strip standalone <function> tool call tags from visible text (#6… · openclaw/openclaw@78df859 fix(agents): preserve cli session metadata before transcript persist … · openclaw/openclaw@898fd04 docs(changelog): move cli transcript entry · openclaw/openclaw@c1817c6 fix(agents): normalize cli transcript api field · openclaw/openclaw@3a3fae0 docs(changelog): note cli transcript persistence · openclaw/openclaw@6c343f1 fix(agents): persist cli transcript turns · openclaw/openclaw@b8ef507 fix(msteams): harden security-sensitive flows (#65841) · openclaw/openclaw@c56b56e [Dashboard] Fix exec approval modal overflow for long command content… · openclaw/openclaw@053c5b0 Docs: remove QA changelog entry · openclaw/openclaw@7fd5771 QA: fix private runtime source loading (#67428) · openclaw/openclaw@d5933af docs(gateway): correct protocol.md schema path, hello-ok example, aut… · openclaw/openclaw@489404d CI: pin Node 22 runners to 22.18.0 · openclaw/openclaw@4ffa621 models.authStatus: normalize provider ids + tighten env-backed escape… · openclaw/openclaw@f2fdb9d Update CHANGELOG.md · openclaw/openclaw@7694a92 test(parallels): clean up npm update guard jobs · openclaw/openclaw@045ea7b Plugins: prefer scanDir override paths · openclaw/openclaw@b2974da fix(dreaming): default storage.mode to "separate" so phase blocks sto… · openclaw/openclaw@8c392f0 fix(memory-core): skip dreaming transcript ingestion via session stor… · openclaw/openclaw@a1b01f0 fix: dedupe replayed exec.finished node events (#67281) · openclaw/openclaw@5dcf526
Extensions/lmstudio: back off inference preload after consecutive fai… · openclaw/openclaw@b555214
2026-04-16 · via Recent Commits to openclaw:main

@@ -15,6 +15,68 @@ type StreamModel = Parameters<StreamFn>[0];

15151616

const preloadInFlight = new Map<string, Promise<void>>();

171718+

/**

19+

* Cooldown state for the LM Studio preload endpoint.

20+

*

21+

* Without this, every chat request would retry preload ~every 2s even when

22+

* LM Studio has rejected the load (for example the memory guardrail will keep

23+

* rejecting until the user adjusts the setting or frees RAM). That produced

24+

* hundreds of `LM Studio inference preload failed` WARN lines per hour without

25+

* actually helping the user. The cooldown applies an exponential backoff per

26+

* preloadKey and, while the cooldown is active, the wrapper skips the preload

27+

* step entirely and proceeds directly to streaming — the model is often

28+

* already loaded from the user's LM Studio UI, so inference can succeed even

29+

* when preload keeps being rejected.

30+

*/

31+

type PreloadCooldownEntry = {

32+

untilMs: number;

33+

consecutiveFailures: number;

34+

};

35+36+

const preloadCooldown = new Map<string, PreloadCooldownEntry>();

37+38+

const PRELOAD_BACKOFF_BASE_MS = 5_000;

39+

const PRELOAD_BACKOFF_MAX_MS = 300_000;

40+41+

function computePreloadBackoffMs(consecutiveFailures: number): number {

42+

const exponent = Math.max(0, consecutiveFailures - 1);

43+

const raw = PRELOAD_BACKOFF_BASE_MS * 2 ** exponent;

44+

return Math.min(PRELOAD_BACKOFF_MAX_MS, raw);

45+

}

46+47+

function recordPreloadSuccess(preloadKey: string): void {

48+

preloadCooldown.delete(preloadKey);

49+

}

50+51+

function recordPreloadFailure(preloadKey: string, now: number): PreloadCooldownEntry {

52+

const existing = preloadCooldown.get(preloadKey);

53+

const consecutiveFailures = (existing?.consecutiveFailures ?? 0) + 1;

54+

const entry: PreloadCooldownEntry = {

55+

consecutiveFailures,

56+

untilMs: now + computePreloadBackoffMs(consecutiveFailures),

57+

};

58+

preloadCooldown.set(preloadKey, entry);

59+

return entry;

60+

}

61+62+

function isPreloadCoolingDown(preloadKey: string, now: number): PreloadCooldownEntry | undefined {

63+

const entry = preloadCooldown.get(preloadKey);

64+

if (!entry) {

65+

return undefined;

66+

}

67+

if (entry.untilMs <= now) {

68+

preloadCooldown.delete(preloadKey);

69+

return undefined;

70+

}

71+

return entry;

72+

}

73+74+

/** Test-only hook for clearing preload cooldown state between cases. */

75+

export function __resetLmstudioPreloadCooldownForTest(): void {

76+

preloadCooldown.clear();

77+

preloadInFlight.clear();

78+

}

79+1880

function normalizeLmstudioModelKey(modelId: string): string {

1981

const trimmed = modelId.trim();

2082

if (trimmed.toLowerCase().startsWith("lmstudio/")) {

@@ -131,29 +193,67 @@ export function wrapLmstudioInferencePreload(ctx: ProviderWrapStreamFnContext):

131193

modelKey,

132194

requestedContextLength,

133195

});

196+197+

const cooldownEntry = isPreloadCoolingDown(preloadKey, Date.now());

134198

const existing = preloadInFlight.get(preloadKey);

135-

const preloadPromise =

199+

const preloadPromise: Promise<void> | undefined =

136200

existing ??

137-

ensureLmstudioModelLoadedBestEffort({

138-

baseUrl: resolvedBaseUrl,

139-

modelKey,

140-

requestedContextLength,

141-

options,

142-

ctx,

143-

modelHeaders: resolveModelHeaders(model),

144-

}).finally(() => {

145-

preloadInFlight.delete(preloadKey);

146-

});

147-

if (!existing) {

148-

preloadInFlight.set(preloadKey, preloadPromise);

149-

}

201+

(cooldownEntry

202+

? undefined

203+

: (() => {

204+

const created = ensureLmstudioModelLoadedBestEffort({

205+

baseUrl: resolvedBaseUrl,

206+

modelKey,

207+

requestedContextLength,

208+

options,

209+

ctx,

210+

modelHeaders: resolveModelHeaders(model),

211+

})

212+

.then(

213+

() => {

214+

recordPreloadSuccess(preloadKey);

215+

},

216+

(error) => {

217+

const entry = recordPreloadFailure(preloadKey, Date.now());

218+

throw Object.assign(new Error("preload-failed"), {

219+

cause: error,

220+

consecutiveFailures: entry.consecutiveFailures,

221+

cooldownMs: entry.untilMs - Date.now(),

222+

});

223+

},

224+

)

225+

.finally(() => {

226+

preloadInFlight.delete(preloadKey);

227+

});

228+

preloadInFlight.set(preloadKey, created);

229+

return created;

230+

})());

150231151232

return (async () => {

152-

try {

153-

await preloadPromise;

154-

} catch (error) {

155-

log.warn(

156-

`LM Studio inference preload failed for "${modelKey}"; continuing without preload: ${String(error)}`,

233+

if (preloadPromise) {

234+

try {

235+

await preloadPromise;

236+

} catch (error) {

237+

const annotated = error as {

238+

cause?: unknown;

239+

consecutiveFailures?: number;

240+

cooldownMs?: number;

241+

};

242+

const cause = annotated.cause ?? error;

243+

const failures = annotated.consecutiveFailures ?? 1;

244+

const cooldownSec = Math.max(

245+

0,

246+

Math.round((annotated.cooldownMs ?? 0) / 1000),

247+

);

248+

log.warn(

249+

`LM Studio inference preload failed for "${modelKey}" (${failures} consecutive failure${

250+

failures === 1 ? "" : "s"

251+

}, next preload attempt skipped for ~${cooldownSec}s); continuing without preload: ${String(cause)}`,

252+

);

253+

}

254+

} else if (cooldownEntry) {

255+

log.debug(

256+

`LM Studio inference preload for "${modelKey}" skipped while backoff active (${cooldownEntry.consecutiveFailures} prior failures)`,

157257

);

158258

}

159259

// LM Studio uses OpenAI-compatible streaming usage payloads when requested via