GitHub - Kareem-Rashed/rubric-eval: Independent framework to test, benchmark, and evaluate LLMs & AI agents locally. - 惯性聚合

推荐订阅源

Security Latest

Threat Intelligence Blog | Flashpoint

Attack and Defense Labs

Security Archives - TechRepublic

News and Events Feed by Topic

Check Point Blog

LINUX DO - 最新话题

Hacker News: Ask HN

Privacy International News Feed

Fortinet All Blogs

Application and Cybersecurity Blog

Threat Research - Cisco Blogs

阮一峰的网络日志

Cyber Attacks, Cyber Crime and Cyber Security

博客园 - 司徒正美

Visual Studio Blog

The Hacker News

CXSECURITY Database RSS Feed - CXSecurity.com

DataBreaches.Net

Privacy & Cybersecurity Law Blog

Google Developers Blog

TaoSecurity Blog

The Blog of Author Tim Ferriss

Engineering at Meta

The Exploit Database - CXSecurity.com

Microsoft Security Blog

酷壳 – CoolShell

Lohrmann on Cybersecurity

cs.CL updates on arXiv.org

Schneier on Security

Netflix TechBlog - Medium

Tor Project blog

MIT News - Artificial intelligence

Proofpoint News Feed

freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More

博客园 - Franky

Google DeepMind News

Proofpoint News Feed

Hacker News - Newest: "LLM"

GitHub - Anhydrite/doc-torn: Project that provides structured documentation skills for AI coding agents. GitHub - kmdupr33/fks2g: A CLI for generating LLM-backed metrics for deciding how closely to review code If an LLM is too expensive it won't be next year StepStone: LLM-Based GPU Kernel Driver Fuzzing via User-Space Libraries [pdf] GitHub - AssimilatedHuman/LLM-Inquisitor: Evaluating AI behaviour under real‑world work conditions to surface issues before they become problems. LLM INQUISITOR identifies failures (drift, instability etc) by observing AI during normal tasks — a tool the industry desperately needs to stem the 85% failure rate. Includes Quick Start, Practitioner’s Guide and Methodology. Creating another MCP server, but this one is for research A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents Sator Arepo - a Hugging Face Space by akolpakov Customizing an LLM for Enterprise Software Engineering Barron AI Solutions Evaluating job search ranking with LLM judged NDCG GitHub - quadracollision/llmisp: JSON AST > Clojure Parity Contracts for Polyglot LLM Commerce: A Case Study GitHub - ndom91/llama-dash: The operations layer for your local LLM stack Ask HN: What's your go-to LLM for coding? How do you reduce LLM spam in PR reviews? Ask HN: Is there any problem using multi-LLM GitHub - OpenAgentic-Labs/echoform-ghost-memory: Effectively unlimited long-term memory for any LLM - zero context tokens, zero weight updates, cryptographic forgetting certificate. GitHub - robertoranon/tokoro: A toolbox for building event publish & discovery web sites, apps, feeds, and more GitHub - sermakarevich/chunker: Agentic approach to chunking a document A new EDIT tool for LLM agents MLSys @ WukLab - Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism Managing metadata is essential in LLM world Fixing LLM Writing with Distribution Fine Tuning The local shape of LLM stable regions GitHub - msunda17/impactarbiter-cli The Infrastructure Behind Making Local LLM Agents Useful PostgreSQL ext makes LLM available as an index for similarity searches,inference GitHub - Tetrahedroned/Agent-Braille: Deterministic 8-bit machine-to-machine protocol for AI agent state. ~92% fewer state-tracking tokens on real Claude Code sessions, a proven single-bit-error-safe command code, fully reproducible. Tell HN: Writing an LLM critique/takedown? – Do not use an LLM to write it 🌱 an LLM models our worst behavior Prompt eval cues predicted refusal shifts across 32k LLM rollouts Ask HN: Is Java the ideal language for LLM-assisted coding? Log in | AI Foundry LLM tracing with MLflow AI Gateway Gert Labs - Games for Machines The LLM Looked Smart. The Metrics Disagreed – tiago.rio.br The Four Horsemen of the LLM Apocalypse GitHub - piqoni/piqo-extension: A good interface is invisible Intro to TLA+ for the LLM Era: Prompt Your Way to Victory Give every tool LLM wiki and bypass Claude Code SSH Throttle The Ultimate LLM Fine-Tuning Guide Ask HN: What LLM models are you using and why? Five Agents, One Browser: Werewolf on Quack + DuckDB LLM models are not ready for orchestrating many agents ClickBook — Offline AI eReader - Apps on Google Play Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention We Built SynapseKit: The Truth About Production LLM Frameworks GitHub - albedan/ai-ml-gpu-bench: A suite to benchmark CPU/GPU Python performance in training ML models and running local LLMs GitHub - chopratejas/headroom: Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server. Most Meaningful Dates on the Web and for an LLM GitHub - Andyyyy64/whichllm: Find the local LLM that actually runs — and performs best — on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly. GitHub - krellixlabs/llm-reasoning-research: Curated, annotated research on reasoning gaps in large language models — temporal reasoning, causal reasoning, and beyond. Agentic evals or LLM as a judge? considering cost, time and quality Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces Add an LLM policy for `rust-lang/rust` by jyn514 · Pull Request #1040 · rust-lang/rust-forge GitHub - nimeshnayaju/markdown-parser: A streaming-capable markdown parser, written in TypeScript Dragos Documents First LLM-Assisted Strike on Water Infrastructure in Mexico Alchemize: PyMC's model to replace Stan/PyMC, etc. with an LLM LLM Witch Hunts are getting F'in Irritating bliki: Interrogatory LLM GitHub - EvanPaules/ctx-opt: Context Window Optimizer Show HN: Local-first Kubernetes YAML visualizer (no server, no LLM) Why Ruby Is the Better Language for LLM-Powered Development Paper page - Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training State media control shapes LLM behaviour by influencing training data Small Model Forensics GitHub - achaljhawar/1rok: Multi-LLM trading harness. GitHub - crawshaw/yeah: yeah: LLM-powered yes/no CLI tool Predicting Rare LLM Failures with 30× Fewer Rollouts — LessWrong Mechanism Design for Quality-Preserving LLM Advertising I tried to put an on-device LLM in an iOS Share Extension. It didn't fit GitHub - mentasystems/gox: Strict static analyzer for Go — designed for LLM-written code. Zero external linter dependencies. GitHub - torrix-ai/install Beyond Similarity Search: Tenure and the Case for Structured Belief State in LLM Memory Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference Hi-Vis: one-shot jailbreak disguised as LLM "software patch" reaching 100% ASR Loading/running every LLM with 4M ctx in 3 clicks GLiGuard: 16x Faster Safety Moderation with a Small Language Model - Pioneer AI by Fastino Labs Are LLM Useful for Solo Founders GitHub - aegis-dq/aegis-dq: Open, audit-grade agentic data quality framework with portable industry packs GitHub - jbethune777/ninchi: Human Software Accountability LLM Research Knowledge Graph — 905 Synthesis Insights. Adrian Chan GitHub - kerv/laze The Close Reader — AI Linguistic Analysis Silent-Bench: Cryptographically-Attested Forensic Auditing of LLM API Gateways GitHub - fu5ha/pi-treebase: Interactive-rebase style tree navigation and compaction for pi Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology Local LLM Proxy: Turn Idle LLM Compute into Universal Credits RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems GitHub - hzw1199/CyberMe-LLM-Wiki: A faithful llm-wiki implementation with agent-maintained Markdown knowledge bases and Wikipedia-style web browsing. Can you help reconcile my first/second-hand LLM Experience with HN's Experience? Graft – semantic memory for AI agents, without the LLM Using LLM in the shebang line of a script The Clipboard Pattern — A Better Way to Compose AI Agents TIL: Using LLM in the shebang line of a script GitHub - uw-syfi/vibe-serve: Can AI Agents Build Bespoke LLM Serving Systems? CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure GitHub - aidarbek/genz-qwen: Post-training Qwen2.5-0.5B-Instruct to talk like GenZ

GitHub - Kareem-Rashed/rubric-eval: Independent framework to test, benchmark, and evaluate LLMs & AI agents locally.

kareemrashed · 2026-06-13 · via Hacker News - Newest: "LLM"

此内容由惯性聚合(RSS阅读器)自动聚合整理，仅供阅读参考。原文来自 — 版权归原作者所有。