惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

美团技术团队
Blog — PlanetScale
Blog — PlanetScale
阮一峰的网络日志
阮一峰的网络日志
M
MIT News - Artificial intelligence
月光博客
月光博客
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
U
Unit 42
博客园_首页
WordPress大学
WordPress大学
H
Hackread – Cybersecurity News, Data Breaches, AI and More
J
Java Code Geeks
F
Fortinet All Blogs
腾讯CDC
罗磊的独立博客
IT之家
IT之家
I
InfoQ
V
V2EX
博客园 - 叶小钗
A
About on SuperTechFans
Y
Y Combinator Blog
C
Check Point Blog
量子位
Martin Fowler
Martin Fowler
Vercel News
Vercel News

GitLab

Rate limits on GitLab.com are changing See who spent your AI credits and set fair caps per team New MCP tools help platform teams scale automation safely GitLab Duo CLI takes a task from goal to done When to use SAST versus an LLM security scanner GitLab Dedicated: Compliance for a new regulatory era How to calculate DevOps platform total cost of ownership GitLab Critical Patch Release: 19.3.2, 19.2.6, 19.1.8 Co-Create: Building GitLab with our users Prepare for the Cyber Resilience Act Bring your own model to GitLab Duo Self-Hosted with Microsoft Foundry GPT-6 Astra on GitLab: Faster runs, fewer tokens used GitLab’s internal playbook to foster AI-fluent technical teams Critical remote code execution in vm2, a widely used Node.js sandbox library GitLab compliance frameworks: Adhere to SOC 2 in minutes How to recognize your team with GitLab Achievements Making room for what GitLab Patch Release: 19.3.1, 19.2.5, 19.1.7 Git was built for humans — agents need an upgrade Scale software delivery without owning the runner fleet When code is abundant When your backlog outgrows your team, GitLab scales remediation Run agentic software delivery inside the boundaries you already trust Build custom flows in minutes with the Flow Creator agent GitLab 19.3 release notes From chaos to context: Building an AI dev workflow From OpenTofu to Argo CD: GitLab as your AWS control plane Avoid the massive end-to-end tax of default full history clones GitLab Critical Patch Release: 19.2.4, 19.1.6, 19.0.8, 18.11.11 Critical remote code execution in Serena, a popular MCP coding agent
Optimize your team
Brittany Lutz · 2026-09-17 · via GitLab

There’s no single best model for every software development task. Implementing a new feature, diagnosing a failed pipeline, and resolving security vulnerabilities all place different demands on the model handling them. GitLab Duo Agent Platform is expanding GitLab-managed model choice with three hosted open weight models: Kimi K3, GLM 5.3, and MiniMax M3.

Together with the frontier models already available in GitLab, your team now has more control over how you optimize for quality, latency, and cost, tuned to the needs of each workload. Point a hard task at Kimi K3 or GLM 5.3, which outperformed comparable frontier models in internal testing and costs less per call, or hand routine, high-volume work to MiniMax M3. Either way, your team gets up to 4x more calls per GitLab Credit than some comparable frontier models, with more AI model options to match cost to task complexity.

GitLab Transcend returns in October

Coding agents are increasing your speed of development, but your reviews, security policies, and release cycles still have to keep pace. Our Transcend event on October 6 will demonstrate how GitLab is helping teams close that gap and explore what it takes to carry the speed of agentic AI across the software lifecycle.

Register for the livestream today!

Why one model has been the easy way out

Different tasks in agentic development need different things from an AI model. A long-running refactor needs a large context window and deeper reasoning, while a routine, high-volume task is often better served by a faster, more cost-efficient model. Optimizing for one task type means giving up ground on the other.

As teams work through the technical tradeoff, whether they can actually access a new model emerges as an additional governance challenge. In regulated environments, every model and the infrastructure it runs on has to clear security, compliance, and internal review before the team can use it. Reviews and approvals move slowly enough that many teams standardize on one approved model, even when it isn’t the best fit for every task.

The cost of that compromise compounds. If your team standardizes on a higher-cost model to cover its hardest tasks, that same model will end up tackling routine work that could have been handled by a faster, lower-cost model. As agent workloads make repeated model calls across long-running, multi-step work, the gap between cost and performance grows.

The cost-performance advantage of open weight models

Open weight models turn model choice into a lever for balancing cost with performance. They can potentially lower the inference costs of many agentic tasks, with performance in line with some frontier models.

The new open weight models hosted in GitLab offer notably more calls with one credit than many comparable frontier models already available on Duo Agent Platform. A call is a single request GitLab Duo Agent Platform sends to a model, and one agent action or chat message can trigger several calls.

Kimi K3 outperforms the default model behind most GitLab Duo Agent Platform features in our internal testing, at 1.82 calls per credit, just under the default model’s 2 calls per credit. GLM 5.3 gets 5 calls per credit while also delivering strong performance in internal testing, and MiniMax M3 builds on that cost efficiency at 8 calls per credit.

Together, they give your team another lever for balancing cost and performance, task by task. For more information on credit multipliers and a breakdown of credit usage per model, visit the GitLab Credits models page.

Each of the hosted open weight models now available in GitLab Duo Agent Platform brings its own strengths.

  • Kimi K3 (Moonshot AI) is the company's newest flagship model, giving teams room to work across large codebases or long-running agent sessions.
  • GLM 5.3 (Z.ai) is built for long-horizon coding tasks, giving teams a model tuned to hold context and accuracy across extended agent sessions.
  • MiniMax M3 (MiniMax) is designed to keep compute costs down on long-running work and tuned for high-volume, multi-step agent work, making it a strong option when cost efficiency matters at scale.

Watch this demonstration of available open weight models:

Model choice with built-in control

With a broad roster spanning open weight and frontier GitLab-managed models, your Duo Agent Platform implementation now has more freedom to optimize model selection for each task.

Your admins maintain centralized control over which models your teams can use, built directly into your software development workflows. While GitLab selects default models based on performance, the owner of your top-level group can set a different default model for each feature and curate which models teams can choose from, settings that apply consistently across every child group and project.

You can align model selection with your security, compliance, and infrastructure requirements through GitLab-managed, self-hosted, or hybrid AI deployment options. GitLab-managed models run entirely within GitLab's environment, with no infrastructure for your team to stand up or maintain. Self-hosted models run on infrastructure your team manages, keeping model traffic inside your own network. Hybrid deployments let your team mix both, running some models through GitLab and others on your own infrastructure, based on what each workload requires.

Before a new model joins the GitLab Duo Agent Platform roster, we evaluate it against internal performance and quality requirements. GitLab-managed models carry their own built-in assurance. Before selecting a vendor to host models for Duo Agent Platform features, GitLab puts it through a third-party risk management process to verify it meets a defined security bar. That includes a review of the vendor's security program and controls, such as access management, data governance, and risk management practices, along with third-party security attestations like a SOC 2 Type 2 report or active ISO 27001 certification, and periodic penetration testing with risk-based vulnerability remediation.

All four hosted open weight models are served through Fireworks AI, which meets that bar. For GitLab Duo Agent Platform requests, GitLab maintains a zero data retention policy with Fireworks, meaning model input and output data is discarded immediately after each response and is not stored for abuse monitoring.

With that governance and security foundation in place, your team can bring open weight models into your workflows with the same confidence you'd expect from any GitLab-managed model.

Start building with open weight models

You can now select Kimi K3, GLM 5.3, and MiniMax M3 as optional models in GitLab Duo Agent Platform, or set one as the default model for a feature so your team uses it automatically. For a deeper look into model support in Duo Agent Platform and guidance on model selection, read the GitLab AI model documentation.

New to Duo Agent Platform? Start a free trial. Already on Premium or Ultimate? Turn on Duo Agent Platform.