惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
爱范儿
爱范儿
罗磊的独立博客
V
V2EX
量子位
Last Week in AI
Last Week in AI
Hugging Face - Blog
Hugging Face - Blog
博客园 - 司徒正美
Jina AI
Jina AI
博客园 - 叶小钗
小众软件
小众软件
博客园 - 【当耐特】
Y
Y Combinator Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
Tailwind CSS Blog
博客园 - 聂微东
Microsoft Security Blog
Microsoft Security Blog
美团技术团队
P
Proofpoint News Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
有赞技术团队
有赞技术团队
MongoDB | Blog
MongoDB | Blog
Recent Announcements
Recent Announcements
酷 壳 – CoolShell
酷 壳 – CoolShell

MachineLearningMastery.com

The Roadmap to Mastering Voice Agents - MachineLearningMastery.com Treating Prompt Templates as Hyperparameters in Scikit-LLM GridSearchCV - MachineLearningMastery.com A Gentle Introduction to Model Distillation - MachineLearningMastery.com Fine-Tuning Agentic AI: A Practical Guide - MachineLearningMastery.com How to Combine Traditional Machine Learning with Agentic Reasoning - MachineLearningMastery.com Chain of Thought vs. Tree of Thoughts: Which is Best for AI Agents? - MachineLearningMastery.com Dataclasses for Structured Application Data - MachineLearningMastery.com Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It - MachineLearningMastery.com AI Agent Memory Design: What Works and What Doesn’t 3 Ways to Enhance Your AI Model's Interpretability - MachineLearningMastery.com Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline - MachineLearningMastery.com Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces - MachineLearningMastery.com Learn Vectorized Thinking in Python Through Examples - MachineLearningMastery.com Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral - MachineLearningMastery.com Integrating Agentic AI with Existing Machine Learning Pipelines - MachineLearningMastery.com How to Build a Robust RAG System with Minimal Resources - MachineLearningMastery.com Managing Small Context Windows in Language Models - MachineLearningMastery.com 7 Regression Tests Every AI Agent Should Pass Before Deploy - MachineLearningMastery.com Understanding the Role of Latent Space in Machine Learning Models - MachineLearningMastery.com Retrieval vs. Memory in Agentic AI System 7 Async Patterns for Running Agents Concurrently in Python - MachineLearningMastery.com Prompt Caching vs. Fine-Tuning: A Cost and Latency Decision Framework - MachineLearningMastery.com Identifying Token Costs Hiding in Your Agentic Loop - MachineLearningMastery.com Designing AI Agents That Can Self-Correct - MachineLearningMastery.com 7 Chunking Strategies That Decide Whether Your RAG Works - MachineLearningMastery.com Measuring Performance of Transformer Inference - MachineLearningMastery.com Static vs. Dynamic vs. Continuous Batching in LLM Inference Decoding Strategies and Output Control - MachineLearningMastery.com Using a Transformer Model: From Training to Inference The End-to-End Agentic AI Pipeline
Versioning and Tracking Scikit-LLM Experiments - MachineL...
Iván Palomares Carrascosa · 2026-09-09 · via MachineLearningMastery.com

In this article, you will learn how to build, track, compare, and register scikit-learn pipelines that integrate large language models using Scikit-LLM and MLflow.

Topics we will cover include:

  • How to configure Scikit-LLM and MLflow to support local large language model execution and experiment tracking.
  • How to log multiple pipeline versions across different large language model backends and compare them using MLflow’s tracking API.
  • How to promote the best-performing pipeline from a tracked experiment into MLflow’s Model Registry for deployment.

Versioning and Tracking Scikit-LLM Experiments

Introduction

Registering, versioning, and comparing scikit-learn-like pipelines that integrate large language models (LLMs) can be made easy with the aid of two cornerstone tools: the Scikit-LLM library and MLflow, an open-source framework for managing the end-to-end lifecycle of machine learning projects.

This article demonstrates the steps to build, log, compare, and register scikit-learn pipelines revolving around LLMs using Scikit-LLM and MLflow. The code shown and described in detail below is designed with the primary purpose of ensuring model versioning and reproducibility across LLM backend updates — a frequent process in real settings that can quickly escalate.

Setup and Initial Configurations

If you haven’t done so before, or if you are running this code on a cloud-based notebook like Google Colab, the first step is to install the key libraries you will need:

pip install "scikit-llm[gpt4all]" mlflow

Make sure to use the extra option in brackets when installing scikit-llm to avoid compatibility issues.

Now, we initialize the configuration of Scikit-LLM with dummy credentials that enable local gpt4all model execution. Meanwhile, the MLflow model registry —the key resource where models will be versioned— relies on a database backend, which is also configured in the code below. Moreover, we initialize an MLflow tracking experiment named "Scikit-LLM-Versioning". Lastly, we define a small labeled dataset for zero-shot classification (more about this LLM-driven form of classification task here).

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

20

21

import mlflow

import mlflow.sklearn

from sklearn.pipeline import Pipeline

from skllm.config import SKLLMConfig

from skllm.models.gpt.classification.zero_shot import ZeroShotGPTClassifier

# 1. Dummy keys required by Scikit-LLM for local gpt4all execution

SKLLMConfig.set_openai_key("local-execution-key")

SKLLMConfig.set_openai_org("local-execution-org")

# 2. Database backend required for the MLflow Model Registry

mlflow.set_tracking_uri("sqlite:///mlflow.db")

mlflow.set_experiment("Scikit-LLM-Versioning")

# Sample dataset for zero-shot classification

X_train = [

    "The application crashed immediately.",

    "Absolutely wonderful support team!",

    "It works fine but is a bit slow."

]

y_train = ["bug", "praise", "feedback"]

Logging the Baseline and Upgraded Pipelines

This is where the real fun starts. We initialize a baseline pipeline that trains a zero-shot classification model using a lightweight pre-trained LLM.

The with block that follows, named after the Orca Mini model selected, enables tracking of the LLM backend type and the model file string as environment parameters, thereby fostering reproducibility. A "cloudpickle" serialization format (a variant of the classic pickle, or .pkl for short, used in smaller machine learning models) is used to log the pipeline. Understanding this block is key to leveraging LLM versioning in MLflow for subsequent experiment tracking. Once execution completes, it outputs a unique MLflow run ID.

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

20

LLM_V1 = "gpt4all::orca-mini-3k-71m-q4_0.gguf"

pipeline_v1 = Pipeline([

    ('llm_classifier', ZeroShotGPTClassifier(model=LLM_V1))

])

with mlflow.start_run(run_name="Baseline_Orca_Mini") as run_v1:

    mlflow.log_param("llm_backend", "gpt4all")

    mlflow.log_param("llm_model_file", LLM_V1)

    pipeline_v1.fit(X_train, y_train)

    # Override strict skops type checking with cloudpickle

    mlflow.sklearn.log_model(

        pipeline_v1,

        "model",

        serialization_format="cloudpickle"

    )

    print(f"V1 Logged - Run ID: {run_v1.info.run_id}")

Output excerpt:

V1 Logged - Run ID: 0852aaec23364725b433f09973a3d911

Next, let’s suppose we create a secondary, upgraded pipeline based on a heavier LLM to demonstrate MLflow’s model-swapping capabilities. Specifically, we now target "gpt4all::ggml-model-gpt4all-falcon-q4_0.bin", which makes for a realistic backend upgrade. The code below isolates this new pipeline inside a separate MLflow run named "Upgraded_Falcon". Everything else is done just as before: pipeline parameterization, model fitting, and logging —just in a distinct MLflow run, yielding a new unique ID.

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

20

LLM_V2 = "gpt4all::ggml-model-gpt4all-falcon-q4_0.bin"

pipeline_v2 = Pipeline([

    ('llm_classifier', ZeroShotGPTClassifier(model=LLM_V2))

])

with mlflow.start_run(run_name="Upgraded_Falcon") as run_v2:

    mlflow.log_param("llm_backend", "gpt4all")

    mlflow.log_param("llm_model_file", LLM_V2)

    pipeline_v2.fit(X_train, y_train)

    # Override strict skops type checking with cloudpickle

    mlflow.sklearn.log_model(

        pipeline_v2,

        "model",

        serialization_format="cloudpickle"

    )

    print(f"V2 Logged - Run ID: {run_v2.info.run_id}")

Output excerpt:

V2 Logged - Run ID: ee892572d0a641f89201c33479b98746

Auditing, Comparing, and Registering Models

Now that we have multiple logged pipeline versions, we invoke the MLflow search API to extract the full versioning experiment and display it as a pandas DataFrame. Note that key auditing columns have been separated for clarity: run ID, MLflow run name, local LLM parameter, and execution status. For a realistic touch, the results below (based on previous runs leading to the final code included in this article) show historical audit information from several executions — displaying not only MLflow tracking of FINISHED pipelines but also early FAILED attempts.

experiment = mlflow.get_experiment_by_name("Scikit-LLM-Versioning")

runs_df = mlflow.search_runs(experiment.experiment_id)

comparison_df = runs_df[['run_id', 'tags.mlflow.runName', 'params.llm_model_file', 'status']]

print("Experiment Tracking Audit:")

display(comparison_df)

run_id tags.mlflow.runName                                params.llm_model_file    status

0  ee892572d0a641f89201c33479b98746     Upgraded_Falcon          gpt4all::ggml-model-gpt4all-falcon-q4_0.bin  FINISHED

1  0852aaec23364725b433f09973a3d911  Baseline_Orca_Mini                  gpt4all::orca-mini-3k-71m-q4_0.gguf  FINISHED

2  68781001dec14e0cbe46e71cf38e92c4     Upgraded_Falcon          gpt4all::ggml-model-gpt4all-falcon-q4_0.bin  FINISHED

3  37e6011d2cca4e43a0a426cb936582e6  Baseline_Orca_Mini                  gpt4all::orca-mini-3k-71m-q4_0.gguf  FINISHED

4  5ccd7e21a74a4fcc8a01633231417b28  Baseline_Orca_Mini                  gpt4all::orca-mini-3k-71m-q4_0.gguf    FAILED

Note that if you run the provided code and all cells execute without errors, you may see a shorter list — ideally containing only two logged runs associated with the two pipelines, both with FINISHED status.

To wrap up, let’s shift from logged to registered. In other words, let’s see how to extract the optimal execution run and promote (formally register) its associated model into MLflow’s Model Registry. The code searches the DataFrame to find the first run matching the "Upgraded_Falcon" label and secures its run ID. This target pipeline is then registered in the backend database, officially recorded as Version 1.

best_run_id = runs_df[runs_df['tags.mlflow.runName'] == 'Upgraded_Falcon'].iloc[0]['run_id']

model_uri = f"runs:/{best_run_id}/model"

registered_model = mlflow.register_model(

    model_uri=model_uri,

    name="Production_ZeroShot_Classifier"

)

print(f"Successfully registered model '{registered_model.name}'")

print(f"Current Registry Version: {registered_model.version}")

Output:

Successfully registered model 'Production_ZeroShot_Classifier'

Current Registry Version: 1

We just performed a hardcoded, manual model selection, but what if you want to find and register the one with the best performance — for instance, the highest accuracy? You could do something like this before calling mlflow.register_model():

# Retrieving runs ordered by accuracy (highest to lowest)

best_runs_df = mlflow.search_runs(

    experiment_ids=[experiment.experiment_id],

    order_by=["metrics.accuracy DESC"]

)

# Extracting the ID of the absolute top performer

metric_winner_id = best_runs_df.iloc[0]

print(metric_winner_id)

Output:

run_id                                      0605f300074a4d91b1e3438348d157f1

experiment_id                                                              1

status                                                              FINISHED

artifact_uri               /content/mlruns/1/0605f300074a4d91b1e3438348d1...

start_time                                  2026-08-29 14:53:18.182000+00:00

end_time                                    2026-08-29 14:53:22.007000+00:00

params.llm_model_file            gpt4all::ggml-model-gpt4all-falcon-q4_0.bin

params.llm_backend                                                   gpt4all

tags.mlflow.source.name             fileId=1GM67JQ61d3Y7eibN63YS6Udjt5qxf14x

tags.mlflow.runName                                          Upgraded_Falcon

tags.mlflow.user                                                        root

tags.mlflow.source.type                                             NOTEBOOK

Wrapping Up

The two-step (logging and registering) workflow for LLM pipeline versioning introduced in this article is designed to prevent your official model registry from becoming cluttered with failed attempts, messy code excerpts, or inferior test runs that led nowhere. The tracking table displaying logged versions is used to compare a set of rough drafts, publishing only the final “winner(s)” to the registry database for deployment or active use.

No comments yet.