惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
V
V2EX
博客园_首页
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Recent Announcements
Recent Announcements
博客园 - 司徒正美
Microsoft Security Blog
Microsoft Security Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Latest news
Latest news
Vercel News
Vercel News
The Register - Security
The Register - Security
T
The Exploit Database - CXSecurity.com
S
Schneier on Security
N
Netflix TechBlog - Medium
WordPress大学
WordPress大学
小众软件
小众软件
L
Lohrmann on Cybersecurity
GbyAI
GbyAI
P
Privacy & Cybersecurity Law Blog
T
Tor Project blog
AWS News Blog
AWS News Blog
美团技术团队
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
K
Kaspersky official blog
B
Blog RSS Feed
G
Google Developers Blog
量子位
大猫的无限游戏
大猫的无限游戏
Google DeepMind News
Google DeepMind News
Scott Helme
Scott Helme
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
I
Intezer
雷峰网
雷峰网
Martin Fowler
Martin Fowler
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Blog — PlanetScale
Blog — PlanetScale
IT之家
IT之家
F
Full Disclosure
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - 【当耐特】
The Hacker News
The Hacker News
U
Unit 42
S
SegmentFault 最新的问题
I
InfoQ
aimingoo的专栏
aimingoo的专栏
Y
Y Combinator Blog
宝玉的分享
宝玉的分享
罗磊的独立博客
Spread Privacy
Spread Privacy
C
CERT Recently Published Vulnerability Notes

博客园 - badwood

ssl自签证书+nginx 备忘 Claude Code私有化使用 mindie推理框架及工程化 minio-2.使用 minio-1.搭建 旧源消失的rpm包安装-devtoolset-9 vllm-ascend 2/2 -双机推理 MCP-1.hello world vllm-ascend 1/2 -单机推理 ESP32-S3玩具1-使用Arduino evalscope使用2-使用自定义数据集压测 设置fdfs自动启动 FunASR mindie开启DeepSeek的128K xinference推理embedding等小模型 离线安装docker 昇腾910b服务器初始化 evalscope使用1-基础压测 大模型围栏-nginx+lua 大模型私有化部署-deepseek671-mindie
dify-2:问题分类器调整
badwood · 2025-05-29 · via 博客园 - badwood

  运行问题分类器时无意间发现数据处理那块有一堆输出,见下。是给大模型的提示词,突发想法用deepseek分析了下,它指出了几个问题:格式污染;冗余输入;中文案例较少;类别设计待优化等。顺便研究了下api源码,发现该起来还挺容易。

{
  "model_mode": "chat",
  "prompts": [
    {
      "role": "system",
      "text": "\n    ### Job Description',\n    You are a text classification engine that analyzes text data and assigns categories based on user input or automatically determined categories.\n    ### Task\n    Your task is to assign one categories ONLY to the input text and only one category may be assigned returned in the output. Additionally, you need to extract the key words from the text that are related to the classification.\n    ### Format\n    The input text is in the variable input_text. Categories are specified as a category list with two filed category_id and category_name in the variable categories. Classification instructions may be included to improve the classification accuracy.\n    ### Constraint\n    DO NOT include anything other than the JSON array in your response.\n    ### Memory\n    Here are the chat histories between human and assistant, inside <histories></histories> XML tags.\n    <histories>\n    \n    </histories>\n",
      "files": []
    },
    {
      "role": "user",
      "text": "\n    { \"input_text\": [\"I recently had a great experience with your company. The service was prompt and the staff was very friendly.\"],\n    \"categories\": [{\"category_id\":\"f5660049-284f-41a7-b301-fd24176a711c\",\"category_name\":\"Customer Service\"},{\"category_id\":\"8d007d06-f2c9-4be5-8ff6-cd4381c13c60\",\"category_name\":\"Satisfaction\"},{\"category_id\":\"5fbbbb18-9843-466d-9b8e-b9bfbb9482c8\",\"category_name\":\"Sales\"},{\"category_id\":\"23623c75-7184-4a2e-8226-466c2e4631e4\",\"category_name\":\"Product\"}],\n    \"classification_instructions\": [\"classify the text based on the feedback provided by customer\"]}\n",
      "files": []
    },
    {
      "role": "assistant",
      "text": "\n```json\n    {\"keywords\": [\"recently\", \"great experience\", \"company\", \"service\", \"prompt\", \"staff\", \"friendly\"],\n    \"category_id\": \"f5660049-284f-41a7-b301-fd24176a711c\",\n    \"category_name\": \"Customer Service\"}\n```\n",
      "files": []
    },
    {
      "role": "user",
      "text": "\n    {\"input_text\": [\"bad service, slow to bring the food\"],\n    \"categories\": [{\"category_id\":\"80fb86a0-4454-4bf5-924c-f253fdd83c02\",\"category_name\":\"Food Quality\"},{\"category_id\":\"f6ff5bc3-aca0-4e4a-8627-e760d0aca78f\",\"category_name\":\"Experience\"},{\"category_id\":\"cc771f63-74e7-4c61-882e-3eda9d8ba5d7\",\"category_name\":\"Price\"}],\n    \"classification_instructions\": []}\n",
      "files": []
    },
    {
      "role": "assistant",
      "text": "\n```json\n    {\"keywords\": [\"bad service\", \"slow\", \"food\", \"tip\", \"terrible\", \"waitresses\"],\n    \"category_id\": \"f6ff5bc3-aca0-4e4a-8627-e760d0aca78f\",\n    \"category_name\": \"Experience\"}\n```\n",
      "files": []
    },
    {
      "role": "user",
      "text": "\n    '{\"input_text\": [\"你是谁\"],',\n    '\"categories\": [{\"category_id\": \"1\", \"category_name\": \"需要运营分析\"}, {\"category_id\": \"2\", \"category_name\": \"执行数据同步\"}, {\"category_id\": \"1748438042043\", \"category_name\": \"用户意图不明\"}], ',\n    '\"classification_instructions\": [\"\"]}'\n",
      "files": []
    },
    {
      "role": "user",
      "text": "你是谁",
      "files": []
    }
  ],
  "usage": {
    "prompt_tokens": 0,
    "prompt_unit_price": "0",
    "prompt_price_unit": "0",
    "prompt_price": "0",
    "completion_tokens": 0,
    "completion_unit_price": "0",
    "completion_price_unit": "0",
    "completion_price": "0",
    "total_tokens": 0,
    "total_price": "0",
    "currency": "USD",
    "latency": 2.3074390459805727
  },
  "finish_reason": "stop"
}

  1、通过关键字定位源码:提示词模板在template_prompts.py;运行代码在question_classifier_node.py

[app@localhost api]$ pwd
/app/dify/api
[app@localhost api]$ grep -rl "You are a text classification engine"
core/workflow/nodes/question_classifier/template_prompts.py
[app@localhost api]$ ls core/workflow/nodes/question_classifier/
entities.py  exc.py  __init__.py  question_classifier_node.py  template_prompts.py

  2、格式污染:单引号包裹'{\"input_text\":...}'可能破坏JSON解析。修改提示词模板:红色粗体部分。也就去掉了单引号。

QUESTION_CLASSIFIER_SYSTEM_PROMPT = """
    ### Job Description',
    You are a text classification engine that analyzes text data and assigns categories based on user input or automatically determined categories.
    ### Task
    Your task is to assign one categories ONLY to the input text and only one category may be assigned returned in the output. Additionally, you need to extract the key words from the text that are related to the classification.
    ### Format
    The input text is in the variable input_text. Categories are specified as a category list with two filed category_id and category_name in the variable categories. Classification instructions may be included to improve the classification accuracy.
    ### Constraint
    DO NOT include anything other than the JSON array in your response.
    ### Memory
    Here are the chat histories between human and assistant, inside <histories></histories> XML tags.
    <histories>
    {histories}
    </histories>
"""  # noqa: E501
...
QUESTION_CLASSIFIER_USER_PROMPT_3 = """
    {{"input_text": ["{input_text}"],
    "categories": {categories},
    "classification_instructions": ["{classification_instructions}"]}}
"""
...
### Memory
Here are the chat histories between human and assistant, inside <histories></histories> XML tags.
<histories>
{histories}
</histories>
### User Input
{{"input_text" : ["{input_text}"], "categories" : {categories},"classification_instruction" : ["{classification_instructions}"]}}
### Assistant Output
"""  # noqa: E501

  3、冗余输入:两个连续的user消息(结构化和纯文本)可能引起歧义。这个需要修改question_classifier_node.py中的代码,见红色粗体部分,原始为sys_query=query

import json
from collections.abc import Mapping, Sequence
from typing import Any, Optional, cast

...

prompt_messages, stop = self._fetch_prompt_messages(
prompt_template=prompt_template,
sys_query="",
memory=memory,
model_config=model_config,
sys_files=files,
vision_enabled=node_data.vision.enabled,
vision_detail=node_data.vision.configs.detail,
variable_pool=variable_pool,
jinja2_variables=[],
)

...
else:
            raise InvalidModelTypeError(f"Model mode {model_mode} not support.")

  4、中文案例较少,这个部分没有做改造,后续考虑是否要把英文提示词全部改成汉语。

  目前的修改就这些,是否对于最终效果能有帮助还有待观察。