> ## Documentation Index
> Fetch the complete documentation index at: https://docs.fastino.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Use https://docs.fastino.ai/llms.txt to discover and navigate pages. Use https://docs.fastino.ai/llms-full.txt when you need the complete documentation corpus. Use https://docs.fastino.ai/openapi.json as the source of truth for customer-facing routes. For GLiDE decision inference, call POST https://api.fastino.ai/v1/systemone with model fastino/GLiDE. Do not infer undocumented routes. Read API keys from FASTINO_API_KEY and never embed credentials in code, logs, or reports.

# 评估和训练 GLiNER 2.5 Decide

> 使用代表性数据评估 GLiNER 2.5 Decide、选择多标签阈值，并针对特定领域微调模型。

评估应用将发送给 `fastino/GLiNER-2.5-Decide` 的确切 `classifications` schema。仅在两次评估运行之间更改标签或任务设置，以确保结果可比。

## 评估 Decide schema

从一个固定任务开始：

```json theme={null}
{
  "task": "support_tags",
  "labels": ["billing", "technical", "account", "security"],
  "multi_label": true,
  "cls_threshold": 0.5
}
```

使用相同的标签名称构建留出集，其中包括：

* 每个标签的正例和反例
* 应返回多个标签的输入
* 没有任何标签应达到阈值的输入
* 模糊文本和域外文本

使用相同的 Decide 模型 ID 和 schema，通过 `POST /v1/chat/completions` 发送每个示例。在对任务结果评分之前，先解析 `choices[0].message.content` 中的 JSON 字符串。

对于单标签 Decide 任务，衡量 top-1 准确率以及每个标签的混淆情况。对于多标签任务，衡量每个标签及整个任务的精确率、召回率和 F1。如果一个请求包含多个任务，请分别评估每个任务，因为 Decide 不会强制规定它们之间的关系。

## 选择 `cls_threshold`

`cls_threshold` 仅影响 `multi_label: true` 的任务。Decide 会返回达到阈值的标签，并省略其余标签：

* 提高阈值会返回更少的标签，通常更偏向精确率。
* 降低阈值会返回更多的标签，通常更偏向召回率。
* 单标签任务始终返回得分最高的标签。
* 阈值只影响推理，不会改变训练。

针对每个候选阈值重新运行留出集，因为响应仅包含在该阈值下被选中的标签。当假阳性和假阴性的成本不同时，请为每个任务选择一个阈值。

不要将 `0.5` 视为通用默认值。对于会触发重大操作的标签，应要求更严格的阈值，或将低置信度结果交由人工复核。

## 训练 GLiNER 2.5 Decide

只有在明确标签分类体系并调整多标签阈值之后，基础模型仍然无法正确处理你的领域时，才进行训练。

当前的在线基础模型目录提供以下信息：

```json theme={null}
{
  "id": "fastino/GLiNER-2.5-Decide",
  "supports_training": true,
  "training_types": ["lora"]
}
```

创建作业前，请检查 `GET /v1/base-models?supports_training=true`，因为目录中的可用模型可能会发生变化。使用 `fastino/GLiNER-2.5-Decide` 作为 `base_model`，使用 `lora` 作为 `training_type`。

### 准备 Decide 分类数据

使用公开的分类 JSONL 格式。每一行都包含 `text`，以及一个 `label` 或多个 `labels`：

```jsonl theme={null}
{"text":"I was charged twice for the same invoice.","label":"billing"}
{"text":"I cannot sign in and suspect my account was compromised.","labels":["account","security"]}
```

数据集标签必须使用与 Decide 推理 schema 中相同的稳定字符串。不要在一个数据集中混用单标签和多标签行格式，并确保上传的训练数据不包含留出评估集。

<Card title="上传分类数据集" icon="database" href="/cn/concepts/datasets">
  创建、上传、处理用于 Decide 训练作业的 JSONL 数据集，并对其进行版本管理。
</Card>

### 创建 Decide 训练作业

将准备好的数据集提交给 Decide 基础模型：

```bash theme={null}
curl -X POST https://api.fastino.ai/v1/training-jobs \
  -H "X-API-Key: $FASTINO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model_name": "support-decide",
    "base_model": "fastino/GLiNER-2.5-Decide",
    "datasets": [{"name": "support-decide-training"}],
    "training_type": "lora",
    "validation_data_percentage": 0.2
  }'
```

除非有证据表明需要覆盖 Decide 训练方案，否则请省略学习率、批次大小和 LoRA 参数。保存返回的训练作业 ID，并轮询作业状态，直到产物可以部署。

### 比较并部署

使用相同的 `classifications` schema，在同一个留出集上运行基础模型和训练后的检查点。比较：

* 单标签任务中每个标签的错误
* 多标签任务中每个标签的精确率、召回率和 F1
* 空结果和标签选择过多的多标签结果
* 每个任务所需的阈值

只有当检查点能够改善你所关注的应用级错误时，才进行部署。部署后，请重新进行阈值选择；即使预测标签保持不变，微调也可能改变 Decide 的置信度值。

<CardGroup cols={2}>
  <Card title="训练作业" icon="activity" href="/cn/training">
    监控 Decide 作业、检查点和可部署状态。
  </Card>

  <Card title="训练 API 参考" icon="book" href="/cn/api-reference/training/overview">
    查看准确的数据集和训练作业契约。
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.