> ## Documentation Index
> Fetch the complete documentation index at: https://docs.fastino.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Use https://docs.fastino.ai/openapi.json as the source of truth for customer-facing Fastino API routes. Only call operations present in that specification. Do not infer or call undocumented routes. Direct Fastino API integrations use https://api.fastino.ai, /v1 routes, and FASTINO_API_KEY.

# 推理 API

> 通过 Fastino 的推理 API 运行基于 schema 的 GLiNER 预测。

Fastino 为 GLiNER 推理提供两种同步接口。请选择与你的模型和负载需求相匹配的接口：

| 功能   | `/v1/chat/completions`                                            | `/v1/gliner-2`                                                           |
| ---- | ----------------------------------------------------------------- | ------------------------------------------------------------------------ |
| 模型   | 选择支持推理的基础模型或已完成训练任务的 UUID                                         | 始终使用 `fastino/gliner2-base-v1`                                           |
| 请求   | 与 OpenAI 兼容的 `model` 和 `messages`，外加 Fastino 的顶层 `schema`         | 原生 `text`、`schema`、`threshold`、`include_confidence` 和 `include_spans` 字段 |
| 响应   | OpenAI chat-completion 信封；将 `choices[0].message.content` 解析为 JSON | 原生 `{ "result": ..., "token_usage": ... }` 响应体                           |
| 批量输入 | 每个请求一个对话                                                          | 接受单个字符串或字符串列表                                                            |
| 异步处理 | 不可用                                                               | 使用 `POST /v1/gliner-2/async` 提交，然后轮询 `GET /v1/gliner-2/jobs/{job_id}`    |
| 适用场景 | 标准集成和微调模型                                                         | 基础 GLiNER2 推理、原生结果、批处理或长时间运行的任务                                          |

<Tip>
  从 `/v1/chat/completions` 开始，可在基础模型和微调模型之间获得一致的 API。
  当你明确需要固定 GLiNER2 基础模型的原生契约、批量输入或异步处理时，
  请使用 `/v1/gliner-2`。
</Tip>

本页下文记录 `/v1/chat/completions` 的原始 HTTP 契约。原生 `/v1/gliner-2` 的 schema
发布在 [OpenAPI 文档](https://docs.fastino.ai/openapi.json)中。SDK 使用、模型训练、
评估、推理历史和反馈不在本页范围内。

<Warning>
  `POST /inference` 已被移除。请勿再发送其旧版字段，例如 `model_id`、`text`、
  `task`、`format_results` 或 `is_warmup`。
</Warning>

## 端点

```text theme={null}
POST https://api.fastino.ai/v1/chat/completions
```

## 身份验证

将 Fastino API 密钥作为 bearer token 发送：

```bash theme={null}
export FASTINO_API_KEY="fast_sk_..."
```

```http theme={null}
Authorization: Bearer $FASTINO_API_KEY
Content-Type: application/json
```

## 请求

<ParamField body="model" type="string" required>
  支持推理的基础模型 ID，或已完成、可部署的 Fastino 训练作业的 UUID。
</ParamField>

<ParamField body="messages" type="object[]" required>
  非空消息列表。对于 GLiNER 推理，请将待分析的文本放入用户消息的 `content` 中。
</ParamField>

<ParamField body="schema" type="object">
  定义自定义编码器任务。提供一个包含 `entities`、`classifications`、`structures` 或 `relations`
  中一个或多个键的字典。仅当所选模型已配置默认任务时才可省略。

  已弃用扁平的实体标签数组，请始终使用字典形式。
</ParamField>

<ParamField body="threshold" type="number" default="0.5">
  置信度阈值，取值范围为 `0` 到 `1`。较低的值更偏向召回率；较高的值更偏向精确率。
</ParamField>

<ParamField body="include_confidence" type="boolean" default="true">
  在提取结果中包含置信度值。
</ParamField>

<ParamField body="include_spans" type="boolean" default="true">
  在实体结果中包含半开区间的字符偏移（`start`、`end`）。
</ParamField>

<ParamField body="store" type="boolean" default="true">
  持久化本次推理。设置为 `false` 可选择不保存。
</ParamField>

### 实体 schema

尽量使用带描述的实体定义：

```json theme={null}
{
  "entities": [
    {
      "name": "organization",
      "description": "business or institution name"
    },
    {
      "name": "product",
      "description": "named commercial product"
    }
  ]
}
```

### 分类 schema

```json theme={null}
{
  "classifications": [
    {
      "task": "sentiment",
      "labels": ["positive", "negative", "neutral"],
      "multi_label": false,
      "top_k": 1
    }
  ]
}
```

### 结构化提取 schema

结构字段使用 `field::type::description` 规范：

```json theme={null}
{
  "structures": {
    "product": [
      "name::str::product name",
      "price::str::listed price",
      "features::list::named features"
    ]
  }
}
```

### 关系 schema

最简单的关系 schema 是关系名称的扁平列表：

```json theme={null}
{
  "relations": ["works_for", "lives_in"]
}
```

也可以使用字典来添加描述或按关系配置：

```json theme={null}
{
  "relations": {
    "works_for": {
      "description": "employment relationship",
      "threshold": 0.6
    }
  }
}
```

不要发送包含 head 和 tail 定义的关系对象。

### 组合 schema

通过组合键在同一段文本上运行多个任务：

```json theme={null}
{
  "entities": [
    {
      "name": "organization",
      "description": "business or institution name"
    },
    {
      "name": "product",
      "description": "named commercial product"
    }
  ],
  "classifications": [
    {
      "task": "sentiment",
      "labels": ["positive", "negative", "neutral"],
      "multi_label": false,
      "top_k": 1
    }
  ]
}
```

schema 会自动识别操作类型。请勿发送 `task` 或 `task_type`。

## 示例

```bash theme={null}
curl -X POST "https://api.fastino.ai/v1/chat/completions" \
  -H "Authorization: Bearer $FASTINO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "fastino/gliner2.5-multi-v1",
    "messages": [
      {
        "role": "user",
        "content": "Apple announced the MacBook Pro at WWDC in Cupertino."
      }
    ],
    "schema": {
      "entities": [
        {
          "name": "organization",
          "description": "business or institution name"
        },
        {
          "name": "product",
          "description": "named commercial product"
        },
        {
          "name": "event",
          "description": "named conference or event"
        },
        {
          "name": "location",
          "description": "city, region, or place"
        }
      ]
    },
    "threshold": 0.5
  }'
```

`model` 值必须当前支持托管推理。请使用实时模型目录，而不要假设每个 Hugging Face 检查点都可用：

```bash theme={null}
curl "https://api.fastino.ai/v1/base-models?supports_inference=true&task_type=encoder"
```

## 响应

该端点返回聊天补全信封。GLiNER 结果以 JSON 字符串形式序列化在 `choices[0].message.content` 中：

```json theme={null}
{
  "id": "chatcmpl_...",
  "object": "chat.completion",
  "created": 1790123456,
  "model": "fastino/gliner2.5-multi-v1",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "{\"entities\":[{\"organization\":[{\"text\":\"Apple\",\"confidence\":0.99,\"start\":0,\"end\":5}],\"product\":[{\"text\":\"MacBook Pro\",\"confidence\":0.98,\"start\":20,\"end\":31}],\"event\":[{\"text\":\"WWDC\",\"confidence\":0.97,\"start\":35,\"end\":39}],\"location\":[{\"text\":\"Cupertino\",\"confidence\":0.99,\"start\":43,\"end\":52}]}],\"classifications\":{},\"structures\":{},\"relations\":{}}"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 0,
    "completion_tokens": 0,
    "total_tokens": 0
  },
  "x_pioneer": {
    "inference_id": "..."
  }
}
```

在读取任务结果之前，请将 `choices[0].message.content` 解析为 JSON。当 `store=true` 时，
`x_pioneer.inference_id` 标识已持久化的推理。

## 重复输入

该端点每个请求接受一段会话。它不支持已移除的原生端点的 `text: string[]` 批量形式。
处理多个输入时，请并发发送多个独立请求。

## 冷启动与重试

空闲或新部署的模型可能会冷启动。请将读取超时设置为至少 300 秒，并对 `425`、`429` 和 `503`
响应进行重试。当响应头中包含 `Retry-After` 时，请遵循该值。

已超时的请求仍可能预热部署，从而让下一次请求成功。对于身份验证、计费、校验、未知模型或
不可部署作业等错误，请在解决根本问题之前不要重试。

```bash theme={null}
curl --max-time 300 -X POST "https://api.fastino.ai/v1/chat/completions" \
  -H "Authorization: Bearer $FASTINO_API_KEY" \
  -H "Content-Type: application/json" \
  -d @request.json
```

## 错误

| 状态码   | 可能原因                        | 处理措施                       |
| ----- | --------------------------- | -------------------------- |
| `400` | 请求或 schema 无效               | 修正请求字段或 schema             |
| `401` | 缺少、格式错误或无效的 API 密钥          | 导出有效的 `fast_sk_...` 密钥     |
| `402` | 额度不足、需要充值，或所有者设置了每日/每月的支出上限 | 添加额度或调整支出上限                |
| `403` | 支付方式、卡片验证、信用额度或账户策略拒绝       | 解决相应的计费或账户要求               |
| `404` | 未知或不受支持的模型                  | 使用支持推理的目录 ID 或可部署作业的 UUID  |
| `409` | 微调模型存在但未部署                  | 等待部署完成或修复模型部署              |
| `422` | 请求校验失败                      | 阅读并修正字段级别的校验详情             |
| `425` | 微调部署正在预热                    | 按提示的延迟时间后重试                |
| `429` | 触发速率或容量限制                   | 遵循 `Retry-After` 并使用退避机制重试 |
| `503` | 冷启动或提供方暂时性容量问题              | 视为终止之前先重试                  |

## 旧版请求迁移

| 已移除的 `/inference` 字段           | 当前字段                                             |
| ------------------------------ | ------------------------------------------------ |
| `model_id`                     | `model`                                          |
| `text`                         | `messages: [{"role": "user", "content": "..."}]` |
| 扁平的 `schema` 实体列表              | `schema` 字典                                      |
| `task` 或 `task_type`           | 删除；schema 会识别操作类型                                |
| `text: string[]`               | 发送多个独立请求                                         |
| `format_results` 或 `is_warmup` | 删除                                               |
| 原生结果体                          | 解析 `choices[0].message.content`                  |
