Skip to main content
This endpoint uses an OpenAI-compatible envelope: an inference-capable model, a messages list, and a top-level schema describing the GLiNER task. See the Inference API overview for authentication and cold-start/retry behavior shared across Fastino’s inference interfaces.
POST /inference has been removed. Do not send its legacy fields such as model_id, text, task, format_results, or is_warmup.

Request

string
required
An inference-capable base-model ID or the UUID of a completed, deployable Fastino training job.
object[]
required
A non-empty list of messages. For GLiNER inference, put the text to analyze in a user message’s content.
object
Defines custom encoder tasks. Supply a dictionary containing one or more of entities, classifications, structures, or relations. Omit it only when the selected model has a configured default task.A flat array of entity labels is deprecated. Always use the dictionary shape.
number
default:"0.5"
Confidence threshold from 0 to 1. Lower values favor recall; higher values favor precision.
boolean
default:"true"
Include confidence values in extracted results.
boolean
default:"true"
Include half-open character offsets (start, end) in entity results.
boolean
default:"true"
Persist the inference. Set to false to opt out.

Entity schema

Use descriptive entity definitions when possible:

Classification schema

Structured extraction schema

Structure fields use field::type::description specifications:

Relation schema

The simplest relation schema is a flat list of relation names:
You can also use a dictionary to add descriptions or per-relation configuration:
Do not send relation objects containing head and tail definitions.

Combined schema

Run multiple tasks over the same text by combining keys:
The schema identifies the operation automatically. Do not send task or task_type.

Example

The model value must currently support hosted inference. Use the live model catalog rather than assuming every Hugging Face checkpoint is available:

Response

The endpoint returns a chat-completion envelope. The GLiNER result is serialized as a JSON string in choices[0].message.content:
Parse choices[0].message.content as JSON before reading task results. When store=true, x_pioneer.inference_id identifies the persisted inference.

Repeated inputs

The endpoint accepts one conversation per request. It does not support the removed native endpoint’s text: string[] batch shape. Send separate requests concurrently when processing multiple inputs.

Errors

See Cold starts and retries for the shared retry pattern.

Legacy request migration