Cognesy\Polyglot\BatchInference\BatchInference is an experimental, opt-in
facade for provider-native batch jobs. Submission returns an observed job and a
durable reference. The provider processes requests later. Your application
decides when to check status, read outcomes, or request cancellation. None of
these operations waits for remote inference to finish.
BatchJob accessors are pure: calling status(), progress(), or
resultsAvailability() does not poll. A completed job may contain failed
items. An in-progress xAI job may already have partial results. Use
capabilities() to inspect cancellation, listing, typed result reading, partial-result support,
input modes, and enforced input limits before submission. Some providers also
publish a minimum item count.
The native drivers currently cover OpenAI Chat Completions and Responses,
Anthropic, Mistral, Gemini, Qwen, xAI/Grok, Groq, and Together. Each driver
has separate wire encoding and status mapping. Public documentation alone
does not qualify every model or account for batch processing; review
provider details before submitting paid work.
Next: submit and resume,
status and cancellation, and
results.