Skip to main content
Cognesy\Polyglot\BatchInference\BatchInference is an experimental, opt-in facade for provider-native batch jobs. Submission returns an observed job and a durable reference. The provider processes requests later. Your application decides when to check status, read outcomes, or request cancellation. None of these operations waits for remote inference to finish.
BatchJob accessors are pure: calling status(), progress(), or resultsAvailability() does not poll. A completed job may contain failed items. An in-progress xAI job may already have partial results. Use capabilities() to inspect cancellation, listing, typed result reading, partial-result support, input modes, and enforced input limits before submission. Some providers also publish a minimum item count. The native drivers currently cover OpenAI Chat Completions and Responses, Anthropic, Mistral, Gemini, Qwen, xAI/Grok, Groq, and Together. Each driver has separate wire encoding and status mapping. Public documentation alone does not qualify every model or account for batch processing; review provider details before submitting paid work. Next: submit and resume, status and cancellation, and results.