/training

Tune it where you serve it.

Fine-tune S Series or open weights on your data. The artifact lands in your catalog with an id of its own, callable and routable like any other model. Training is by invitation. It is not self-serve GA. Join the waitlist for managed and API jobs, or get a quote for custom work.

Three ways to train

/surfaces

LoRA · no custom loop

Managed training

Give Grid your data and the method. We schedule the job, train, checkpoint, and return a catalog id. SFT, DPO, ORPO, and reinforcement fine-tuning run here. Console or API. OpenAI-compatible chat JSONL, so an existing SFT dataset does not need a conversion.

Custom loop · LoRA · per-token

Serverless training API

Write the loop in Python. Grid attaches to shared trainers. LoRA only. You pay for training tokens, not idle GPUs. Use this when a standard managed job is not enough and you do not need a dedicated cluster.

Custom loop · LoRA or full-parameter

Dedicated training API

Your trainers, held for the job. Custom loss, distillation, DPO, ORPO, and full-parameter updates run here. Time-based capacity, quoted. Tear it down when the run ends.

Methods

/methods

The same method set as a full-spectrum training stack: supervised, preference, reinforcement, distillation, and full-parameter. Eligibility is per base model. Qwen 3.8-27B is the fine-tune base in the public catalog today.

Managed or Training API

SFT

Supervised fine-tuning

Train on labeled examples of the output you want. Text and vision. The default first job. LoRA on managed. Full-parameter on the dedicated API.

Managed or Dedicated API

DPO

Direct preference optimization

Train on preferred and non-preferred pairs for the same prompt. DPO compares the policy to a reference model. Use it for tone, format, and answers that have no single gold string.

Managed or Dedicated API

ORPO

Odds ratio preference optimization

The same preference pairs as DPO. ORPO folds a supervised objective in, so you do not hold a separate reference model. Same dataset shape: input, preferred_output, non_preferred_output.

Managed

RFT

Reinforcement fine-tuning

Train against a reward function, not only paired text. Grid runs the loop: rollouts, scoring, checkpoint, promote. Bring an evaluator. For teams that already have a trainer, use RL rollouts instead.

Dedicated API

Distillation

Teacher to student

A custom loop that transfers behavior from a larger teacher into a smaller student. Not a managed preset. Quoted with the dedicated API.

Dedicated API

Full-parameter

Update the base weights

LoRA is the fast start. Full-parameter is the ceiling when adapters cannot move the behavior. Managed jobs stay LoRA. Full-parameter is dedicated only.

Custom training

Our team will train it for you.

If you do not want to run the job, Forward Intelligence will. You bring the task, the data, and the success bar. We pick the method, run SFT, DPO, ORPO, or RFT, and hand back a catalog id on Grid. Quoted. Not a packaged SKU and not a public rate.

If the work is a whole workload on the factory, not one training job, that is a custom project. Talk to us there.

Keep your trainer

RL rollouts

Point your own RL loop at Grid for sampling. We run the rollout inference. You keep the trainer. OpenAI-compatible. Reserved capacity, quoted. Not the same product as managed RFT, which runs the whole loop for you.

RL rollouts

What you train is what you serve

Same factory

Checkpoints promote onto Grid Serverless, Dedicated, or Private. Smart Router can include the new id if you put it in the mix. Training is metered as jobs, not as inference tokens. Batch is not the way to train.

Training docs

Your first call is one base_url away.

Request access

Grid is invite only while we scale.

Talk to us

Tell us who you are and what this is about.