/training/rl-rollouts

Keep your trainer. Scale the rollouts.

Dedicated rollout inference for teams running their own RL. Point the fleet at Grid. We serve the samples. You own the algorithm. Reserved capacity. Talk to us to size it.

All training

What you get

/rollouts

Keep the trainer

Your loop stays where it is. Grid is the sampling fleet: OpenAI-compatible completions, sticky sessions for multi-turn trajectories, and a weight-update path we size with you.

Same inference stack

Rollouts run on the factory that already serves Grid. You do not stand up a second engine beside production. Capacity can move between product inference and rollout windows.

Not managed RFT

Managed reinforcement fine-tuning runs the whole loop for you: data, reward, train, promote. RL rollouts are the layer underneath, for teams that already have a trainer and do not want to migrate it.

Quoted capacity

Rollouts run on reserved capacity, not on anonymous Serverless. We start with a proof of concept, then size regions. There is no public rate card.

Questions

/faq

Can I keep my own trainer?

Yes. That is the product. Upload checkpoints to object storage we agree on, sample through Grid, and push weight updates on a schedule we set with you.

How is this different from managed RFT?

Managed RFT is the full job: you bring data and a reward, Grid trains and deploys. RL rollouts are sampling and fleet orchestration only. Pick rollouts if the trainer is already yours.

Is this live self-serve?

No. It is reserved capacity, by conversation. Talk to us to size a proof of concept.

Your first call is one base_url away.

Request access

Grid is invite only while we scale.

Talk to us

Tell us who you are and what this is about.