Yooo
@ProCreations @Banaxi-Tech could you guys have a look at the code if there’s any errors? I’ve got a 8xH200 run on the 5th from the super micro jumpstart thing and it would be good if no errors occur lol. I’m waiting for my usage to reset
@Bc-AI here are the issues:
The startup gradient check will reject a healthy model. In gradient_health_summary() names are checked against prefixes such as "vault" and "embedding". FSDP adds _fsdp_wrapped_module., so the gradients are incorrectly classified as cortex gradients. all four other groups become zero, causing the startup probe to raise. Also, gradient norms need to be combined across ranks because individual ranks hold only portions of the parameters.
Telemetry and ablation probes bypass FSDP. model.run_telemetry_probe() and model.run_ablation_probe() resolve to methods on the underlying model. Their internal self(...) calls bypass the FSDP wrapper, so parameters aren’t gathered before computation. This can produce tensor-shape errors. The probes should perform forward passes through the wrapped model(...).
Checkpoint resume is broken in two ways. In restore_checkpoint() Model weights are saved as .safetensors but loaded with torch.load(). Use the already-imported load_file() instead. Only model_shards[0] and optimizer_shards[0] are loaded. The saving function splits larger states into multiple files, so restoration must load and merge every listed shard.
Periodic checkpointing can deadlock the eight GPUs. Each rank independently evaluates the checkpoint condition, although saving requires every rank to participate together. Worse, last_save_time and dirty are updated only on rank zero. Other ranks can therefore enter another save while rank zero continues training. Rank zero should decide when to save, broadcast that decision, and update checkpoint bookkeeping on every rank.
Data-pipeline failure recovery can also deadlock. If rank zero’s stream.batch() raises inside distributed_batch(), the other ranks are already waiting for its broadcast. Rank zero then attempts a collective recovery save that they never enter. Broadcast a success/failure status before broadcasting the batch.
The suggested OOM fix does nothing. The error messages say to set MICRO_BATCH_SIZE=56, but main() overwrites that setting with:
MICRO_BATCH_SIZE = GLOBAL_MICRO_BATCH_SIZE // WORLD_SIZE
With the current settings, that always becomes 14 per GPU. To halve the physical batch while preserving tokens per update, change GLOBAL_MICRO_BATCH_SIZE from 112 to 56; accumulation then becomes 2.
I could be wrong on some of those so take my advice with a grain of salt (or 10)
@GGUFGuy @Hoglet-33 completely free. i am a year 7 student and i got no job so i cant do paid lol. you have to use a like custom email like yourname@yourwebsite.com for sign up i think https://www.supermicro.com/en/jumpstart
you can also acces B300s and blah at least the person said so the email said if i wanted to trial other stuff i email them.
@Hoglet-33 yeah you can do multiple times I think I applied for B300x8 after this. So u go to the site find one u like click apply button fill the form our put in our custom email domain and uh my one is for 5 day idk for others
Shhh! 🤫
@Banaxi-Tech AI SLOP BATTLE
Lol
