Reproduce "SMELT: Scaling Laws for MoE Looped Transformers" at small scale

Reading the paper, then pricing the smallest run that can test its claim.

hf_jobs run · l40sx1 · 45m · $1.35 reserved

Four compute-matched runs at 120M params, 2 looped / 2 flat.

trackio dashboard is live — loss is tracking the paper's curve.

looped-120mflat-120m
train/loss
3.42.92.31.801k2k3k4k

At 1.2k steps the looped variant is 0.06 ahead, same as Fig. 3

HuggingChat
ML Intern papers · training · spaces · datasets · eval · hub
1 2 3 4
Training
Describe the ML task — paper URL, model id, or dataset
Reproduce "SMELT: Scaling Laws for MoE Looped Transformers" at small scale
ML Intern NEW Tools
Finetune @google/gemma-4-E2B-it on the Qwen3.8 distillation set

LoRA on 40k reasoning traces, bf16, packing on — 2h ceiling.

hf_jobs uv run · a10g-large · 2h · $3.00 reserved

Pushing checkpoints to the Hub every 500 steps.

Eval split held back so the score at the end means something.

lora-r16
train/loss
2.31.91.61.306001.2k1.8k2.5k
eval/loss
2.31.91.61.306001.2k1.8k2.5k

Step 900: eval loss 1.42, down from 1.96 at the base model

HuggingChat
ML Intern papers · training · spaces · datasets · eval · hub
1 2 3 4
Training
Describe the ML task — paper URL, model id, or dataset
Finetune @google/gemma-4-E2B-it on the Qwen3.8 distillation set
ML Intern NEW Tools
Generate a synthetic function-calling dataset with @zai-org/GLM-5.3-Flash

Seeding 40 tool schemas, then sampling multi-turn calls against them.

hf_jobs run · cpu-upgrade · 30m · $0.02 reserved

Dropping malformed calls and near-duplicates as they come in.

18,240 rows kept of 21,000 sampled — 87% pass rate.

18,240 rows · 4 cols
toolarguments
search_flights{"from":"SFO","to":"CDG","date":"2026-10-02"}
get_forecast{"city":"Lisbon","units":"metric","days":5}
create_issue{"repo":"hf/chat-ui","title":"Budget pill …"}
run_sql{"query":"select count(*) from events where …"}

Pushing to the Hub with a datasheet and the generation config

HuggingChat
ML Intern papers · training · spaces · datasets · eval · hub
1 2 3 4
Generating
Describe the ML task — paper URL, model id, or dataset
Generate a synthetic function-calling dataset with @zai-org/GLM-5.3-Flash
ML Intern NEW Tools
Build a Gradio demo Space for @MiniMaxAI/MiniMax-H3

Image-to-video, so the UI is one image drop plus a prompt box.

hf_spaces create · zerogpu · gradio 6.2

Warming the weights on build so the first click isn't a cold start.

Smoke-tested one generation end to end before handing it over.

minimax-h3-demo
Generate video

Live at hf.co/spaces/ml-intern-explorers/minimax-h3-demo

HuggingChat
ML Intern papers · training · spaces · datasets · eval · hub
1 2 3 4
Deploying
Describe the ML task — paper URL, model id, or dataset
Build a Gradio demo Space for @MiniMaxAI/MiniMax-H3
ML Intern NEW Tools
Reproducepapers
Finetunemodels
Synthesizedatasets
BuildanddeployMLapps
ML Intern
in Hugging Chat
hf.co/chat
0.00s space play · ←→ step · H hide