tinygrad

mirror of https://github.com/tinygrad/tinygrad.git synced 2026-06-13 08:28:55 +08:00

Files

chenyu 45baec1aab model parallel llama (#11588 )

MP=8 GRADIENT_ACC_STEPS=3 BS=1 DEFAULT_FLOAT=bfloat16 OPTIM_DTYPE=bfloat16 LLAMA3_SIZE=70B SEQLEN=512 PYTHONPATH=. MODEL=llama3 python3 examples/mlperf/model_train.py

2025-08-09 16:54:27 -04:00

scripts

UNet3D MLPerf (#3470 )

2024-09-10 04:37:28 -04:00

training_submission_v4.0/tinycorp

copy mlperf 4.0 to mlperf 4.1 (#5614 )

2024-07-20 16:12:00 -04:00

training_submission_v4.1/tinycorp

update mlperf systems and copy 4.1 to 5.0 (#7004 )

2024-10-11 16:20:34 -04:00

training_submission_v5.0/tinycorp

mlperf system updates (#10550 )

2025-05-28 16:15:46 -04:00

training_submission_v5.1/tinycorp

remove FUSE_ARANGE_UINT (#11567 )

2025-08-07 16:49:06 -04:00

dataloader.py

feat: generate blend index (#11566 )

2025-08-07 14:20:28 -04:00

helpers.py

ruff check whole examples/mlperf/ (#10979 )

2025-06-25 12:57:48 -04:00

initializers.py

ruff check whole examples/mlperf/ (#10979 )

2025-06-25 12:57:48 -04:00

losses.py

cleanups on losses and dataset tests (#9538 )

2025-03-21 17:03:18 -04:00

lr_schedulers.py

CosineAnnealingLRWithWarmup (#10981 )

2025-06-25 17:45:21 -04:00

metrics.py

log_perplexity metrics (#10912 )

2025-06-21 10:44:47 -04:00

model_eval.py

feat: llama3 dataloader (#11340 )

2025-07-30 13:27:55 -07:00

model_spec.py

remove Tensor.no_grad, it's meaningless now [pr] (#10556 )

2025-05-28 22:20:02 -07:00

model_train.py

model parallel llama (#11588 )

2025-08-09 16:54:27 -04:00

README

start on mlperf models

2023-05-10 16:30:49 -07:00

README

Each model should be a clean single file.
They are imported from the top level `models` directory

It should be capable of loading weights from the reference imp.

We will focus on these 5 models:

# Resnet50-v1.5 (classic) -- 8.2 GOPS/input
# Retinanet
# 3D UNET (upconvs)
# RNNT
# BERT-large (transformer)

They are used in both the training and inference benchmark:
https://mlcommons.org/en/training-normal-21/
https://mlcommons.org/en/inference-edge-30/
And we will submit to both.

NOTE: we are Edge since we don't have ECC RAM