<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->
synpre_mix_v2_1M_t5-base
This model is a fine-tuned version of t5-base on the tyzhu/synpre_mix_v2_1M dataset. It achieves the following results on the evaluation set:
- Loss: 0.0230
- Bleu: 97.6147
- Gen Len: 88.3084
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 0.0001
- train_batch_size: 128
- eval_batch_size: 128
- seed: 42
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lr_scheduler_type: constant_with_warmup
- lr_scheduler_warmup_steps: 10000
- training_steps: 80000
Training results
Training Loss | Epoch | Step | Validation Loss | Bleu | Gen Len |
---|---|---|---|---|---|
9.5227 | 0.64 | 5000 | 9.3562 | 0.1747 | 104.0258 |
9.1242 | 1.28 | 10000 | 8.8917 | 0.7778 | 77.3082 |
8.7226 | 1.92 | 15000 | 8.2870 | 0.798 | 53.5217 |
6.4581 | 2.56 | 20000 | 5.3528 | 3.5251 | 55.7876 |
0.1451 | 3.2 | 25000 | 0.0601 | 94.2354 | 90.0891 |
0.0491 | 3.84 | 30000 | 0.0299 | 97.9265 | 87.9334 |
0.0366 | 4.48 | 35000 | 0.0280 | 97.6519 | 88.2664 |
0.0328 | 5.12 | 40000 | 0.0258 | 97.4026 | 88.5963 |
0.0289 | 5.76 | 45000 | 0.0252 | 97.7039 | 88.0322 |
0.0269 | 6.4 | 50000 | 0.0233 | 97.8004 | 88.1797 |
0.0263 | 7.04 | 55000 | 0.0242 | 97.3222 | 88.5001 |
0.024 | 7.68 | 60000 | 0.0222 | 97.649 | 88.3117 |
0.0244 | 8.32 | 65000 | 0.0235 | 97.7653 | 88.1559 |
0.0282 | 8.96 | 70000 | 0.0245 | 97.7309 | 88.2283 |
0.0311 | 9.6 | 75000 | 0.0223 | 97.6517 | 88.2709 |
0.0321 | 10.24 | 80000 | 0.0230 | 97.6147 | 88.3084 |
Framework versions
- Transformers 4.34.0
- Pytorch 2.1.0+cu121
- Datasets 2.14.5
- Tokenizers 0.14.1