Instructions to use eric-ml-nlp/Delveta-LayaChoice-v2-checkpoints with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use eric-ml-nlp/Delveta-LayaChoice-v2-checkpoints with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Delveta LayaChoice v2 β historical training checkpoints
Scope. This repository is intended to hold only the non-selected historical training checkpoints of the V2 run (epochs 1β6), plus a record of the full 10-epoch training history.
- The selected model is Epoch 7 and is published separately at
eric-ml-nlp/Delveta-LayaChoice-v2.- This repository is not the inference entry point. Use the selected model repository for inference.
- Epochs 1β6 are historical training artifacts. They are not independent evaluation results and are not claimed to be better than Epoch 7; Epoch 7 was selected because it had the highest validation top-1.
Publication status. The epoch 1β6 weight files are published in this repository (see Checkpoint Inventory and SHA256 checksums). Weight files for epochs 7β10 are intentionally not included β Epoch 7 is published as the selected model, and epochs 8β10 are recorded here as history only.
1. Repository Purpose
This repository exists so the training history of the V2 run is auditable: it records every epoch's training loss and validation metrics, identifies which checkpoint was selected, and (once staged) publishes the non-selected epoch weights that were kept.
It complements:
- the selected model (Epoch 7): https://huggingface.co/eric-ml-nlp/Delveta-LayaChoice-v2
- the frozen dataset: https://huggingface.co/datasets/eric-ml-nlp/Delveta-LayaChoice-v2-Data
2. Relationship to the Selected Model
| Selected checkpoint | Epoch 7 |
| Selection rule | best_epoch = argmax(validation top-1), ties resolved to the earlier epoch; train loss never used |
| Epoch-7 validation top-1 | 290/300 = 96.67 % |
| Selected model repository | https://huggingface.co/eric-ml-nlp/Delveta-LayaChoice-v2 |
Epochs 8β10 are not better than Epoch 7: their validation top-1 is 96.33 %, below Epoch 7's 96.67 %.
3. Checkpoint Inventory
The epoch 1β6 model.safetensors weights are published. Each is the raw fine-tuned
weight file for that epoch; sizes and digests are the real published values. Epoch 7's
validation top-1 (96.67 %) was the highest, so Epoch 7 was selected.
| Epoch | Path | Size | Validation top-1 |
|---|---|---|---|
| 1 | epoch-1/model.safetensors |
1.29 GB | 85.67% |
| 2 | epoch-2/model.safetensors |
1.29 GB | 90.00% |
| 3 | epoch-3/model.safetensors |
1.29 GB | 94.00% |
| 4 | epoch-4/model.safetensors |
1.29 GB | 89.67% |
| 5 | epoch-5/model.safetensors |
1.29 GB | 94.67% |
| 6 | epoch-6/model.safetensors |
1.29 GB | 94.33% |
Epochs 7β10 weight files are intentionally out of scope for this repository: Epoch 7 is the selected model (published in the model repository) and Epochs 8β10 are recorded here as history only. No weight file from epochs 7β10 is placed here, and no epoch's file was copied or renamed to stand in for another epoch.
4. Validation Metrics by Epoch
Validation v3, 300 rows, card view B_noprov. Top-1 is invariant under the calibration
temperature. This table records the training history; it does not imply this
repository holds the weights of every epoch.
| Epoch | Train Loss | Val Top-1 | Val ECE | REJECT Recall | REJECT FPR |
|---|---|---|---|---|---|
| 1 | 0.641706 | 85.67% | 0.080275 | 81.25% | 8.45% |
| 2 | 0.384151 | 90.00% | 0.036385 | 93.75% | 8.10% |
| 3 | 0.256366 | 94.00% | 0.056254 | 93.75% | 2.46% |
| 4 | 0.132143 | 89.67% | 0.101781 | 100.00% | 8.10% |
| 5 | 0.079365 | 94.67% | 0.054404 | 100.00% | 2.82% |
| 6 | 0.039875 | 94.33% | 0.056691 | 100.00% | 2.82% |
| 7 | 0.013119 | 96.67% | 0.030319 | 100.00% | 1.76% |
| 8 | 0.000358 | 96.33% | 0.035276 | 100.00% | 2.11% |
| 9 | 0.000000 | 96.33% | 0.035276 | 100.00% | 2.11% |
| 10 | 0.000000 | 96.33% | 0.035276 | 100.00% | 2.11% |
- Epoch 7 is the selected model (validation top-1 = 96.67 % = 290/300).
- Validation
REJECTcounts are small (16 goldREJECTrows per epoch), so per-epochREJECTrecall/FPR are noisy and should not be read as precise estimates. - Epochs 9β10 are identical to Epoch 8 because the training loss had reached 0 and the weights stopped moving.
5. Training Provenance
| Base model | convaiinnovations/laya |
| Base revision | 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851 |
| Base subfolder | multilingual |
| Encoder | jhu-clsp/mmBERT-base |
| Method | full-parameter fine-tuning (no adapters) |
| Runtime | laya==0.3.21 |
| Dataset | https://huggingface.co/datasets/eric-ml-nlp/Delveta-LayaChoice-v2-Data |
| Candidate count | K = 3 |
| Card representation | B_noprov |
| Decision setup | four-way decision including REJECT |
Optimizer details, exact hardware and wall-clock training time are not asserted here β they are not recorded in the artifacts this card was built from. See the selected model card for the training configuration that is recorded.
6. File Contents and Resumability
Each epoch-N/ directory is expected to hold a model.safetensors weight file. A
weight file alone is not a fully resumable training checkpoint: full resumption
requires the optimizer state, LR scheduler state, step counter and RNG state from the
same step. Unless those are also published alongside the weights, this repository
supports inference-style weight loading, not resumption of training from the
mid-run state.
7. SHA256 Checksums
30f28e23cf8b3e0fa76fd7b1f74337b49d69b6eec3932e21ebc221314be8203f epoch-1/model.safetensors
ea0f866aed715141733b25b00a643f46392b3a61230ee55e399ce1688866c2d2 epoch-2/model.safetensors
27dd2756ec2a7547555fcae2b2acf6d37cc10d456a132fc6754357f1487fcb5b epoch-3/model.safetensors
e04b5375cf46273744d7a6e47698a258278735d74c0061eed227921fc80bc56f epoch-4/model.safetensors
4572dd2d3986b58476685bdc9151458a308d15808904a54d9b09b39120d832dc epoch-5/model.safetensors
a2f1046c5f8068ac13afe86516bc5ece7961eea410adabf590c1ca829218a207 epoch-6/model.safetensors
These are the SHA256 digests of the published weight files, cross-checked against the
published repository's LFS metadata. Epoch 7's published weight is the same bytes as
the model repository's model.safetensors (c6331bd5β¦).
8. Reproduction Notes
- To reproduce the selected model's behaviour, use Epoch 7 from
Delveta-LayaChoice-v2, not a checkpoint from this repository. - The training/validation metrics above were read from the run's per-epoch metric records, not reconstructed from memory.
- The dataset used for training is frozen and published (see Β§5).
9. Limitations
- Historical checkpoints are superseded by the selected Epoch 7 model and are provided for audit / research only.
- Validation
REJECTmetrics are computed over only 16 goldREJECTrows per epoch and are correspondingly uncertain. - Validation top-1 is a 300-row estimate; small differences between epochs (e.g. 96.33 % vs 96.67 %) are within the noise of that split.
- No claim is made that any unselected epoch generalizes better than Epoch 7.
10. License
No license is asserted for these weights. The laya software package is
Apache-2.0, but the license of the base model weights
(convaiinnovations/laya) could not be determined at release time. Downstream use
should confirm terms with the upstream rights holder.
- Downloads last month
- -
Model tree for eric-ml-nlp/Delveta-LayaChoice-v2-checkpoints
Base model
convaiinnovations/laya