HPLT v2.0 - Cleaned - Slovak

This is one of the decoder-only language models trained on HPLT2.0_cleaned.

All the HPLT decoder-only models use the same hyper-parameters, roughly following the llama architecture with 2.15B parameters in total:

hidden size: 2048
attention heads: 32
layers: 24
sequence length: 2048

Intermediate checkpoints

We are releasing intermediate checkpoints for each model at intervals of every 1000 training steps in separate branches. The naming convention is checkpoint_00xxxx00: for example, checkpoint_0005000. The checkpoints range from checkpoint_0001000 to checkpoint_0047684 and the latter is in the main branch.

Cite us

@misc{burchell2025expandedmassivemultilingualdataset,
      title={An Expanded Massive Multilingual Dataset for High-Performance Language Technologies},
      author={Laurie Burchell and Ona de Gibert and Nikolay Arefyev and Mikko Aulamo and Marta Bañón and Pinzhen Chen and Mariia Fedorova and Liane Guillou and Barry Haddow and Jan Hajič and Jindřich Helcl and Erik Henriksson and Mateusz Klimaszewski and Ville Komulainen and Andrey Kutuzov and Joona Kytöniemi and Veronika Laippala and Petter Mæhlum and Bhavitvya Malik and Farrokh Mehryary and Vladislav Mikhailov and Nikita Moghe and Amanda Myntti and Dayyán O'Brien and Stephan Oepen and Proyag Pal and Jousia Piha and Sampo Pyysalo and Gema Ramírez-Sánchez and David Samuel and Pavel Stepachev and Jörg Tiedemann and Dušan Variš and Tereza Vojtěchová and Jaume Zaragoza-Bernabeu},
      year={2025},
      eprint={2503.10267},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2503.10267},
}

Downloads last month: 190

Inference Providers NEW

This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train HPLT/hplt2c_slk_checkpoints

Collection including HPLT/hplt2c_slk_checkpoints

HPLT 2.0 Monolingual reference models

Collection

This is a collection of decoder-only language models trained on HPLT2.0_cleaned. • 39 items • Updated 4 days ago • 8