Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines Paper • 2605.01077 • Published May 1 • 2
Sabiá-2: A New Generation of Portuguese Large Language Models Paper • 2403.09887 • Published Mar 14, 2024
TiEBe: A Benchmark for Assessing the Current Knowledge of Large Language Models Paper • 2501.07482 • Published Jan 13, 2025
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data Paper • 2510.10159 • Published Oct 11, 2025 • 3
Measuring what Matters: Construct Validity in Large Language Model Benchmarks Paper • 2511.04703 • Published Nov 3, 2025 • 8
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment Paper • 2601.12910 • Published Jan 19 • 3
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment Paper • 2601.12910 • Published Jan 19 • 3
Structured Extraction from Business Process Diagrams Using Vision-Language Models Paper • 2511.22448 • Published Nov 27, 2025 • 2
Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings Paper • 2509.14405 • Published Sep 17, 2025 • 2
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Paper • 2506.22439 • Published May 29, 2025 • 3
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments Paper • 2509.14233 • Published Sep 17, 2025 • 24
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America Paper • 2507.00999 • Published Jul 1, 2025 • 3
The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors Paper • 2509.04484 • Published Aug 31, 2025 • 2
LCFO: Long Context and Long Form Output Dataset and Benchmarking Paper • 2412.08268 • Published Dec 11, 2024
Large Concept Models: Language Modeling in a Sentence Representation Space Paper • 2412.08821 • Published Dec 11, 2024 • 18
Exploring Methods for Cross-lingual Text Style Transfer: The Case of Text Detoxification Paper • 2311.13937 • Published Nov 23, 2023 • 1
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation Paper • 2502.04314 • Published Feb 6, 2025
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation Paper • 2504.07072 • Published Apr 9, 2025 • 9
It's the same but not the same: Do LLMs distinguish Spanish varieties? Paper • 2504.20049 • Published Apr 8, 2025