ยท
AI & ML interests
Language, lingo, jargon, parlance...
Recent Activity
reacted to tardellirs's post with ๐ฅ 3 days ago Robotics is now the second most downloaded dataset category on the Hub, after text generation.
Robotics datasets got 13.7M downloads in September, ahead of text classification and question answering. Two years ago the category ranked 23rd. One in 7 new datasets is now robotics, mostly LeRobot recordings: typically a few dozen demos, about half of them on low-cost SO-100/SO-101 arms.
I found this after adding datasets and Spaces to Model Pulse, which rebuilds daily history from @cfahlgren1's hub-stats snapshots. Two more findings:
- In 2022, 36% of authors who list training data cited classic NLP sets like IMDb, SQuAD and GLUE. In 2026 it's 2.4%. Reasoning traces distilled from models like DeepSeek-R1 and Claude are now the most cited kind.
- In October 2025, 122K Spaces were created, 71K of them websites built with DeepSite. That's about 6x the monthly pace of late 2024, while likes given per month fell from about 35K to about 20K.
New in the app: a page for every dataset, with daily downloads and the models trained on it (636 list FineWeb), a page for every Space, and rankings for both.
Thank you to everyone who liked Model Pulse this week: it made Spaces of the Week and is #7 on trending. Thanks also to @dipankarsarkar, whose comments on the last post fixed three data issues. If a number looks wrong, tell me.
Spaces: https://huggingface.co/spaces/tardellirs/model-pulse reacted to Undi95's post with ๐ฅ 5 days ago Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek.
I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use?
No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note.
The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match.
Later, code will be checked by actually running tests.
The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights.
Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again.
The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move.
I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks.
Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try.
Did you already tried something like that? What was your result? I'm curious! View all activity Organizations