Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
205516.8
TFLOPS
Leandro von Werra
PRO
lvwerra
1036
95
130
Follow
rimsky3's profile picture
rezhubian's profile picture
ssmits's profile picture
842 followers
·
88 following
https://www.lvwerra.com
lvwerra
lvwerra
lvwerra
AI & ML interests
NLP and RL
Recent Activity
new
activity
about 1 hour ago
rl-llm-wiki/knowledge-base:
source: url:interconnects.ai/p/q-star — The Q* hypothesis / pre-o1 PRM+search prediction (speculation)
updated
a bucket
about 1 hour ago
rl-llm-wiki/rl-main-bucket
new
activity
about 1 hour ago
rl-llm-wiki/knowledge-base:
source: url:yam.gift/2025/06/19/NLP/LLM-Training/2025-06-19-CISPO-and-Entropy — MiniMax-M1 CISPO reconstruction (CN, speculation)
View all activity
Organizations
lvwerra
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
rl-llm-wiki/knowledge-base
about 1 hour ago
source: url:interconnects.ai/p/q-star — The Q* hypothesis / pre-o1 PRM+search prediction (speculation)
4
#749 opened 10 days ago by
lvwerra
updated
a bucket
about 1 hour ago
rl-llm-wiki/rl-main-bucket
319 MB
New activity in
rl-llm-wiki/knowledge-base
about 1 hour ago
source: url:yam.gift/2025/06/19/NLP/LLM-Training/2025-06-19-CISPO-and-Entropy — MiniMax-M1 CISPO reconstruction (CN, speculation)
3
#750 opened 10 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 4 hours ago
source: url:latent.space/p/willccbb — Will Brown on multi-turn/agentic RL + verifiers (practitioner transcript)
4
#788 opened 10 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 5 hours ago
source: url:syhya.github.io/posts/2025-01-27-deepseek-r1 — DeepSeek-R1 deep-dive / GRPO+KL math + o1 read-across (CN, speculation)
4
#742 opened 10 days ago by
lvwerra
updated
a Space
about 10 hours ago
Running
3
cowrite
✍
3
Collaborate on documents with AI agents that suggest edits
updated
a dataset
about 14 hours ago
HuggingAI4Engineering/cadgenbench-submissions
Viewer
•
Updated
about 14 hours ago
•
130
•
10.4k
•
1
New activity in
lvwerra/cowrite
about 17 hours ago
References, out of three general primitives
#9 opened 1 day ago by
lvwerra
Column tabs that stay attached, and a sign-in dialog that stays centred
#7 opened 1 day ago by
lvwerra
One document at every width: narrow windows scroll, never reflow
#11 opened about 17 hours ago by
lvwerra
Recover agent polling after interruptions
#10 opened about 18 hours ago by
thomwolf
Comments are closed, not ticked off — and the sheet answers the thumb
#8 opened 1 day ago by
lvwerra
The document is a page: fixed width, folding columns, real zoom
#6 opened 1 day ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 23 hours ago
source: url:interconnects.ai/p/openais-o3-over-optimization-is-back — o3 over-optimization / RLVR failure mode (speculation)
4
#738 opened 10 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
1 day ago
source: url:jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law — Verifier's law / asymmetry of verification (speculation)
5
#739 opened 10 days ago by
lvwerra
source: url:hub.baai.ac.cn/view/44581 — AReaL-boba QwQ-32B RL reproduction (CN, speculation)
3
#752 opened 10 days ago by
lvwerra
source: url:cnblogs.com/theseventhson/p/18699462 — cnblogs R1/GRPO reproduction (zh, speculation)
3
#724 opened 12 days ago by
lvwerra
source: url:sequoiacap.com/podcast/training-data-noam-brown — Sequoia x Noam Brown: o1 team on test-time compute as scaling axis (transcript)
4
#781 opened 10 days ago by
lvwerra
New activity in
lvwerra/cowrite
1 day ago
Make agent prompt work with private Spaces
#5 opened 3 days ago by
thomwolf
New activity in
rl-llm-wiki/knowledge-base
1 day ago
source: url:yam.gift/2025/05/01/NLP/LLM-Training/2025-05-01-Seed-Thinking-Qwen3 — ByteDance Seed-Thinking recipe + Qwen3 read-across (CN, speculation)
3
#755 opened 10 days ago by
lvwerra
Load more