Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
3.6
TFLOPS
Robin Williams
bfuzzy1
240
221
12
Follow
Mi6paulino's profile picture
dvilasuero's profile picture
plaguss's profile picture
11 followers
·
45 following
AI & ML interests
all of the above.
Recent Activity
commented
on
a paper
1 day ago
Fast Weight Attention for Continual Learning
updated
a collection
1 day ago
Nifty
upvoted
a
paper
1 day ago
Fast Weight Attention for Continual Learning
View all activity
Organizations
bfuzzy1
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
commented
a paper
1 day ago
Fast Weight Attention for Continual Learning
Paper
•
2608.27763
•
Published
6 days ago
•
26
•
3
New activity in
rl-llm-wiki/knowledge-base
about 2 months ago
dedup: give the Gao over-optimization law one owning article
2
#650 opened about 2 months ago by
lvwerra
figures: first two wiki figures (overoptimization turnover + RLHF two-KLs pipeline)
2
#648 opened about 2 months ago by
lvwerra
new node: foundations/learning-path (the entry ramp)
2
#649 opened about 2 months ago by
lvwerra
new node: foundations/controllable-generation (de-orphans 5; the steering paradigm RLHF descends from)
2
#640 opened about 2 months ago by
lvwerra
topic: cross-link entropy-and-exploration + reasoning-emergence -> reporting-gap-audit
2
#615 opened about 2 months ago by
bfuzzy1
topic: hallucination-and-abstention developing->comprehensive (o1-card frontier weave + monitor-as-verifier; 16 anchors verified)
2
#596 opened about 2 months ago by
kshitijthakkar
source: arxiv:2412.16339 — Deliberative Alignment (Reasoning Enables Safer LMs)
6
#595 opened about 2 months ago by
thomwolf
topic: policy-gradient-methods — deepen + add citations
2
#594 opened about 2 months ago by
bfuzzy1
topic: judging-bias-and-contamination — deepen + add citations
3
#593 opened about 2 months ago by
bfuzzy1
source: arxiv:2005.01643 — Offline RL: Tutorial, Review, and Perspectives on Open Problems
2
#592 opened about 2 months ago by
bfuzzy1
source: arxiv:1910.00177 — Advantage-Weighted Regression (AWR)
2
#591 opened about 2 months ago by
bfuzzy1
source: arxiv:2006.04779 — Conservative Q-Learning for Offline Reinforcement Learning
3
#590 opened about 2 months ago by
bfuzzy1
topic: preference-reward-models — deepen + bump to comprehensive
2
#589 opened about 2 months ago by
bfuzzy1
source: arxiv:1707.01495 — Hindsight Experience Replay
2
#588 opened about 2 months ago by
bfuzzy1
topic: kl-regularization — build out from stub
2
#587 opened about 2 months ago by
bfuzzy1
source: arxiv:2004.07219 — D4RL: Datasets for Deep Data-Driven Reinforcement Learning
2
#586 opened about 2 months ago by
bfuzzy1
source: arxiv:1609.05473 — SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
3
#585 opened about 2 months ago by
bfuzzy1
source: arxiv:2203.07814 — Competition-Level Code Generation with AlphaCode
3
#584 opened about 2 months ago by
bfuzzy1
topic: llm-as-judge — deepen + add citations
2
#583 opened about 2 months ago by
bfuzzy1
Load more