Forest
Deventhedude
AI & ML interests
None yet
Recent Activity
updated a bucket 10 days ago
Deventhedude/parakeet-redux-bucket published a bucket 10 days ago
Deventhedude/parakeet-redux-bucket updated a collection about 2 months ago
computer useOrganizations
computer use
grounding data
Distillation
web_use
-
tiiuae/viscon-1m
Viewer • Updated • 1.01M • 87 • 2 -
tiiuae/falcon-refinedweb
Viewer • Updated • 968M • 83.4k • 966 -
The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
Paper • 2306.01116 • Published • 45 -
openbmb/Ultra-FineWeb
Viewer • Updated • 1.29B • 67.2k • 450
READ
-
Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling
Paper • 2601.02346 • Published • 28 -
unsloth/alpaca-cleaned
Viewer • Updated • 51.8k • 2.74k • 29 -
Hierarchical Reasoning Model
Paper • 2506.21734 • Published • 54 -
Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
Paper • 2507.07955 • Published • 27
Long context
Vision
CyberSecurity
Benchmarks
Embedding Fintune Datasets
-
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
Paper • 2310.19923 • Published • 15 -
allenai/c4
Viewer • Updated • 10.4B • 925k • 677 -
jina-reranker-v3: Last but Not Late Interaction for Document Reranking
Paper • 2509.25085 • Published • 13 -
jinaai/negation-dataset
Viewer • Updated • 10.5k • 173 • 24
AXM finetune
-
Learning a Canonical Basis of Human Preferences from Binary Ratings
Paper • 2503.24150 • Published • 1 -
google-research-datasets/go_emotions
Viewer • Updated • 265k • 18.6k • 269 -
id4thomas/emotion-prediction-comet-atomic-2020
Viewer • Updated • 134k • 85 -
filnow/LAIGAI-Emotion-Prediction
Viewer • Updated • 480 • 5
Tools dataset
reasoning
Terminus
Mobile
Quantization
Video
V3_SCIT_ANCHORS
Agent
Time Datasets
TRADING/FINANCES DATASETS
Safety Training
Multi Language
coding data
Finetune data
-
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
Paper • 2505.10597 • Published -
COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values
Paper • 2504.05535 • Published • 44 -
nvidia/HelpSteer3
Viewer • Updated • 133k • 9.11k • 119 -
nvidia/Nemotron-RL-instruction_following
RL Environment • Updated • 1.08k • 23
human_mouse_movment
reasoning
computer use
Terminus
grounding data
Mobile
Distillation
Quantization
web_use
-
tiiuae/viscon-1m
Viewer • Updated • 1.01M • 87 • 2 -
tiiuae/falcon-refinedweb
Viewer • Updated • 968M • 83.4k • 966 -
The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
Paper • 2306.01116 • Published • 45 -
openbmb/Ultra-FineWeb
Viewer • Updated • 1.29B • 67.2k • 450
Video
READ
-
Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling
Paper • 2601.02346 • Published • 28 -
unsloth/alpaca-cleaned
Viewer • Updated • 51.8k • 2.74k • 29 -
Hierarchical Reasoning Model
Paper • 2506.21734 • Published • 54 -
Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
Paper • 2507.07955 • Published • 27
V3_SCIT_ANCHORS
Long context
Agent
Vision
Time Datasets
CyberSecurity
TRADING/FINANCES DATASETS
Benchmarks
Safety Training
Embedding Fintune Datasets
-
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
Paper • 2310.19923 • Published • 15 -
allenai/c4
Viewer • Updated • 10.4B • 925k • 677 -
jina-reranker-v3: Last but Not Late Interaction for Document Reranking
Paper • 2509.25085 • Published • 13 -
jinaai/negation-dataset
Viewer • Updated • 10.5k • 173 • 24
Multi Language
AXM finetune
-
Learning a Canonical Basis of Human Preferences from Binary Ratings
Paper • 2503.24150 • Published • 1 -
google-research-datasets/go_emotions
Viewer • Updated • 265k • 18.6k • 269 -
id4thomas/emotion-prediction-comet-atomic-2020
Viewer • Updated • 134k • 85 -
filnow/LAIGAI-Emotion-Prediction
Viewer • Updated • 480 • 5
coding data
Tools dataset
Finetune data
-
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
Paper • 2505.10597 • Published -
COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values
Paper • 2504.05535 • Published • 44 -
nvidia/HelpSteer3
Viewer • Updated • 133k • 9.11k • 119 -
nvidia/Nemotron-RL-instruction_following
RL Environment • Updated • 1.08k • 23