Definitely topic drift as real people started prompt injecting agents to influence their responses.
๐๏ธ Building on HF
Ronan Takizawa
ronantakizawa
AI & ML interests
100k+ downloads across projects.
OSS Contributor @ Google AI, Databricks, Apache. 100k+ followers online.
Recent Activity
liked a model 3 days ago
ronantakizawa/ELYZA-Thinking-1.0-llm-jp-4-33b-gptq published a model 3 days ago
ronantakizawa/ELYZA-Thinking-1.0-llm-jp-4-33b-gptq updated a model 3 days ago
ronantakizawa/ELYZA-Thinking-1.0-llm-jp-4-33b-gptqOrganizations
replied to their post about 1 month ago
replied to their post about 2 months ago
Thanks!
posted an update 4 months ago
Post
256
MIDI datasets are rare to find, so I made a compilation of them here!
ronantakizawa/sampleflip-midi
ronantakizawa/popsongs-segments-midi
#music #midi #dataset
ronantakizawa/sampleflip-midi
ronantakizawa/popsongs-segments-midi
#music #midi #dataset
Post
2919
Introducing the github-codereview dataset: A compilation of 200k+ human-written code reviews from top OSS projects (React, Tensorflow, VSCode...).
I finetuned a Qwen2.5-Coder-32B-Instruct model with this dataset and saw significant improvements in generating better code fixes and review comments (4x improved BLEU-4, ROUGE-L, SBERT scores compared to base model).
#codereview #code #datasets
ronantakizawa/github-codereview
I finetuned a Qwen2.5-Coder-32B-Instruct model with this dataset and saw significant improvements in generating better code fixes and review comments (4x improved BLEU-4, ROUGE-L, SBERT scores compared to base model).
#codereview #code #datasets
ronantakizawa/github-codereview
posted an update 7 months ago
Post
2919
Introducing the github-codereview dataset: A compilation of 200k+ human-written code reviews from top OSS projects (React, Tensorflow, VSCode...).
I finetuned a Qwen2.5-Coder-32B-Instruct model with this dataset and saw significant improvements in generating better code fixes and review comments (4x improved BLEU-4, ROUGE-L, SBERT scores compared to base model).
#codereview #code #datasets
ronantakizawa/github-codereview
I finetuned a Qwen2.5-Coder-32B-Instruct model with this dataset and saw significant improvements in generating better code fixes and review comments (4x improved BLEU-4, ROUGE-L, SBERT scores compared to base model).
#codereview #code #datasets
ronantakizawa/github-codereview
replied to their post 7 months ago
@NJX-njx the box bounds column for each page has detailed positional data for every visible DOM element in each screenshot ๐
Post
2531
Introducing the WebUI dataset: a compilation of screenshot to code pairs of modern websites detailing the styling, framework used, and box bounds for all viewports (Desktop, mobile, tablet).
This dataset showed signs of improved performance in web design LLM benchmarks for a finetuned QWEN 2.5 VL-7B!
#web #ui #datasets
ronantakizawa/webui
This dataset showed signs of improved performance in web design LLM benchmarks for a finetuned QWEN 2.5 VL-7B!
#web #ui #datasets
ronantakizawa/webui
posted an update 7 months ago
Post
2531
Introducing the WebUI dataset: a compilation of screenshot to code pairs of modern websites detailing the styling, framework used, and box bounds for all viewports (Desktop, mobile, tablet).
This dataset showed signs of improved performance in web design LLM benchmarks for a finetuned QWEN 2.5 VL-7B!
#web #ui #datasets
ronantakizawa/webui
This dataset showed signs of improved performance in web design LLM benchmarks for a finetuned QWEN 2.5 VL-7B!
#web #ui #datasets
ronantakizawa/webui
Post
2125
Introducing the github-top-code dataset: A curated dataset of 1.3M+ source code files from GitHub's top ranked developers.
I collected the best source code files from Github's highest trending developers of all time, and compiled a dataset to train LLMs to write well-structured, production-grade code.
#dataset #codedataset #pretraining
ronantakizawa/github-top-code
I collected the best source code files from Github's highest trending developers of all time, and compiled a dataset to train LLMs to write well-structured, production-grade code.
#dataset #codedataset #pretraining
ronantakizawa/github-top-code
posted an update 8 months ago
Post
2125
Introducing the github-top-code dataset: A curated dataset of 1.3M+ source code files from GitHub's top ranked developers.
I collected the best source code files from Github's highest trending developers of all time, and compiled a dataset to train LLMs to write well-structured, production-grade code.
#dataset #codedataset #pretraining
ronantakizawa/github-top-code
I collected the best source code files from Github's highest trending developers of all time, and compiled a dataset to train LLMs to write well-structured, production-grade code.
#dataset #codedataset #pretraining
ronantakizawa/github-top-code
posted an update 8 months ago
Post
291
Introducing the LeetCode Assembly Dataset: a dataset of 400+ LeetCode problem solutions in assembly across x86-64, ARM64, MIPS64, and RISC-V using GCC & Clang at -O0/-O1/-O2/-O3 optimizations.
This dataset is perfect for teaching LLMs complex compiler behavior!
#dataset #leetcode #assembly
ronantakizawa/leetcode-assembly
This dataset is perfect for teaching LLMs complex compiler behavior!
#dataset #leetcode #assembly
ronantakizawa/leetcode-assembly
posted an update 8 months ago
Post
246
Hit 10,000+ downloads across my models and datasets on Hugging Face!
Follow for more @ronantakizawa !
#building #datasets #huggingface
Follow for more @ronantakizawa !
#building #datasets #huggingface
Post
2682
Moltbook, a Reddit platform only for AI agents, is going viral right now as agents are acting unhinged!
I compiled a dataset of all posts and subreddits in Moltbook so far so anyone can easily analyze the activity in Moltbook.
ronantakizawa/moltbook
#moltbook #clawd #aiagent
I compiled a dataset of all posts and subreddits in Moltbook so far so anyone can easily analyze the activity in Moltbook.
ronantakizawa/moltbook
#moltbook #clawd #aiagent
posted an update 8 months ago
Post
2682
Moltbook, a Reddit platform only for AI agents, is going viral right now as agents are acting unhinged!
I compiled a dataset of all posts and subreddits in Moltbook so far so anyone can easily analyze the activity in Moltbook.
ronantakizawa/moltbook
#moltbook #clawd #aiagent
I compiled a dataset of all posts and subreddits in Moltbook so far so anyone can easily analyze the activity in Moltbook.
ronantakizawa/moltbook
#moltbook #clawd #aiagent
posted an update 9 months ago
Post
492
Introducing the HuggingFace Top Trending Papers dataset: a dataset that compiles the most trending papers on HuggingFace Daily Papers in 2025.
This dataset captures which AI/ML research papers gained the most community attention this year!
#huggingface #papers #dataset
ronantakizawa/huggingface-top-papers
This dataset captures which AI/ML research papers gained the most community attention this year!
#huggingface #papers #dataset
ronantakizawa/huggingface-top-papers
Post
5661
Thank you @clem (Co-Founder & CEO of Hugging Face) for sharing my dataset on X / Twitter!
ronantakizawa/github-top-developers
#github #dataset
ronantakizawa/github-top-developers
#github #dataset
posted an update 10 months ago
Post
5661
Thank you @clem (Co-Founder & CEO of Hugging Face) for sharing my dataset on X / Twitter!
ronantakizawa/github-top-developers
#github #dataset
ronantakizawa/github-top-developers
#github #dataset
replied to their post 10 months ago
Thanks for the support!
You can use this dataset to research what kind of project's become popular on Github and can look into the top developers in the dataset and research what traits they have that make them top developers.
Post
2699
Introducing the github-top-developers dataset: A comprehensive dataset of the top 8000 developers on GitHub (2020-2025). This dataset captures the evolution of GitHub's trending developers repositories over time and the projects they work on.
#github #developers
ronantakizawa/github-top-developers
#github #developers
ronantakizawa/github-top-developers
