Β·
AI & ML interests
Mechanistic Interpretability (MI) Research & sp00ky code stuff
Recent Activity
reacted to theirpost with π₯ about 13 hours ago What can you actually build with a cybersecurity dataset?
I've been updating a few of mine on Hugging Face, and they now cover some pretty different parts of the security workflow.
- open malsec has 1,104 defensive security scenarios across 20 subsets covering phishing, malware, scams, cloud security, API security, AI security and more
- opensec triage has 50,000 contextual alert examples, plus compact model and edge training sets for testing whether models classify from the evidence around an event
- infosec tool output has 1,004 examples across 19 tools for turning raw security output into evidence backed explanations, limitations and defensive next steps
You could use them for:
* phishing and scam explainers
* alert triage tools
* SOC assistants
* scanner output explainers
* analyst training
* model comparisons
* grounding and hallucination tests
* small specialised security models
* edge and local model experiments
Or combine them into something like:
scenario β evidence β triage β explanation β next action
You also don't need to train anything straight away.
Grab a few examples, run them through whatever model you already use and see where it gets confused :)
Datasets:
https://huggingface.co/datasets/tegridydev/open-malsec
https://huggingface.co/datasets/tegridydev/opensec-triage
https://huggingface.co/datasets/tegridydev/infosec-tool-output
reacted to theirpost with π§ about 13 hours ago What can you actually build with a cybersecurity dataset?
I've been updating a few of mine on Hugging Face, and they now cover some pretty different parts of the security workflow.
- open malsec has 1,104 defensive security scenarios across 20 subsets covering phishing, malware, scams, cloud security, API security, AI security and more
- opensec triage has 50,000 contextual alert examples, plus compact model and edge training sets for testing whether models classify from the evidence around an event
- infosec tool output has 1,004 examples across 19 tools for turning raw security output into evidence backed explanations, limitations and defensive next steps
You could use them for:
* phishing and scam explainers
* alert triage tools
* SOC assistants
* scanner output explainers
* analyst training
* model comparisons
* grounding and hallucination tests
* small specialised security models
* edge and local model experiments
Or combine them into something like:
scenario β evidence β triage β explanation β next action
You also don't need to train anything straight away.
Grab a few examples, run them through whatever model you already use and see where it gets confused :)
Datasets:
https://huggingface.co/datasets/tegridydev/open-malsec
https://huggingface.co/datasets/tegridydev/opensec-triage
https://huggingface.co/datasets/tegridydev/infosec-tool-output
reacted to theirpost with π about 13 hours ago What can you actually build with a cybersecurity dataset?
I've been updating a few of mine on Hugging Face, and they now cover some pretty different parts of the security workflow.
- open malsec has 1,104 defensive security scenarios across 20 subsets covering phishing, malware, scams, cloud security, API security, AI security and more
- opensec triage has 50,000 contextual alert examples, plus compact model and edge training sets for testing whether models classify from the evidence around an event
- infosec tool output has 1,004 examples across 19 tools for turning raw security output into evidence backed explanations, limitations and defensive next steps
You could use them for:
* phishing and scam explainers
* alert triage tools
* SOC assistants
* scanner output explainers
* analyst training
* model comparisons
* grounding and hallucination tests
* small specialised security models
* edge and local model experiments
Or combine them into something like:
scenario β evidence β triage β explanation β next action
You also don't need to train anything straight away.
Grab a few examples, run them through whatever model you already use and see where it gets confused :)
Datasets:
https://huggingface.co/datasets/tegridydev/open-malsec
https://huggingface.co/datasets/tegridydev/opensec-triage
https://huggingface.co/datasets/tegridydev/infosec-tool-output
View all activity Organizations
view article So, letβs make our own dataset
tegridydev
β’ β’ 4
view article minecraft time with astra | [td]
tegridydev
β’ β’ 1
view article π Research Papers Dataset
tegridydev
β’ β’ 5
published an article over 1 year ago view article Open Source AI Agents | Github/Repo List | [2025]
tegridydev
β’ β’ 31
published an article over 1 year ago view article WTF is Fine-Tuning? (intro4devs) | [2025]
tegridydev
β’ β’ 7
published an article over 1 year ago view article LLM Dataset Formats 101: A NoβBS Guide for Hugging Face Devs
tegridydev
β’ β’ 16