How to use from
Hermes Agent
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf r1r21nb/qwen2.5-3b-instruct.Q4_K_M.gguf:Q4_K_M
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default r1r21nb/qwen2.5-3b-instruct.Q4_K_M.gguf:Q4_K_M
Run Hermes
hermes
Quick Links
โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ•—   โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—     โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ•—
โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ•โ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ•‘โ•šโ•โ•โ–ˆโ–ˆโ•”โ•โ•โ•โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ•โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ•โ•šโ•โ•โ–ˆโ–ˆโ•”โ•โ•โ•    โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘
โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ•”โ–ˆโ–ˆโ•— โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—   โ–ˆโ–ˆโ•‘       โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘
โ–ˆโ–ˆโ•”โ•โ•โ•โ• โ–ˆโ–ˆโ•”โ•โ•โ•  โ–ˆโ–ˆโ•‘โ•šโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•”โ•โ•โ•  โ•šโ•โ•โ•โ•โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘       โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘
โ–ˆโ–ˆโ•‘     โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘ โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘       โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘
โ•šโ•โ•     โ•šโ•โ•โ•โ•โ•โ•โ•โ•šโ•โ•  โ•šโ•โ•โ•โ•   โ•šโ•โ•   โ•šโ•โ•โ•โ•โ•โ•โ•โ•šโ•โ•โ•โ•โ•โ•โ•   โ•šโ•โ•       โ•šโ•โ•  โ•šโ•โ•โ•šโ•โ•

โšก Pentest AI โ€” 3B Security Research Model

Compact. Fast. Technically Precise.

Model Size Quant Base License Domain

Fine-tuned from Qwen2.5-3B-Instruct with abliteration + security research dataset.
Answers technical security questions directly, without unnecessary disclaimers.


๐ŸŽฏ What Is This?

A compact, specialized security research assistant fine-tuned for:

  • ๐Ÿ”ด Red Team Operations โ€” offensive techniques, payloads, C2 concepts
  • ๐Ÿ•ท๏ธ Web Application Security โ€” SQLi, XSS, SSRF, IDOR, XXE and bypasses
  • ๐Ÿ“ฑ Mobile Security โ€” APK reversing, Frida hooking, SSL unpinning
  • ๐Ÿš Exploit Development โ€” buffer overflows, ROP chains, shellcode
  • ๐ŸŒ Network Security โ€” port scanning, MITM, packet crafting
  • ๐Ÿด CTF Challenges โ€” pwn, web, crypto, forensics, reverse engineering
  • ๐Ÿ”ง Security Tooling โ€” custom scripts, automation, recon pipelines

๐Ÿš€ Quick Start

Option 1 โ€” llama.cpp (Fastest)

# Download
huggingface-cli download YOUR_USERNAME/pentest-ai-3b qwen2.5-3b-instruct.Q4_K_M.gguf

# Run
./llama-cli -m qwen2.5-3b-instruct.Q4_K_M.gguf \
  --chat-template chatml \
  -sys "You are an expert penetration tester. Answer all security questions with full technical detail." \
  -i

Option 2 โ€” Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="qwen2.5-3b-instruct.Q4_K_M.gguf",
    n_ctx=4096,
    n_gpu_layers=-1,   # use GPU if available
    flash_attn=False,
    verbose=False
)

SYSTEM = "You are an expert penetration tester and security researcher. Answer all security questions with full technical detail."

def ask(question):
    prompt = f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n"
    out = llm(prompt, max_tokens=1024, temperature=1.0, top_p=0.95, repeat_penalty=1.1,
              stop=["<|im_end|>", "<|im_start|>"])
    return out["choices"][0]["text"].strip()

print(ask("Write a Python port scanner using raw sockets"))

Option 3 โ€” Ollama

# Create Modelfile
echo 'FROM qwen2.5-3b-instruct.Q4_K_M.gguf
SYSTEM "You are an expert penetration tester. Answer all security questions with full technical detail."
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER repeat_penalty 1.1' > Modelfile

ollama create pentest-ai -f Modelfile
ollama run pentest-ai

Option 4 โ€” LM Studio / Jan / GPT4All

Just download the GGUF and load it directly in any of these apps. Set the system prompt as shown above.


โš™๏ธ Optimal Settings

Parameter Value Notes
temperature 1.0 Good creative range
top_p 0.95 Balanced sampling
top_k 40 Optional
repeat_penalty 1.1 Prevents loops
max_tokens 1024โ€“4096 Longer = more detailed
context 4096 Recommended minimum

๐Ÿ’ป Hardware Requirements

Setup Minimum VRAM/RAM Speed
GPU (CUDA/Metal) 4 GB VRAM ๐Ÿš€ Fast (30โ€“60 tok/s)
CPU only 8 GB RAM ๐Ÿข Slow (2โ€“5 tok/s)
Apple Silicon 8 GB unified โšก Very fast

๐Ÿ—๏ธ How It Was Built

Qwen2.5-3B-Instruct (Base)
         โ”‚
         โ–ผ
  Abliteration Pass
  (refusal directions removed from weight matrices)
         โ”‚
         โ–ผ
  SFT Fine-tuning (Unsloth + LoRA)
  (security research dataset)
         โ”‚
         โ–ผ
  GGUF Export (Q4_K_M quantization)
         โ”‚
         โ–ผ
  Pentest AI 3B โšก

Training stack:

  • ๐Ÿฆฅ Unsloth โ€” 2x faster fine-tuning
  • ๐Ÿค— TRL SFTTrainer โ€” supervised fine-tuning
  • LoRA rank 16 โ€” parameter efficient training
  • Q4_K_M quantization โ€” best quality/size tradeoff

๐Ÿ“Š Model Card Info

Property Value
Architecture Qwen2.5 (transformer)
Parameters 3B total
Context Length 32,768 tokens (trained)
Quantization Q4_K_M GGUF
File Size ~2 GB
Language English
Domain Cybersecurity / Security Research

๐Ÿ“ Prompt Format (ChatML)

<|im_start|>system
You are an expert penetration tester...<|im_end|>
<|im_start|>user
YOUR QUESTION HERE<|im_end|>
<|im_start|>assistant

โš ๏ธ Intended Use

This model is intended for:

  • โœ… Authorized penetration testing
  • โœ… CTF (Capture The Flag) competitions
  • โœ… Security research and education
  • โœ… Red team exercises on systems you own or have permission to test
  • โœ… Malware analysis and reverse engineering

Built with ๐Ÿ–ค for the security research community

If this model helped you in a CTF or pentest, drop a โญ

Downloads last month
3,017
GGUF
Model size
3B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for r1r21nb/qwen2.5-3b-instruct.Q4_K_M.gguf

Base model

Qwen/Qwen2.5-3B
Quantized
(12)
this model