bencodez commited on
Commit
6236247
·
verified ·
1 Parent(s): 3e1bb3e

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +72 -33
README.md CHANGED
@@ -1,12 +1,21 @@
1
  ---
2
  license: apache-2.0
3
- base_model: Qwen/Qwen2.5-Coder-0.5B-Instruct
 
 
 
 
 
 
 
 
 
 
 
4
  tags:
5
  - code
6
  - security
7
  - secure-coding
8
- - lora
9
- - qwen2.5-coder
10
  language:
11
  - en
12
  pipeline_tag: text-generation
@@ -14,34 +23,37 @@ pipeline_tag: text-generation
14
 
15
  # Cipheron
16
 
17
- **Cipheron** is a small, LoRA fine-tuned coding model specialized in **secure code review** — given a piece of code, it tries to spot common security vulnerabilities and suggest a fixed, secure version.
 
 
18
 
19
- - **Base model**: [Qwen/Qwen2.5-Coder-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B-Instruct) (Apache 2.0)
20
- - **Method**: LoRA fine-tuning (r=16, alpha=32), 3 epochs, ~830 steps
21
- - **Training data**: [CyberNative/Code_Vulnerability_Security_DPO](https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO) (~4.6k vulnerable/secure code pairs across 11 languages), trained on the secure ("chosen") responses only
22
- - **Size**: 0.5B parameters
23
- - **Formats**: full-precision merged model (this repo) and a `Cipheron-Q8_0.gguf` quantized file for on-device / CPU / phone use via llama.cpp, Ollama, or similar runners
24
 
25
- ## What it's good at
26
 
27
- In testing, Cipheron reliably identifies and correctly fixes:
28
- - **SQL injection** (rewrites string-concatenated queries as parameterized queries)
29
- - **Command injection** (rewrites `os.system`/shell string concatenation as safer `subprocess` calls)
30
 
31
- These categories are well-represented in the training data.
32
 
33
- ## Known limitations
 
 
 
 
34
 
35
- The training dataset is heavily imbalanced (e.g. ~30% buffer-overflow examples, mostly in memory-unsafe languages like C/C++, largely irrelevant to Python; some important categories like path traversal, hardcoded secrets, and weak cryptography have only a handful of examples total). As a result, in testing Cipheron **failed to correctly fix**:
36
  - Path traversal
37
- - Hardcoded secrets / API keys
38
- - Weak hashing (e.g. MD5 for passwords)
39
- - Insecure deserialization (`pickle.loads` on untrusted input)
40
  - Reflected XSS
 
 
 
41
 
42
- For these categories it tends to produce superficial, security-irrelevant changes (e.g. wrapping code in try/except, adding default arguments) rather than the actual fix. **Do not rely on this model as a substitute for a real security review or a larger model.** It's best used as a lightweight, offline first-pass check for the vulnerability classes listed above under "What it's good at," not as a general-purpose security auditor.
43
 
44
- This is a small (0.5B parameter) educational/experimental model, not a production security tool.
45
 
46
  ## Usage
47
 
@@ -50,19 +62,46 @@ from transformers import AutoModelForCausalLM, AutoTokenizer
50
  import torch
51
 
52
  tokenizer = AutoTokenizer.from_pretrained("bencodez/Cipheron")
53
- model = AutoModelForCausalLM.from_pretrained("bencodez/Cipheron", torch_dtype=torch.bfloat16)
 
 
 
 
54
 
55
  messages = [
56
- {"role": "system", "content": "You are a secure coding assistant. Review code for security vulnerabilities and provide fixed, secure versions."},
57
- {"role": "user", "content": "Review this code for security issues and fix it:\n\ndef get_user(username):\n query = \"SELECT * FROM users WHERE username = '\" + username + \"'\"\n return db.execute(query)"},
 
 
 
 
 
 
 
 
 
 
 
 
 
 
58
  ]
59
- input_ids = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=False)
60
- out = model.generate(input_ids, max_new_tokens=250)
61
- print(tokenizer.decode(out[0][input_ids.shape[1]:], skip_special_tokens=True))
62
- ```
63
-
64
- Or with the GGUF file via `llama-cpp-python` / llama.cpp / Ollama for lightweight CPU/on-device inference.
65
-
66
- ## License
67
 
68
- Apache 2.0, inherited from the base model (Qwen2.5-Coder-0.5B-Instruct).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ tags:
4
+ - code
5
+ - security
6
+ - secure-coding
7
+ language:
8
+ - en
9
+ pipeline_tag: text-generation
10
+ ---
11
+ from pathlib import Path
12
+
13
+ readme = """---
14
+ license: apache-2.0
15
  tags:
16
  - code
17
  - security
18
  - secure-coding
 
 
19
  language:
20
  - en
21
  pipeline_tag: text-generation
 
23
 
24
  # Cipheron
25
 
26
+ **Cipheron** is a lightweight coding model designed for **secure code review**.
27
+
28
+ It analyzes source code for common security vulnerabilities and attempts to explain the issue and provide a safer implementation.
29
 
30
+ ## What Cipheron Is Good At
 
 
 
 
31
 
32
+ Cipheron performs particularly well on:
33
 
34
+ - **SQL injection** identifying unsafe query construction and recommending parameterized queries.
35
+ - **Command injection** identifying unsafe shell command construction and recommending safer subprocess-based approaches.
 
36
 
37
+ These vulnerability classes are strongly represented in its evaluation data.
38
 
39
+ ## Known Limitations
40
+
41
+ Cipheron has limited reliability across many security vulnerability categories.
42
+
43
+ In testing, it struggled with:
44
 
 
45
  - Path traversal
46
+ - Hardcoded secrets and API keys
47
+ - Weak password hashing
48
+ - Insecure deserialization
49
  - Reflected XSS
50
+ - Complex multi-step security vulnerabilities
51
+
52
+ For these cases, the model may produce changes that appear security-related but do not actually eliminate the underlying vulnerability.
53
 
54
+ **Do not rely on Cipheron as a replacement for professional security review, static analysis, penetration testing, or a larger security-focused model.**
55
 
56
+ Cipheron is best considered a lightweight, experimental tool for first-pass security analysis and secure-coding experimentation.
57
 
58
  ## Usage
59
 
 
62
  import torch
63
 
64
  tokenizer = AutoTokenizer.from_pretrained("bencodez/Cipheron")
65
+
66
+ model = AutoModelForCausalLM.from_pretrained(
67
+ "bencodez/Cipheron",
68
+ torch_dtype=torch.bfloat16
69
+ )
70
 
71
  messages = [
72
+ {
73
+ "role": "system",
74
+ "content": (
75
+ "You are a secure coding assistant. "
76
+ "Review code for security vulnerabilities "
77
+ "and provide fixed, secure versions."
78
+ )
79
+ },
80
+ {
81
+ "role": "user",
82
+ "content": """Review this code for security issues and fix it:
83
+
84
+ def get_user(username):
85
+ query = "SELECT * FROM users WHERE username = '" + username + "'"
86
+ return db.execute(query)"""
87
+ },
88
  ]
 
 
 
 
 
 
 
 
89
 
90
+ input_ids = tokenizer.apply_chat_template(
91
+ messages,
92
+ add_generation_prompt=True,
93
+ return_tensors="pt",
94
+ return_dict=False
95
+ )
96
+
97
+ out = model.generate(
98
+ input_ids,
99
+ max_new_tokens=250
100
+ )
101
+
102
+ print(
103
+ tokenizer.decode(
104
+ out[0][input_ids.shape[1]:],
105
+ skip_special_tokens=True
106
+ )
107
+ )