arxiv:2609.03887
🤝 Open to Collab
Hoang Cuong Nguyen
HoangCuongNguyen
AI & ML interests
Natural Language Processing in Cybersecurity/ Safety Alignment for LLMs
Recent Activity
authored a paper 13 days ago
Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness published a dataset 13 days ago
HoangCuongNguyen/refusal-direction-results updated a collection 22 days ago
EMNLP 2026 Post-Training AnalysisOrganizations
None yet