Spaces:
Running
Running
Request for Kernel Publishing Access โ Showcasing Manusagents 18M+ Codex & Dataset Collections
#18
by Manusagents - opened
Hi everyone and the Review Team,
I have submitted a request for Kernel Publishing Access and would like to showcase my work to help speed up the review process.
๐ About My Work & Core Contributions
I specialize in building, curating, and fine-tuning large-scale open-source datasets and agentic pipelines. My primary project focus includes:
The Open Distillation Codex (
Manusagents18M+ Dataset):- A massive 76GB+ / 18.5M+ row multi-domain distillation dataset.
- Covers 8 curated categories including Agentic Coding, Cybersecurity (Red/Blue team & Exploit analysis), Mathematics, Natural Sciences, and Multilingual Reasoning.
- Contains over 7,090 repository-scale code archives for deep context LLM training.
31+ Curated Dataset Collections:
- High-depth reasoning datasets (Claude Opus / DeepSeek-R1 distilled chains).
- Specialist datasets for multilingual instruction-tuning, formal mathematical proofs, and safety/security assessment frameworks.
๐ก Objective & Community Value
With Kernel Publishing access granted, I plan to:
- Share optimized dataset streaming pipelines and data-processing notebooks.
- Publish benchmark implementations, fine-tuning scripts, and model evaluation workflows.
- Provide open-source tools to help the community easily train and evaluate next-gen LLMs.
I would greatly appreciate it if the team could review my profile and grant publishing access. Looking forward to contributing to the community!
Thank you!
Manusagents / JD