Request for Kernel Publishing Access โ€” Showcasing Manusagents 18M+ Codex & Dataset Collections

#18
by Manusagents - opened

Hi everyone and the Review Team,

I have submitted a request for Kernel Publishing Access and would like to showcase my work to help speed up the review process.

๐Ÿš€ About My Work & Core Contributions

I specialize in building, curating, and fine-tuning large-scale open-source datasets and agentic pipelines. My primary project focus includes:

  • The Open Distillation Codex (Manusagents 18M+ Dataset):

    • A massive 76GB+ / 18.5M+ row multi-domain distillation dataset.
    • Covers 8 curated categories including Agentic Coding, Cybersecurity (Red/Blue team & Exploit analysis), Mathematics, Natural Sciences, and Multilingual Reasoning.
    • Contains over 7,090 repository-scale code archives for deep context LLM training.
  • 31+ Curated Dataset Collections:

    • High-depth reasoning datasets (Claude Opus / DeepSeek-R1 distilled chains).
    • Specialist datasets for multilingual instruction-tuning, formal mathematical proofs, and safety/security assessment frameworks.

๐Ÿ’ก Objective & Community Value

With Kernel Publishing access granted, I plan to:

  1. Share optimized dataset streaming pipelines and data-processing notebooks.
  2. Publish benchmark implementations, fine-tuning scripts, and model evaluation workflows.
  3. Provide open-source tools to help the community easily train and evaluate next-gen LLMs.

I would greatly appreciate it if the team could review my profile and grant publishing access. Looking forward to contributing to the community!

Thank you!
Manusagents / JD

Sign up or log in to comment