Carbon-A-1.2B

A DNA annotation model from the Carbon family.

Carbon-A predicts protein-coding sequence (CDS) at single-base resolution in eukaryotic genomes. It produces separate forward- and reverse-strand probabilities, preserving coding regions that overlap on opposite strands. The annotation pipeline in this repository turns these probabilities into gene models with a confidence score per gene.

Facts

  • 1.2B-parameter DNA annotation model with two strand-specific classification heads.
  • Tokenizer: non-overlapping 6-mers; each DNA token represents six bases.
  • Sequence length: 16,384 tokens (98,304 bp). The annotation pipeline reads longer records in windows that overlap by 1,024 tokens (6,144 bp).
  • Inputs: genomic DNA. The annotation pipeline reads FASTA and GenBank files, optionally gzip-compressed.
  • Outputs: per-base probabilities for non-coding and CDS on each strand. The annotation pipeline turns them into gene models with a confidence score per gene.
  • Training data: 2,055 RefSeq GCF assemblies (2,047 species) from mammals, other vertebrates, invertebrates, plants, fungi, and protozoa.
  • Transformers interface: AutoModelForTokenClassification with trust_remote_code=True; Transformers 4.56 or later, including 5.x. FlashAttention-2 needs Transformers 4.x.
  • Inference: BF16 with FlashAttention-2, the setting of all reported results. PyTorch SDPA (--attn-implementation sdpa) needs no FlashAttention installation, but its probabilities are not bit-identical.

How to use

Install

The annotation pipeline runs on Linux. It needs Python 3.10 or later, a CUDA build of PyTorch, and a C++ compiler with OpenMP. If you do not have a suitable compiler, we recommend GCC, for example on Debian or Ubuntu:

sudo apt install build-essential

Then download this repository and install its requirements:

pip install -U huggingface_hub
hf download HuggingFaceBio/Carbon-A-1.2B --exclude model.safetensors --local-dir Carbon-A-1.2B
pip install -r Carbon-A-1.2B/requirements.txt
pip install --no-build-isolation flash-attn

The weights (4.6 GB) are downloaded to the Hugging Face cache on first use.

Annotate a genome

bash Carbon-A-1.2B/annotation_pipeline/run_annotation.sh --input genome.fna.gz

The input is a FASTA or GenBank file, optionally gzip-compressed. Carbon-A reads it in 98,304-bp windows that overlap by 6,144 bp and averages the overlapping predictions. The pipeline then decodes the gene structures and writes the results to results/<accession>/:

  • prediction.gff: genes, mRNAs, and CDS. Every gene carries confidence, the probability that its structure is exactly right.
  • prediction.fna and prediction.faa: the CDS and protein sequences.
  • *.raw_probs.parquet: the per-base CDS probabilities.

All genes are written. The reported results count the genes with a confidence of at least 0.1.

The standard genetic code is used by default. For another code, pass --codon-table. Tetrahymena thermophila (GCF_000189635.1) uses translation table 6 automatically. A likely codon-table mismatch is detected automatically: the pipeline adds a warning to results/WARNINGS.txt and renames the result directory to results/<accession>_WARNING.

The pipeline README lists all options. It also describes the evaluation against a reference annotation and BUSCO.

Test on a benchmark genome

The benchmark genomes and their NCBI annotations are in the data bucket under eval/. To test the pipeline on yeast, download its genome and annotation:

url=https://huggingface.co/buckets/HuggingFaceBio/Carbon-A-training-data/resolve/eval/seen/GCF_000146045.2
mkdir -p GCF_000146045.2
for file in GCF_000146045.2_R64_genomic.fna genomic.gff cds_from_genomic.fna sequence_report.jsonl; do
    curl -fsSL -o GCF_000146045.2/$file $url/$file
done

Annotate the genome and compare the result with the NCBI annotation:

bash Carbon-A-1.2B/annotation_pipeline/run_annotation.sh \
    --input GCF_000146045.2/GCF_000146045.2_R64_genomic.fna
bash Carbon-A-1.2B/annotation_pipeline/run_evaluation.sh \
    --prediction-prefix results/GCF_000146045.2/prediction \
    --genome-dir GCF_000146045.2

results/GCF_000146045.2/evaluation_metrics.json then reports a nucleotide F1 of 0.998, an exon F1 of 0.943, and a gene F1 of 0.953.

License

Apache 2.0.

Downloads last month
126
Safetensors
Model size
1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using HuggingFaceBio/Carbon-A-1.2B 1

Collection including HuggingFaceBio/Carbon-A-1.2B

Article mentioning HuggingFaceBio/Carbon-A-1.2B