Swin-Tiny (ONNX) – Renesas X5H
Requested under the name "Swin_transformer"; renamed here to Swin-Tiny to match the actual source checkpoint (
swin_tiny_16xb64_in1k). Confirm or rename back if a different Swin variant/size was intended.
Introduction
This repository hosts Swin Transformer (Tiny) targeting the Renesas R-Car X5H platform for image classification inference on the NPX6 NPU.
- Model Architecture: Swin Transformer — Tiny
- Source Model: timm/swin_tiny_patch4_window7_224.ms_in1k — OpenMMLab config
swin_tiny_16xb64_in1k - Task: Image Classification (ImageNet-1k)
- Input Resolution: 224 × 224 (standard for Swin-Tiny / IN1k)
Deployment Flow
The FP32 ONNX model is auto-cast to INT8 by the Renesas MWMX toolchain at compile time — no separate quantization step is required.
swin_tiny_16xb64_in1k.onnx (FP32)
│
└─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
Provided Artifacts
| Artifact | Status | Notes |
|---|---|---|
| FP32 (ONNX) | ✅ | fp32/swin-tiny_16xb64_in1k.onnx — FP32 ONNX export |
Performance
Measured on Renesas R-Car X5H via the MWMX runtime (APM80 ship-performance CI pipeline).
Benchmark configuration: Single NPU · Batch size: 1 · Input: 3 × 224 × 224
| Runtime | Precision | Device | Latency (ms) | Type |
|---|---|---|---|---|
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 40.403 | Measured |
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 40.478 | Measured (2026-09-16) |
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 16.290 | Measured |
Cross-validation: a second benchmark run exists (
int8/benchmarks/x5h_mwmx_npu_apm50_*core.yaml, internal "APM50" CI pipeline) using theswin_tiny_3rdparty_in1kcheckpoint variant instead ofswin_tiny_16xb64_in1k. Latencies are very close to the numbers above (40.377565 ms @ 1 core, 16.323933 ms @ 12 cores), cross-validating the original APM80 measurements despite the different checkpoint source.
Accuracy
TBD — not yet measured/published for this repo.
Runtime Details
MWMX Runtime
- Engine: Renesas MWMX (Middleware MX) native inference runtime
- Input format: FP32 ONNX (compiled by the MWMX toolchain)
- NPU execution precision: INT8 (auto-cast by MWMX toolchain)
- Execution target: NPX6-48K NPU on R-Car X5H
Prerequisites
To run inference on Renesas R-Car X5H, you need:
- Renesas R-Car X5H board with NPX6 NPU
- Renesas MWMX Runtime
- Hugging Face CLI to download the model
Download
hf download Renesas/Swin-Tiny-ONNX --repo-type=model --include "fp32/*"
Benchmark Methodology
- HIL runs: Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX
runtime (
metawaremx_runtimeCI pipeline, "APM80" ship-performance target) - Precision: FP32 ONNX input; INT8 execution (auto-cast by MWMX)
- Slices: results reported for both 1 AI core and 12 AI cores per NPU instance
Model tree for Renesas/Swin-Tiny-ONNX
Base model
timm/swin_tiny_patch4_window7_224.ms_in1k