Running this model locally is fastest when deployed through a PowerShell script.
Refer to the instructions below to proceed.
An automated background process downloads all required large-scale files.
The installer will automatically analyze your hardware and select the optimal configuration.
The Qwen3-ASR-0.6B model is a compact speech recognition system designed for realātime transcription across multiple languages. It contains 0.6āÆbillion parameters, striking a balance between accuracy and onādevice deployment feasibility. The architecture leverages efficient attention mechanisms to achieve low inference latency, making it suitable for realātime applications. A dedicated languageāagnostic encoder enables robust performance on languages not commonly represented in largeāscale datasets. The modelās lightweight footprint is highlighted in the comparison table below, which outlines key metrics such as parameter count, word error rate, and inference time.
| Metric | Value |
|---|---|
| Parameters | 0.6āÆB |
| Word Error Rate | 6.2% |
| Inference Latency | 12āÆms |