AIarXiv
Heuristic editor, no API keyVerdict: NotableMultilingual GSM-Symbolic: What determines capability transfer across languages?
We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants…
Key numbers
- 92% of between-language variation
- 23% of the model-by-language variation
VerdictWorth a reader's time today.
Abstract
We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size (β= 1.77), language resource level (β= 0.77), reasoning (β= 0.67) and typological distance (β= -0.25). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages (β= -0.27 and β= -0.20, respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model's performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 24% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ███░░ 3 | 18% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ███░░ 3 | 14% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ██░░░ 2 | 16% | A new combination of known ideas. |
| Trajectory | ██░░░ 2 | 18% | Some room to improve with obvious engineering. |
| Stakes | ███░░ 3 | 10% | Meaningful benefit to many people within a few years. |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (relative gain); verification (multiple benchmarks, held-out test); stakes (global scale, general AI).
How the score was computed
- Merit
- 5.3 / 10
- Adjusted merit
- 4.6 / 10
- Attention
- 71%
- Freshness
- 38%
- Hugging Face upvotes39 (reference 25, via hf-daily, Oct 6, 2026, 01:15 UTC)
- GitHub stars4 (reference 250, via hf-daily, Oct 6, 2026, 01:15 UTC)