Editorial cover for ai story

Ant Group’s Ling-3.0-flash-VL Scores 25 on Artificial Analysis, Lands #2 in Size Class

Ant Group Ling-3.0-flash-VL Artificial Analysis is the focus of this story. On September 12, 2026, Ant Group’s inclusionAI lab published an independent intelligence reading for its new natively multimodal model, Ling-3.0-flash-VL. The model registered a 25 on the Artificial Analysis Intelligence Index v4.3, placing it at #2 of 64 entries in its size class against a class median of 8. The score was independently run and released alongside a model page that lists speed, latency, and price.

Ant Group Ling-3.0-flash-VL Artificial Analysis benchmark numbers

The independent 25 is the headline figure. It is a real number on the current v4.3 ruler. Ant Group’s own model card claims 42 on Artificial Analysis Intelligence Index protocol v4.1.1. That vendor number is unreproduced. The two protocols do not measure the same thing: v4.3 incorporates AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, and Humanity’s Last Exam. Two things are true at once: the 42 is unreproduced, and it is not evidence of anything dishonest, because v4.1.1 and v4.3 do not measure the same thing.

Architecture details place the model in a competitive spot on the intelligence-versus-active-parameter curve. Ling-3.0-flash-VL is a sparse mixture-of-experts with 124B total parameters and roughly 5.5B active per token, released under the MIT license with BF16 and FP8 weights, accepting text, images and video and returning text. The class median on this index is 8; Ling-3.0-flash-VL scores 25; and it does so while activating 5.5B parameters per token. The model’s served identifier, Ling-3.0-flash-VL-rc1, carries a release-candidate marker, and it remains unclear whether the served checkpoint is byte-identical to the September 4 open weights.

Speed, context, and verbosity

Throughput and latency numbers are solid. Output speed reached 142.9 tokens per second, ranking #10 of 64 against a peer median of 105.1. Time to first token was 2.21 seconds, just below the peer median of 2.29 seconds. The context window measured 262K tokens by Artificial Analysis, versus 256K stated on the model card. The minor gap is consistent with measurement differences rather than a marketing issue.

Verbosity is the gotcha. Artificial Analysis records Ling-3.0-flash-VL spending 160M output tokens to complete the Intelligence Index, ranked #8 of 64 against a median of 84M, and labels it “very verbose.” A model that activates only 5.5B parameters per token enjoys a cheap per-token cost, but doubling the median output token count gives back some of that advantage on a per-task basis. Buyers running batch jobs should price on tokens produced, not tokens consumed.

Pricing, serving, and the OrcaRouter gap

The Artificial Analysis page lists Ling-3.0-flash-VL at $0.00 per 1M input and $0.00 per 1M output tokens. That is not a price. It is the appearance of one. The $0.00 figure reflects a free trial via Ling Studio, not a published tariff. The text-only sibling Ling-3.0-flash has been listed around $0.075 input and $0.22 output per 1M tokens, which is closer to a workable reference point for anyone budgeting inference spend.

Serving is available through several routes. Ant runs an OpenAI-compatible endpoint at api.ant-ling.com/v1/chat/completions, an Anthropic-compatible endpoint, and a free trial on Ling Studio. The model is self-hostable. Ant has published an SGLang cookbook and a Ling-tuned fork of vLLM. Weights hit disk on September 4; the FP8 quantization arrived September 8 at roughly 126 GB; the independent Artificial Analysis score was published September 10.

One commercial note: OrcaRouter does not carry this model on its catalogue. Buyers routing through aggregator providers will need to point directly at Ant’s endpoints or self-host. The vendor’s own demonstration scored 1,441 on the Image-to-WebDev Arena under the handle “linthium,” narrowly edging GPT-5.4’s 1,440, though that is a vendor-run benchmark on a vendor-chosen prompt set.

What to watch

The MIT-licensed weights, the FP8 release, and the 5.5B-active-per-token profile make Ling-3.0-flash-VL a useful option for teams that want a multimodal model they can self-host without paying frontier-tier inference rates. The verbosity profile and the release-candidate served-checkpoint ambiguity are the two items to monitor before committing production traffic. The 25 score on v4.3 is the number to plan around, and the protocol-version split between Ant’s claim and the independent run is the caveat that needs to travel with it. Ant Group Ling-3.0-flash-VL Artificial Analysis reporting will continue to evolve as more checkpoints and pricing tiers land.

Source: https://www.orcarouter.ai/blog/ling-3-0-flash-vl-intelligence-index

The 25 score lands in a crowded neighborhood. On the v4.3 ruler, the bracket just above Ling-3.0-flash-VL is occupied by a handful of dense 30B-to-70B models from western labs, and the bracket just below it is dense with smaller mixture-of-experts releases that activate between 7B and 12B per token. What makes Ling-3.0-flash-VL structurally interesting is the active-parameter efficiency: a 25 on the index from 5.5B active parameters per token is a better ratio than several denser peers that score in the same band. For self-hosters sizing GPUs, that ratio matters more than the headline score, because H100 and H200 capacity budgets are denominated in active compute, not total weights.

The protocol-version discrepancy deserves a closer look, because it will recur. Ant’s model card cites v4.1.1, which pre-dates the introduction of GDPval-AA v2, AutomationBench-AA, and the revised Humanity’s Last Exam weighting that arrived in v4.2 and v4.3. A 42 on v4.1.1 is not directly comparable to a 25 on v4.3, but the gap is also not purely an artifact of stricter tests; the index family was rebalanced between revisions, and several previously-leading models saw their v4.1.1 scores drop by 30 percent or more when re-evaluated under v4.3. Buyers who keep internal leaderboards should pin a protocol version to every entry and avoid mixing v4.1.1 numbers with v4.3 numbers in the same chart.

The 160M output tokens used to complete the Intelligence Index is worth contextualizing. The Artificial Analysis suite runs a fixed battery of agentic and reasoning tasks, and the output-token count reflects how much the model chose to write. A “very verbose” label in this dataset correlates weakly with raw quality and strongly with response style; some top-quartile models on the index spend fewer than 50M output tokens across the same battery. For production deployments, this matters because API bills, latency tail behavior, and downstream parsing costs all scale with output length. Teams building agent loops on top of Ling-3.0-flash-VL should budget for roughly twice the output-token volume they would see from a comparable dense model in the same intelligence bracket, and they should consider prompt-level interventions to constrain verbosity where the use case allows. That is the practical shape of the Ant Group Ling-3.0-flash-VL Artificial Analysis release as the data currently stands.

Source: https://example.com

Leave a Comment

Your email address will not be published. Required fields are marked *