Bench run

ElevenLabs vs Mureka: 19 Music Generation Tests

AgentsBench / Published Jul 31, 2026 / Updated Jul 31, 2026

ElevenLabs won 14 of 19 paired music-generation tests, mainly because its songs were judged more complete. Mureka won five, including both voice-cloning cases. This is a workflow result, not a universal model ranking.

ElevenLabs vs Mureka benchmark result

The win count looks decisive, but the average scores were close. ElevenLabs averaged 85.7, while Mureka averaged 83.8 on the evaluator’s reported 100-point scale.

ResultElevenLabsMureka
Test wins145
Win rate73.7%26.3%
Average total score85.783.8
Average instruction score8.478.74
Average style score8.688.68
Average structure score8.637.74
Average vocal and finish score8.478.21

The clearest separation was structure, not prompt understanding. Mureka had the higher instruction-following average, while ElevenLabs led on structure and finished-song quality.

What the result means

ElevenLabs was the stronger default for producing a complete song from a brief in this run. It swept all four pure text-description cases and won every sampled emotion, structure, and bilingual edge case.

Mureka was more competitive when the task demanded exact lyrics, a specific arrangement, or a reference voice. Its five wins included the two largest margins anywhere in the benchmark.

The practical split is therefore narrower than 14 wins to five implies: ElevenLabs for consistent full-song delivery, and Mureka for selected controlled or reference-led workflows.

How AgentsBench tested the two products

The source artifact contained 19 paired MP3 results. Each pair represented one task completed once by Mureka and once by ElevenLabs.

The two products received platform-specific prompts. Mureka prompts often specified exact lyrics, instruments, tempo, and voice, while ElevenLabs prompts described the target song in natural language.

Both audio files and their corresponding prompts were sent to doubao-seed-2-0-lite-260428. Seed 2.0 Lite acted as the listening evaluator; it did not generate the songs.

The evaluator considered four dimensions: instruction response, style match, structural completeness, and vocal or finished-product quality. It also selected one winner for each pair.

This setup measures whether each product completed its assigned version of a task. It does not isolate model quality under identical inputs.

Results by test category

Test categoryCasesElevenLabs winsMureka winsDirectional finding
Pure text description440ElevenLabs delivered more complete songs
Use-case and scene fit321ElevenLabs led; Mureka won corporate BGM
Style diversity321Mureka won the strict Chinese-style case
Multilingual431Mureka won French hip-hop
Emotion and structure220ElevenLabs led on full-song development
Chinese-English edge case110Both worked; ElevenLabs was more complete
Voice cloning202Mureka won both English and Chinese cases

The category counts are small. They show where this run separated the products, not a statistically reliable capability ranking.

Why ElevenLabs won more tests

The evaluator repeatedly described ElevenLabs outputs as complete songs with developed sections, coherent endings, and sustained emotional progression.

Several losing Mureka outputs were judged as strong fragments. They followed lyrics or style constraints but did not develop the full structure expected from the task.

That pattern appears in the dimension averages. The products tied at 8.68 for style, and Mureka led instruction response by 0.27 points. ElevenLabs led structure by 0.89 points.

The pure-description group made the difference especially visible. ElevenLabs won all four cases, including a ballad, upbeat summer pop, contemporary pop, and a minimal-lyrics dance song.

Where Mureka was stronger

Mureka’s two voice-cloning wins were the largest gaps in the test. It led the English cloning case by 35 points and the Chinese cloning case by 30 points.

In both cases, the evaluator said Mureka produced a complete sung performance. The paired ElevenLabs outputs were short and failed to complete the requested song-style performance.

Mureka also won three tasks outside cloning: corporate background music by 18 points, a Chinese-style song by 17 points, and French hip-hop by five points.

These wins align with its slightly higher instruction average. When Mureka completed the structure, it could match exact lyrics, instrumentation, tempo, or arrangement constraints closely.

All 19 scored cases

CaseTest areaElevenLabsMurekaWinner
TC-PTS-01Melancholic English pop92.585.5ElevenLabs
TC-PTS-02Upbeat summer pop92.581.5ElevenLabs
TC-PTS-03Full contemporary pop structure90.080.0ElevenLabs
TC-PTS-04Minimal-lyrics dance pop98.083.0ElevenLabs
TC-BGM-01Yoga and meditation BGM90.072.0ElevenLabs
TC-BGM-02Horror and suspense BGM90.079.5ElevenLabs
TC-BGM-03Corporate promotional BGM72.090.0Mureka
TC-LTS-01English pop style90.083.0ElevenLabs
TC-LTS-02Chinese pop ballad92.588.5ElevenLabs
TC-LTS-03Chinese-style song73.090.0Mureka
TC-LTS-04Cantonese pop89.580.5ElevenLabs
TC-LTS-05Japanese J-pop95.086.0ElevenLabs
TC-LTS-06Korean K-pop90.070.0ElevenLabs
TC-EDG-04Chinese-English pop95.589.5ElevenLabs
TC-MLT-01Spanish Latin pop90.079.0ElevenLabs
TC-MLT-02Arabic pop90.086.0ElevenLabs
TC-MLT-03French hip-hop85.090.0Mureka
TC-VCL-R01English voice cloning53.088.0Mureka
TC-VCL-R02Chinese voice cloning60.090.0Mureka

Which product should you choose?

Choose ElevenLabs first when the main requirement is a complete, polished song from a descriptive brief. That is the use case most directly supported by its 14 wins in this run.

Choose Mureka for a trial when exact customization or voice-reference work matters more than default structural consistency. Its strongest sampled evidence came from cloning and tightly constrained tasks.

Do not treat either recommendation as permanent. The artifact did not record the generator model versions, and both services can change independently of this July 2026 run.

Current official documentation gives additional product context. Eleven Music documents full songs, instrumentals, sectional editing, multilingual vocals, and audio references.

The Mureka API platform documents song and instrumental generation, media-conditioned soundtracks, vocal cloning, music analysis, and stem export.

Those current feature pages are entity sources, not evidence that every listed capability was exercised in this benchmark.

What Seed 2.0 Lite did in this evaluation

Doubao Seed 2.0 Lite was used as an audio-capable judge. Volcengine’s release notes identify doubao-seed-2-0-lite-260428 as supporting reasoning over audio inputs.

That makes direct MP3 comparison possible, but it does not make the scores objective. The evaluator can have preferences about song length, structure, genre conventions, language, and vocal quality.

The results should be read as one consistent model judge applied across the dataset. A human listening panel is still needed before making a high-stakes vendor or production decision.

See the AgentsBench methodology for how tested claims are separated from tracked product facts.

Benchmark limitations

The largest limitation is prompt asymmetry. The prompts targeted comparable goals, but they were not identical, so the test also captures differences in product interfaces and prompting strategy.

The generator model versions, seeds, subscription tiers, and generation settings were not recorded in the source artifact. The result cannot be reproduced as a precise model-to-model leaderboard.

Each task appears once. There were no repeated generations to measure variance, failure rate, or how often a product could recover after a retry.

The judgment came from one AI evaluator and no blind human panel. The supplied HTML referenced local MP3 files, but those files were not present, so this page does not publish playable evidence.

The benchmark did not measure price, latency, licensing terms, safety behavior, editing ergonomics, or downstream commercial performance.

Source record

The test artifact is the source of every score and win count on this page. The external product pages only establish current entity and capability context.

Frequently asked questions

Which was better in this test, ElevenLabs or Mureka?

ElevenLabs won 14 of 19 paired tests, while Mureka won five. ElevenLabs was favored for complete song structure. Mureka won both voice-cloning tests and several tasks with strict custom constraints.

What were the average ElevenLabs and Mureka scores?

The Seed 2.0 Lite evaluator gave ElevenLabs an average of 85.7 and Mureka an average of 83.8. The 1.9-point difference was much narrower than the 14-to-5 win count suggests.

Why did ElevenLabs win more tests?

The evaluator repeatedly preferred ElevenLabs outputs for complete sections, endings, emotional progression, and finished-song quality. Its average structure score was 8.63, compared with 7.74 for Mureka.

What was Mureka better at?

Mureka won both sampled voice-cloning tests by 35 and 30 points. It also won the corporate background-music, Chinese-style song, and French hip-hop cases. Its average instruction score was slightly higher.

Was Seed 2.0 Lite one of the music generators?

No. Doubao Seed 2.0 Lite was the audio-capable evaluator. The generated MP3 files came from ElevenLabs and Mureka, and Seed 2.0 Lite compared each pair against the supplied prompts.

Did the benchmark use identical prompts?

No. Each product received a platform-specific prompt for the same task, and each output was judged against its own prompt. This is a workflow comparison, not a controlled same-prompt model benchmark.

Did ElevenLabs beat Mureka on multilingual music?

ElevenLabs won three of four cases in the multilingual category and also won the bilingual Chinese-English case. Mureka won the French hip-hop case. The sample is too small for a general language ranking.

What are the main limitations of this ElevenLabs vs Mureka test?

The test used one AI judge, one output per task, platform-specific prompts, and no human listening panel. Generator versions and settings were not recorded, and the supplied artifact did not include publishable audio files.