โšก Overview

We tested three frontier AI models on a brutal coding challenge: generate a complete, movie-quality Three.js solar system simulator in a single HTML file โ€” all procedural textures, custom shaders, bloom post-processing, and interactive controls, zero external assets.

๐ŸŽฏ The Challenge
Same prompt, three models, head-to-head. The prompt specified 15+ celestial bodies with exact shader requirements (multi-layer FBM noise, vortex noise, cellular noise), 3 glow layers on the sun, Jupiter's Great Red Spot with spiral edges, Saturn's Cassini division, Earth's city lights on the dark side, and more. No shortcuts allowed.

๐Ÿค– Models Tested

ModelToolTask 1: Solar SystemTask 2: City Renderer
GPT-5.6 SolCodex11 min11 + 33 min (2 rounds)
Kimi K3Kimi Code45 min3+ hours, round 2 abandoned
Qwen3.8MaxDirectโ€”<10 min (2 rounds)
โš ๏ธ Methodology Note
This is not a formal benchmark โ€” it's a real-world stress test by GitHub user OBdangshang07. Same prompt, same expectations, measured wall-clock time from tool invocation to complete output. Quality judged by visual fidelity and code completeness.

๐ŸŒŒ Task 1: Solar System Simulator

The prompt demanded a cinematic-grade solar system with 15+ celestial bodies, each requiring custom ShaderMaterial with multi-layer noise. Key requirements:

  • Sun: Dynamic plasma texture with 3 glow layers, sunspots, solar wind particles
  • Jupiter: Multi-layer stretched noise bands, Great Red Spot with spiral edges, 3+ smaller vortices
  • Saturn: Multi-layer translucent rings with Cassini division, particle scatter via shader discard
  • Earth: Ocean specular highlights, city lights on dark side, Fresnel atmosphere shell
  • Moons: 4 Galilean moons + 4 Saturnian moons, each with unique procedural textures

Results

AspectGPT-5.6 Sol (11 min)Kimi K3 (45 min)
Visual fidelityโญโญโญโญโญ Excellent shaders, Jupiter bands look organicโญโญโญโญ Good but some bodies use simpler noise
Code completenessAll 15+ bodies, all moons includedAll bodies present, some moons simplified
InteractivitySmooth orbit controls, info panels, time warpSimilar controls, slightly less polished UI
Post-processingUnrealBloomPass well-tunedBloom present but less refined
StabilityStable at 60fpsOccasional frame drops on lower GPUs

๐Ÿ™๏ธ Task 2: City Renderer

A harder challenge: generate a real-time metropolitan renderer with 50+ adjustable parameters, procedural infinite city, dynamic day/night cycle, atmospheric scattering, and dense 3D building clusters with unique textures.

Results

ModelRound 1Round 2 (Refinement)Building DensityTexture Quality
Qwen3.8Max<10 min<10 minโญโญโญโญโญ Dense, variedโญโญโญโญ Good procedural
GPT-5.6 Sol11 min33 minโญโญโญโญ Denseโญโญโญโญโญ Excellent detail
Kimi K33+ hoursโŒ Abandonedโญโญโญ Moderateโญโญโญ Decent
๐Ÿšจ Kimi K3 Struggled
Kimi K3 took over 3 hours for the first round and couldn't complete the second. The city renderer task exposed K3's weakness in generating complex, performance-sensitive 3D code with many interdependent systems (roads, traffic, lighting, building generation).

๐Ÿ“Š Key Findings

๐ŸŽ๏ธ Speed

Qwen3.8Max is the speed king โ€” both rounds under 10 minutes total. GPT-5.6 Sol is consistent (~11 min per task). Kimi K3 is 3-4x slower on complex tasks.

๐ŸŽจ Quality

GPT-5.6 Sol produces the best visual quality โ€” its shaders are more refined, noise layers more organic, and post-processing better tuned. Qwen3.8Max trades some visual polish for raw speed. Kimi K3's output is functional but less cinematic.

๐Ÿ”ง Reliability

GPT-5.6 Sol is the most reliable โ€” it completed both tasks fully. Qwen3.8Max is fast but may need a refinement round. Kimi K3 can timeout on complex multi-system tasks.

๐ŸŽฎ Live Demos โ€” Try Them Yourself

Below are the actual outputs from each model. Open them in your browser and compare the visual quality, interactivity, and performance:

โ˜€๏ธ Solar System Simulator

ModelGeneration TimeDemo
GPT-5.6 Sol11 min๐Ÿš€ Launch Demo โ†’
Kimi K345 min๐Ÿš€ Launch Demo โ†’

๐ŸŒƒ City Renderer

ModelGeneration TimeDemo
GPT-5.6 Sol11 + 33 min๐Ÿš€ Launch Demo โ†’
Kimi K33+ hours๐Ÿš€ Launch Demo โ†’
Qwen3.8Max<10 min ร— 2๐Ÿš€ Launch Demo โ†’

๐ŸŽฏ What This Means for Developers

๐Ÿ’ก Model Selection Guide
  • Need the best visual quality? โ†’ GPT-5.6 Sol (Codex). Best shaders, most polished output.
  • Need speed? โ†’ Qwen3.8Max. Blazing fast, good enough quality for prototyping.
  • Complex multi-system 3D? โ†’ Avoid Kimi K3 for now. It excels at simpler tasks but struggles with interconnected systems.
  • Cost-sensitive? โ†’ Qwen3.8Max delivers the best time-to-quality ratio.

The gap between models is not just speed โ€” it's about the complexity ceiling. GPT-5.6 Sol can handle prompts with 15+ interdependent shader systems. Qwen3.8Max is fast but may need refinement passes. Kimi K3 has a lower ceiling on complex 3D tasks.

For single-file HTML generation โ€” the emerging "vibe coding" benchmark โ€” GPT-5.6 Sol currently leads in quality, while Qwen3.8Max leads in throughput. Choose based on your priority.

๐Ÿ“Ž Source

All code, prompts, and test results are open source: github.com/OBdangshang07/AI_project

Test conducted by GitHub user OBdangshang07. Demos reproduced with original HTML output files.