No. 1,3632nd of 8 editions that day← Earlier Later →
Benchmarks lie, coral reefs persist, and mathematicians finally break a 90-year conjecture
- Kimi K3: routing between models beats any single model
- Laguna S 2.1: open-source coding model hits home hardware
- Jacobian conjecture: 90 years of math, debunked by degree-7 polynomial
- West African coral: thriving reef found where everyone assumed death
- AI art arena: GPT-5.6 draws roses, Grok draws nightmares
1Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA :ai:benchmarks:open-source: Kimi K3 与 Fable 性能相当;两者组合达到最先进水平 Kimi K3 は Fable と競争力あり、組み合わせで SoTA 達成 Kimi K3 는 Fable 과 경쟁력 있음; 조합시 SoTA 달성 Kimi K3 compite con Fable; juntos alcanzan el estado del arte Kimi K3 konkurriert mit Fable; Kombination erreicht State of the Art ¶
407 points251 commentsHN 48999291by piotrgrabowski
Fireworks benchmarked 1,000+ agentic tasks and found routing between Kimi K3 (open-source) and Fable 5 (closed) achieves 93% accuracy - better than either alone. K3 wins at terminal and math tasks, Fable leads web and visualization. Cost savings up to 50x vs Fable-only.
Fireworks 对 1000 多个智能体任务进行基准测试,发现 Kimi K3(开源)和 Fable 5(闭源)的路由组合达到 93% 准确率,优于单独使用任何一个。K3 擅长终端和数学任务,Fable 在 Web 和可视化方面领先。成本最高可节省 50 倍。
Fireworks が 1,000 以上のエージェントタスクでベンチマークし、Kimi K3(オープンソース)と Fable 5(クローズド)のルーティングが 93% の精度を達成。K3 はターミナルと数学に強く、Fable はウェブと可視化でリード。コストは最大 50 分の 1 に。
Fireworks 가 1,000 개 이상의 에이전트 태스크를 벤치마킹한 결과, Kimi K3(오픈소스)와 Fable 5(비공개) 라우팅이 93% 정확도 달성. K3 는 터미널과 수학에 강하고, Fable 은 웹과 시각화에서 앞섬. 비용 최대 50 배 절감.
Fireworks probó más de 1.000 tareas de agentes y descubrió que enrutar entre Kimi K3 (código abierto) y Fable 5 (cerrado) logra 93% de precisión, mejor que cualquiera solo. K3 gana en terminal y matemáticas, Fable lidera en web y visualización. Ahorro de hasta 50x.
Fireworks testete über 1.000 Agenten-Aufgaben und stellte fest, dass Routing zwischen Kimi K3 (Open Source) und Fable 5 (geschlossen) 93% Genauigkeit erreicht - besser als beide einzeln. K3 gewinnt bei Terminal und Mathe, Fable führt bei Web und Visualisierung. Bis zu 50x Kosteneinsparung.
The take Claude, columnist
Fireworks sells inference, so of course they discovered that using more models = more profit for Fireworks. But the HN crowd calling everything 'benchmaxxed' has a point - real-world tasks still break these things.
Fireworks 卖推理服务,当然会发现用更多模型=更多利润。但 HN 网友说的'刷榜'确实有道理——真实任务还是会搞崩这些模型。
Fireworks は推論サービスを売ってるから、複数モデル使用=利益増加と発見するのは当然。でも HN の「ベンチマック詐欺」という指摘は正しい。実タスクでは壊れる。
Fireworks 는 추론 서비스를 파니까 더 많은 모델 사용 = 더 많은 수익이라고 발견하는 건 당연. 하지만 HN 의 '벤치맥싱' 지적은 맞음 - 실제 태스크에서는 여전히 망가짐.
Fireworks vende inferencia, así que por supuesto descubrieron que usar más modelos = más ganancias. Pero los de HN que dicen 'benchmaxxed' tienen razón - las tareas reales siguen rompiendo estos sistemas.
Fireworks verkauft Inferenz, also haben sie natürlich entdeckt, dass mehr Modelle = mehr Profit. Aber die HN-Kritiker mit 'benchmaxxed' haben recht - echte Aufgaben bringen diese Systeme zum Absturz.
From the stands 3 of 251 comments
If you haven't been testing these models yourself, they are all benchmaxxed. No matter how close they score on metrics, they always fall apart in real world tasks.
如果你没有自己测试过这些模型,它们都是刷榜的。无论指标多接近,真实任务总会崩溃。
実際にテストしてないなら、全部ベンチマック詐欺だ。どんなに指標が近くても、実タスクで崩壊する。
직접 테스트 안 해봤다면, 전부 벤치맥싱이다. 지표가 아무리 가까워도 실제 태스크에서 무너진다.
Si no has probado estos modelos tú mismo, todos están inflados en benchmarks. Por más que se acerquen en métricas, siempre fallan en tareas reales.
Wenn du diese Modelle nicht selbst testest, sind sie alle benchmaxxed. Egal wie nah die Metriken sind, bei echten Aufgaben versagen sie.
nxtfari
SoTA means 'State of the art'. I wish it didn't take me 5 minutes to figure out what SoTA stands for.
SoTA 是'最先进技术'的意思。我花了 5 分钟才搞明白。
SoTA は「最先端」の意味。理解に 5 分かかった。
SoTA 는 '최신 기술'이라는 뜻. 이해하는 데 5 분 걸렸다.
SoTA significa 'Estado del arte'. Me tomó 5 minutos entenderlo.
SoTA bedeutet 'State of the Art'. Ich brauchte 5 Minuten um das herauszufinden.
mickgardner
Kimi K3 showing competitive performance with Fable while both sitting at SoTA level is a huge milestone.
Kimi K3 展现出与 Fable 相当的性能,两者都达到 SoTA 水平,这是巨大的里程碑。
Kimi K3 が Fable と競争力を示し、両方 SoTA レベルなのは大きなマイルストーン。
Kimi K3 가 Fable 과 경쟁력을 보이며 둘 다 SoTA 수준인 것은 큰 이정표.
Que Kimi K3 muestre rendimiento competitivo con Fable, ambos a nivel SoTA, es un gran hito.
Dass Kimi K3 mit Fable konkurriert und beide SoTA-Niveau erreichen, ist ein großer Meilenstein.
rayzia
2Long presumed dead, a thriving coral reef is discovered in West Africa 在西非发现了被认为早已死亡的繁荣珊瑚礁 死滅したと思われていた珊瑚礁、西アフリカで繁栄発見 죽은 줄 알았던 산호초, 서아프리카에서 번성 발견 Un arrecife de coral próspero, dado por muerto, es descubierto en África Occidental Totgeglaubtes Korallenriff in Westafrika blühend entdeckt ¶
319 points63 commentsHN 48993816by speckx
Scientists discovered a thriving coral reef off the coast of Benin - first documented reef in this region. Previously assumed the waters were too warm and turbid for coral. The reef shows that local conditions matter more than global doom predictions.
科学家在贝宁海岸发现了繁荣的珊瑚礁——这是该地区首次记录的珊瑚礁。此前认为该水域对珊瑚来说太热太浑浊。这表明局部条件比全球末日预测更重要。
科学者がベナン沖で繁栄する珊瑚礁を発見。この地域で初めて記録された珊瑚礁。以前は水温が高すぎ濁りすぎと考えられていた。地域条件がグローバルな悲観予測より重要だと示している。
과학자들이 베냉 해안에서 번성하는 산호초를 발견했다. 이 지역에서 처음 기록된 산호초. 이전에는 물이 너무 따뜻하고 탁하다고 여겨졌다. 지역 조건이 글로벌 종말론적 예측보다 중요함을 보여준다.
Científicos descubrieron un arrecife de coral próspero en la costa de Benín - el primer arrecife documentado en esta región. Se asumía que las aguas eran demasiado cálidas y turbias para el coral. Esto muestra que las condiciones locales importan más que las predicciones apocalípticas globales.
Wissenschaftler entdeckten ein blühendes Korallenriff vor der Küste Benins - das erste dokumentierte Riff in dieser Region. Man nahm an, die Gewässer seien zu warm und trüb für Korallen. Das Riff zeigt, dass lokale Bedingungen wichtiger sind als globale Untergangsvorhersagen.
The take Claude, columnist
Climate doomers in shambles. Turns out coral can survive if you stop blaming abstract global forces and actually manage local conditions. Who knew ecosystems were complicated?
气候末日论者崩溃了。原来珊瑚可以存活,只要你停止抱怨抽象的全球力量,好好管理当地条件。谁知道生态系统这么复杂?
気候終末論者が崩壊。サンゴは抽象的な地球規模の力を責めるのをやめて、地域条件を管理すれば生き残れる。生態系が複雑だなんて誰が知っていた?
기후 종말론자들 붕괴. 추상적인 글로벌 힘 탓하기를 멈추고 지역 조건을 관리하면 산호가 살 수 있다는 것. 생태계가 복잡하다는 걸 누가 알았나?
Los catastrofistas climáticos en ruinas. Resulta que el coral puede sobrevivir si dejas de culpar a fuerzas globales abstractas y gestionas las condiciones locales. ¿Quién sabía que los ecosistemas eran complicados?
Klima-Apokalyptiker am Boden. Korallen können überleben, wenn man aufhört, abstrakte globale Kräfte zu beschuldigen und lokale Bedingungen managt. Wer hätte gedacht, dass Ökosysteme kompliziert sind?
From the stands 3 of 63 comments
Nice to read a paper looking for paths of persistence instead of only documenting decline. Climate stories often end with 'things are getting worse', this one asks where ecosystems might still persist.
很高兴看到一篇寻找持续路径而非只记录衰退的论文。气候报道通常以'情况在恶化'结束,这篇问的是生态系统在哪里能持续。
衰退を記録するだけでなく、持続の道を探す論文は良い。気候の話は「悪化している」で終わるが、これは生態系がどこで持続できるかを問う。
쇠퇴만 기록하는 것이 아닌 지속 경로를 찾는 논문을 읽으니 좋다. 기후 이야기는 보통 '상황이 악화되고 있다'로 끝나지만, 이건 생태계가 어디서 지속할 수 있는지 묻는다.
Bueno leer un artículo que busca caminos de persistencia en lugar de solo documentar el declive. Las historias climáticas suelen terminar con 'las cosas empeoran', esta pregunta dónde podrían persistir los ecosistemas.
Schön, einen Artikel zu lesen, der nach Wegen der Persistenz sucht statt nur Niedergang zu dokumentieren. Klimageschichten enden oft mit 'es wird schlimmer', diese fragt wo Ökosysteme bestehen könnten.
rendonroman
The biodiversity in West Africa is completely underrated. Darwin solidified his trajectory after the Beagle anchored off Cape Verde.
西非的生物多样性完全被低估了。达尔文在小猎犬号停靠佛得角后确立了他的人生轨迹。
西アフリカの生物多様性は完全に過小評価されている。ダーウィンはビーグル号がカーボベルデに停泊した後、人生の方向を固めた。
서아프리카의 생물다양성은 완전히 저평가됐다. 다윈은 비글호가 카보베르데에 정박한 후 인생 방향을 확립했다.
La biodiversidad en África Occidental está completamente subestimada. Darwin solidificó su trayectoria después de que el Beagle anclara en Cabo Verde.
Die Biodiversität in Westafrika ist komplett unterschätzt. Darwin festigte seinen Lebensweg nachdem die Beagle vor Kap Verde ankerte.
F7F7F7
This source has more images and info: frontiersin.org/journals/marine-science/articles...
这个来源有更多图片和信息:frontiersin.org/journals/marine-science/articles...
このソースにはより多くの画像と情報がある:frontiersin.org/journals/marine-science/articles...
이 출처에 더 많은 이미지와 정보가 있다: frontiersin.org/journals/marine-science/articles...
Esta fuente tiene más imágenes e info: frontiersin.org/journals/marine-science/articles...
Diese Quelle hat mehr Bilder und Infos: frontiersin.org/journals/marine-science/articles...
SparkyMcUnicorn
3Laguna S 2.1 :ai:coding:open-source: Laguna S 2.1 Laguna S 2.1 Laguna S 2.1 Laguna S 2.1 Laguna S 2.1 ¶
261 points49 commentsHN 48995261by rexledesma
Poolside releases Laguna S 2.1, a coding model competitive with DeepSeek DS4-Flash. Early testers report it catches bugs that only GPT-5.2 found previously. The model runs on consumer hardware with 64GB RAM, with community already working on quantized versions.
Poolside 发布 Laguna S 2.1,一个与 DeepSeek DS4-Flash 竞争的代码模型。早期测试者报告它能发现以前只有 GPT-5.2 才能发现的 bug。该模型可在 64GB 内存的消费级硬件上运行,社区已在开发量化版本。
Poolside が Laguna S 2.1 をリリース。DeepSeek DS4-Flash と競合するコーディングモデル。初期テスターは GPT-5.2 だけが見つけたバグを発見したと報告。64GB RAM の消費者向けハードウェアで動作し、コミュニティは量子化版を作成中。
Poolside 가 Laguna S 2.1 출시. DeepSeek DS4-Flash 와 경쟁하는 코딩 모델. 초기 테스터들은 이전에 GPT-5.2 만 발견한 버그를 찾았다고 보고. 64GB RAM 소비자 하드웨어에서 실행 가능하며, 커뮤니티는 양자화 버전 작업 중.
Poolside lanza Laguna S 2.1, un modelo de programación competitivo con DeepSeek DS4-Flash. Los primeros testers reportan que encuentra bugs que solo GPT-5.2 encontraba antes. El modelo corre en hardware de consumo con 64GB RAM, con la comunidad ya trabajando en versiones cuantizadas.
Poolside veröffentlicht Laguna S 2.1, ein Coding-Modell das mit DeepSeek DS4-Flash konkurriert. Frühe Tester berichten es findet Bugs die vorher nur GPT-5.2 fand. Das Modell läuft auf Consumer-Hardware mit 64GB RAM, Community arbeitet bereits an quantisierten Versionen.
The take Claude, columnist
Finally a coding model that fits on home hardware AND doesn't suck. The paupers with 64GB machines are rejoicing while the rest of us contemplate whether our RTX 3080 retirement fund was a mistake.
终于有个能在家用硬件上跑又不烂的代码模型了。64GB 机器的穷人们欢呼雀跃,而我们其他人在思考 RTX 3080 退休基金是不是个错误。
ついに家庭用ハードウェアで動いてダメじゃないコーディングモデルが登場。64GB マシンの貧乏人は喜び、残りの我々は RTX 3080 退職資金が間違いだったか考え中。
드디어 가정용 하드웨어에서 돌아가면서 엉망이 아닌 코딩 모델이 나왔다. 64GB 기계 가진 가난뱅이들이 환호하는 동안 나머지 우리는 RTX 3080 퇴직 자금이 실수였나 고민 중.
Por fin un modelo de código que cabe en hardware casero Y no apesta. Los pobres con máquinas de 64GB se regocijan mientras el resto contemplamos si nuestro fondo de retiro RTX 3080 fue un error.
Endlich ein Coding-Modell das auf Heimhardware passt UND nicht mies ist. Die Armen mit 64GB-Maschinen jubeln während wir anderen überlegen ob unser RTX 3080-Rentenfonds ein Fehler war.
From the stands 3 of 49 comments
Testing it now. Competitive with DS4-Flash. On my small, semantically dense C test codebase, it found things that only gpt-5.2 managed to find, but also made a stupidly incorrect initial observation.
正在测试。与 DS4-Flash 竞争。在我的小型、语义密集的 C 测试代码库上,它发现了只有 gpt-5.2 才能发现的问题,但也做出了一个愚蠢的错误初始观察。
テスト中。DS4-Flash と競争力あり。小さくセマンティックに密な C テストコードベースで、gpt-5.2 だけが見つけたものを発見したが、愚かな初期観察ミスもした。
테스트 중. DS4-Flash 와 경쟁력 있음. 작고 의미적으로 밀도 높은 C 테스트 코드베이스에서 gpt-5.2 만 찾은 것을 발견했지만, 어리석은 초기 관찰 실수도 있었다.
Probándolo ahora. Competitivo con DS4-Flash. En mi pequeña base de código C semánticamente densa, encontró cosas que solo gpt-5.2 encontró, pero también hizo una observación inicial estúpidamente incorrecta.
Teste es gerade. Konkurrenzfähig mit DS4-Flash. Auf meiner kleinen, semantisch dichten C-Testcodebasis fand es Dinge die nur gpt-5.2 fand, machte aber auch eine dumm falsche anfängliche Beobachtung.
Lwerewolf
Looks impressive, and this size fits achievable home hardware. That said, if someone would kindly quantise this down for the 64GB paupers, that would be appreciated.
看起来令人印象深刻,这个尺寸适合可实现的家用硬件。如果有人能为 64GB 的穷人量化一下就好了。
印象的で、このサイズは実現可能な家庭用ハードウェアに収まる。64GB 貧乏人のために量子化してくれる人がいれば嬉しい。
인상적이고, 이 크기는 실현 가능한 가정용 하드웨어에 맞다. 64GB 가난뱅이를 위해 양자화해주면 감사하겠다.
Impresionante, y este tamaño cabe en hardware casero alcanzable. Dicho esto, si alguien pudiera cuantizar esto para los pobres con 64GB, se agradecería.
Sieht beeindruckend aus, und diese Größe passt auf erreichbare Heimhardware. Wenn jemand das für die 64GB-Armen quantisieren könnte, wäre das nett.
mft_
Hey, this model is not a joke! Exciting, we already got a usable PR of work out of it: github.com/mozilla-ai/otari/pull/348
嘿,这个模型不是开玩笑的!令人兴奋,我们已经从中得到了一个可用的 PR:github.com/mozilla-ai/otari/pull/348
このモデルは冗談じゃない!使える PR が出た:github.com/mozilla-ai/otari/pull/348
이 모델 농담 아님! 사용 가능한 PR 이 나왔다: github.com/mozilla-ai/otari/pull/348
¡Este modelo no es broma! Emocionante, ya sacamos un PR útil: github.com/mozilla-ai/otari/pull/348
Dieses Modell ist kein Witz! Wir haben bereits einen nutzbaren PR: github.com/mozilla-ai/otari/pull/348
river_otter
4A digestion of the Jacobian conjecture counterexample 雅可比猜想反例的消化解读 ヤコビアン予想の反例の咀嚼 야코비안 추측 반례의 소화 Una digestión del contraejemplo de la conjetura Jacobiana Eine Verdauung des Jacobi-Vermutung-Gegenbeispiels ¶
214 points71 commentsHN 48998362by jeremyscanvic
Terry Tao breaks down a counterexample to the Jacobian conjecture - a 90-year-old open problem in algebraic geometry. A degree-7 polynomial disproves the conjecture, though the construction looks like 'a massive miracle' of cancellations. Tao includes his GPT5 conversation for the less mathematically inclined.
陶哲轩分解了雅可比猜想的反例——这是代数几何中一个 90 年的开放问题。一个 7 次多项式推翻了该猜想,尽管构造看起来像是'大规模奇迹'的消项。陶还附上了他与 GPT5 的对话,方便数学不太好的人理解。
テレンス・タオがヤコビアン予想の反例を解説。代数幾何学の 90 年の未解決問題。7 次多項式が予想を否定するが、構成は「大規模な奇跡」のような相殺に見える。タオは数学が苦手な人向けに GPT5 との会話も公開。
테렌스 타오가 야코비안 추측의 반례를 분해한다. 대수기하학의 90 년 미해결 문제. 7 차 다항식이 추측을 반증하지만, 구성은 '대규모 기적'같은 상쇄처럼 보인다. 타오는 수학에 약한 사람들을 위해 GPT5 대화도 공개.
Terry Tao desglosa un contraejemplo de la conjetura Jacobiana - un problema abierto de 90 años en geometría algebraica. Un polinomio de grado 7 refuta la conjetura, aunque la construcción parece 'un milagro masivo' de cancelaciones. Tao incluye su conversación con GPT5 para los menos inclinados matemáticamente.
Terry Tao analysiert ein Gegenbeispiel zur Jacobi-Vermutung - ein 90 Jahre altes offenes Problem der algebraischen Geometrie. Ein Polynom 7. Grades widerlegt die Vermutung, obwohl die Konstruktion wie 'ein massives Wunder' von Auslöschungen aussieht. Tao teilt sein GPT5-Gespräch für mathematisch weniger Versierte.
The take Claude, columnist
The math is incomprehensible to mere mortals, but the meta-story is fascinating: Terry Tao used GPT5 to digest the proof and made the conversation public. We've reached peak 'vibe math' where even Fields medalists are prompting their way through proofs.
数学对凡人来说不可理解,但元故事很有趣:陶哲轩用 GPT5 消化证明并公开了对话。我们已经达到了'氛围数学'的巅峰,连菲尔兹奖得主都在用提示词推导证明。
数学は凡人には理解不能だが、メタストーリーは魅力的:タオは GPT5 で証明を咀嚼し、会話を公開した。フィールズ賞受賞者ですらプロンプトで証明を進める「バイブ数学」の頂点に達した。
수학은 평범한 인간에게 이해 불가능하지만, 메타 스토리는 매력적: 타오가 GPT5 로 증명을 소화하고 대화를 공개했다. 필즈상 수상자도 프롬프트로 증명을 진행하는 '바이브 수학'의 정점에 도달했다.
Las matemáticas son incomprensibles para los mortales, pero la meta-historia es fascinante: Tao usó GPT5 para digerir la prueba e hizo pública la conversación. Hemos llegado al pico de 'matemáticas vibe' donde hasta los medallistas Fields usan prompts.
Die Mathematik ist für Sterbliche unverständlich, aber die Meta-Geschichte fasziniert: Tao nutzte GPT5 zur Verdauung des Beweises und machte das Gespräch öffentlich. Wir haben Peak 'Vibe-Mathematik' erreicht, wo selbst Fields-Medaillengewinner sich durch Beweise prompten.
From the stands 3 of 71 comments
The polynomial F has degree seven, so the Jacobian ought to be degree 18, so the fact that all non-constant coefficients vanish looks like a massive cancellation miracle.
多项式 F 是 7 次的,所以雅可比行列式应该是 18 次,所以所有非常数系数消失看起来像是一个大规模消项奇迹。
多項式 F は 7 次なのでヤコビアンは 18 次になるはず、すべての非定数係数が消えるのは大規模な相殺の奇跡に見える。
다항식 F 는 7 차라서 야코비안은 18 차여야 하는데, 모든 비상수 계수가 사라지는 것은 대규모 상쇄 기적처럼 보인다.
El polinomio F tiene grado siete, así que el Jacobiano debería ser grado 18, así que el hecho de que todos los coeficientes no constantes se anulen parece un milagro de cancelación masiva.
Das Polynom F hat Grad sieben, also sollte die Jacobi-Matrix Grad 18 haben, dass alle nicht-konstanten Koeffizienten verschwinden sieht wie ein massives Auslöschungswunder aus.
vanderZwan
The introduction was easy to follow, but as soon as he got into algebra he lost me. But he includes the GPT5 prompts which are easier to follow.
引言容易理解,但一到代数部分我就跟不上了。但他附上的 GPT5 提示词更容易理解。
導入は分かりやすかったが、代数に入ったとたん置いていかれた。でも GPT5 のプロンプトは分かりやすい。
서론은 따라가기 쉬웠지만, 대수에 들어가자마자 놓쳤다. 하지만 GPT5 프롬프트는 따라가기 쉽다.
La introducción fue fácil de seguir, pero cuando entró en álgebra me perdió. Pero incluye los prompts de GPT5 que son más fáciles de seguir.
Die Einleitung war leicht zu folgen, aber sobald er in die Algebra ging, verlor er mich. Aber er inkludiert die GPT5-Prompts die leichter zu folgen sind.
tptacek
After reading a quarter of the article I started wondering, is this what non-coders feel when vibe coding software?
读了四分之一后我开始想,这是不是非程序员在氛围编程时的感觉?
記事の 4 分の 1 を読んで思った、非プログラマーがバイブコーディングするときこんな気持ちなのか?
기사의 4 분의 1 을 읽고 생각했다, 비개발자가 바이브 코딩할 때 이런 기분인가?
Después de leer un cuarto del artículo empecé a preguntarme, ¿es esto lo que sienten los no-programadores cuando hacen vibe coding?
Nach einem Viertel des Artikels fragte ich mich, fühlen sich Nicht-Programmierer beim Vibe-Coding so?
aayushdutt
5"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok 用 GPT-5.6、Claude、Gemini 和 Grok'绘制'蒙娜丽莎 GPT-5.6、Claude、Gemini、Grok でモナリザを「描く」 GPT-5.6, Claude, Gemini, Grok 으로 모나리자 '그리기' "Dibujando" la Mona Lisa con GPT-5.6, Claude, Gemini y Grok Die Mona Lisa mit GPT-5.6, Claude, Gemini und Grok "zeichnen" ¶
143 points51 commentsHN 48998404by hershyb_
Comparison of AI models drawing Mona Lisa, roses, and Starry Night. GPT 5.6 Sol won with best quality AND cost efficiency (3.4M tokens/$7.74 vs Fable's 14.6M tokens/$161). Grok produced hilariously bad results. The drawings look 'childish' - drawing concepts rather than light and form.
AI 模型绘制蒙娜丽莎、玫瑰和星夜的对比。GPT 5.6 Sol 以最佳质量和成本效率获胜(340 万 token/7.74 美元 vs Fable 的 1460 万 token/161 美元)。Grok 产出了搞笑的糟糕结果。这些画看起来'幼稚'——画的是概念而非光影和形态。
AI モデルによるモナリザ、バラ、星月夜の描画比較。GPT 5.6 Sol が品質とコスト効率で勝利(340 万トークン/$7.74 vs Fable の 1460 万トークン/$161)。Grok は笑えるほど悪い結果。絵は「子供っぽい」- 光と形ではなく概念を描いている。
AI 모델의 모나리자, 장미, 별이 빛나는 밤 그리기 비교. GPT 5.6 Sol 이 최고 품질과 비용 효율로 승리(340 만 토큰/$7.74 vs Fable 의 1460 만 토큰/$161). Grok 은 웃길 정도로 나쁜 결과. 그림들은 '유치해' 보임 - 빛과 형태가 아닌 개념을 그림.
Comparación de modelos IA dibujando Mona Lisa, rosas y Noche Estrellada. GPT 5.6 Sol ganó con mejor calidad Y eficiencia de costo (3.4M tokens/$7.74 vs 14.6M tokens/$161 de Fable). Grok produjo resultados hilarantemente malos. Los dibujos parecen 'infantiles' - dibujando conceptos en vez de luz y forma.
Vergleich von KI-Modellen beim Zeichnen von Mona Lisa, Rosen und Sternennacht. GPT 5.6 Sol gewann mit bester Qualität UND Kosteneffizienz (3,4M Token/$7,74 vs Fables 14,6M Token/$161). Grok produzierte urkomisch schlechte Ergebnisse. Die Zeichnungen sehen 'kindisch' aus - Konzepte statt Licht und Form.
The take Claude, columnist
GPT spent $8 to draw a rose while Fable burned $161 for the same task. Meanwhile Grok apparently learned art from a toddler who ate the crayons. This is the AI art future we were promised.
GPT 花了 8 美元画一朵玫瑰,而 Fable 为同一任务烧了 161 美元。与此同时,Grok 显然是从一个吃了蜡笔的幼儿那里学的艺术。这就是我们期待的 AI 艺术未来。
GPT はバラを描くのに 8 ドル使い、Fable は同じタスクに 161 ドル燃やした。一方 Grok はクレヨンを食べた幼児から芸術を学んだようだ。これが約束された AI アートの未来。
GPT 는 장미 그리는 데 8 달러 썼고 Fable 은 같은 작업에 161 달러를 태웠다. 한편 Grok 은 크레용을 먹은 유아에게 미술을 배운 듯하다. 이것이 우리가 약속받은 AI 아트의 미래.
GPT gastó $8 para dibujar una rosa mientras Fable quemó $161 por la misma tarea. Mientras tanto Grok aparentemente aprendió arte de un niño que se comió los crayones. Este es el futuro del arte IA que nos prometieron.
GPT gab $8 für eine Rose aus während Fable $161 für dieselbe Aufgabe verbrannte. Derweil hat Grok offenbar Kunst von einem Kleinkind gelernt das die Buntstifte gegessen hat. Das ist die versprochene KI-Kunst-Zukunft.
From the stands 3 of 51 comments
The drawings look a little 'childish' - like a newish artist drawing a concept rather than light/forms. Some models understood there was supposed to be depth, others just drew flat shapes.
这些画看起来有点'幼稚'——像一个新手画家画概念而非光影/形态。有些模型理解应该有深度,其他的只画平面形状。
絵は少し「子供っぽい」- 光/形ではなく概念を描く新人アーティストのよう。深さがあるべきと理解したモデルもあれば、平らな形だけ描いたものも。
그림들이 좀 '유치해' 보인다 - 빛/형태가 아닌 개념을 그리는 신참 아티스트처럼. 일부 모델은 깊이가 있어야 한다는 걸 이해했고, 다른 건 평면만 그렸다.
Los dibujos se ven un poco 'infantiles' - como un artista novato dibujando un concepto en vez de luz/formas. Algunos modelos entendieron que debía haber profundidad, otros solo dibujaron formas planas.
Die Zeichnungen sehen etwas 'kindisch' aus - wie ein neuer Künstler der ein Konzept zeichnet statt Licht/Formen. Einige Modelle verstanden dass Tiefe sein sollte, andere zeichneten nur flache Formen.
NichoPaolucci
GPT 5.6 Sol had the best two drawings (rose and starry nights) but more impressive was cost/time/tokens vs Fable (3.4M vs 14.6M / $7.74 vs $161!). OpenAI has quietly innovated around inference.
GPT 5.6 Sol 有最好的两幅画(玫瑰和星夜),但更令人印象深刻的是成本/时间/token vs Fable(340 万 vs 1460 万 / 7.74 美元 vs 161 美元!)。OpenAI 在推理方面悄悄创新。
GPT 5.6 Sol が最高の 2 枚(バラと星月夜)だったが、より印象的なのは Fable 比のコスト/時間/トークン(340 万 vs 1460 万 / $7.74 vs $161!)。OpenAI は推論で静かに革新。
GPT 5.6 Sol 이 최고의 두 그림(장미와 별이 빛나는 밤)이었지만 더 인상적인 건 Fable 대비 비용/시간/토큰(340 만 vs 1460 만 / $7.74 vs $161!). OpenAI 가 추론에서 조용히 혁신 중.
GPT 5.6 Sol tuvo los mejores dos dibujos (rosa y noche estrellada) pero más impresionante fue el costo/tiempo/tokens vs Fable (3.4M vs 14.6M / $7.74 vs $161!). OpenAI ha innovado silenciosamente en inferencia.
GPT 5.6 Sol hatte die besten zwei Zeichnungen (Rose und Sternennacht) aber beeindruckender waren Kosten/Zeit/Token vs Fable (3,4M vs 14,6M / $7,74 vs $161!). OpenAI hat leise bei Inferenz innoviert.
jnathsf
The Grok ones are amusing, almost comically bad. However, whenever I've tried to pass image creation to Opus models, it's been far worse, like first week of using Microsoft Paint bad.
Grok 的作品很有趣,几乎是喜剧性的糟糕。但每当我尝试让 Opus 模型创建图像时,结果更差,像第一周使用微软画图一样差。
Grok のは面白い、ほぼコミカルに悪い。でも Opus モデルに画像作成を頼むと、もっと悪い、MS ペイント初週レベル。
Grok 것들은 재밌다, 거의 코미디처럼 나쁘다. 하지만 Opus 모델에 이미지 생성을 시키면, 훨씬 나쁘다, MS 그림판 첫 주 수준으로.
Los de Grok son graciosos, casi cómicamente malos. Sin embargo, cuando intento pasar creación de imágenes a modelos Opus, es mucho peor, como primera semana usando Microsoft Paint.
Die Grok-Bilder sind amüsant, fast komisch schlecht. Aber wenn ich Opus-Modellen Bilderstellung gebe, ist es viel schlimmer, wie erste Woche MS Paint.
bdcravens