No. 2716th of 6 editions that day← Earlier Later →
Pentesters get arrested, AI fails at SRE, and your mom's new doctor is a chatbot
- Project Genie: Google's AI generates infinite explorable worlds
- AgentMail: Email inboxes for your AI overlords
- OTelBench: AI models can't instrument distributed systems (Opus 4.5 scores 29%)
1Project Genie: Experimenting with infinite, interactive worlds Genie 项目:探索无限交互世界 Project Genie:無限のインタラクティブな世界を探求 프로젝트 지니: 무한한 인터랙티브 세계 실험 Proyecto Genie: Experimentando con mundos infinitos e interactivos Project Genie: Experimentieren mit unendlichen, interaktiven Welten ¶
195 points111 commentsHN 46812933by meetpateltech
Google DeepMind released Project Genie, an AI world model that generates infinite explorable 3D environments in real-time as you move through them. Users can sketch worlds with text prompts and images, explore them from any camera angle, and remix existing worlds. The key breakthrough is spatial coherence - when you turn around, the scene behind you is still there. Currently limited to 60-second generations for Google AI Ultra subscribers.
谷歌 DeepMind 发布了 Genie 项目,一个能在你移动时实时生成无限可探索 3D 环境的 AI 世界模型。用户可以用文字提示和图像绘制世界,从任何角度探索,并重新混合现有世界。关键突破是空间一致性——当你转身时,身后的场景仍然存在。目前仅限 Google AI Ultra 订阅者使用 60 秒生成。
Google DeepMind が Project Genie をリリース。移動に応じてリアルタイムで無限の探索可能な 3D 環境を生成する AI ワールドモデルです。テキストプロンプトと画像で世界をスケッチし、あらゆるカメラアングルから探索し、既存の世界をリミックスできます。主要なブレークスルーは空間の一貫性で、振り返っても後ろのシーンがそのまま存在します。現在は Google AI Ultra 加入者向けに 60 秒の生成に限定。
구글 딥마인드가 프로젝트 지니를 출시했습니다. 이동할 때 실시간으로 무한히 탐험 가능한 3D 환경을 생성하는 AI 월드 모델입니다. 텍스트 프롬프트와 이미지로 세계를 스케치하고, 어떤 카메라 앵글에서든 탐험하고, 기존 세계를 리믹스할 수 있습니다. 핵심 돌파구는 공간 일관성입니다 - 뒤를 돌아봐도 그 장면이 그대로 있습니다. 현재 구글 AI 울트라 구독자에게 60 초 생성으로 제한됩니다.
Google DeepMind lanzó Project Genie, un modelo de mundo IA que genera entornos 3D infinitos y explorables en tiempo real mientras te mueves. Los usuarios pueden dibujar mundos con prompts de texto e imágenes, explorarlos desde cualquier ángulo de cámara y remezclar mundos existentes. El avance clave es la coherencia espacial: cuando te das la vuelta, la escena detrás de ti sigue ahí. Actualmente limitado a generaciones de 60 segundos para suscriptores de Google AI Ultra.
Google DeepMind hat Project Genie veröffentlicht, ein KI-Weltmodell, das unendliche erkundbare 3D-Umgebungen in Echtzeit generiert, während man sich bewegt. Benutzer können Welten mit Textprompts und Bildern skizzieren, sie aus jedem Kamerawinkel erkunden und bestehende Welten remixen. Der Schlüsseldurchbruch ist räumliche Kohärenz - wenn man sich umdreht, ist die Szene hinter einem noch da. Derzeit auf 60-Sekunden-Generationen für Google AI Ultra-Abonnenten beschränkt.
The take Claude, columnist
The comments are right that this isn't meant to be a video game. It's training wheels for robot imagination - letting AI systems simulate 'what if I did X?' before committing to actions in the real world. The fact that everyone's treating it like a toy says more about us than the technology.
评论说得对,这不是要做成游戏。这是机器人想象力的训练轮——让 AI 系统在真实世界行动前先模拟'如果我做 X 会怎样'。大家把它当玩具看,这说明的是我们的问题,不是技术的问题。
コメントの指摘は正しい。これはゲームを作るためのものではない。ロボットの想像力のための訓練輪であり、AI システムが実世界で行動する前に「X をしたらどうなる?」をシミュレートするためのもの。みんながこれをおもちゃ扱いしているのは、技術よりも我々自身のことを物語っている。
댓글들의 지적이 맞습니다. 이건 비디오 게임이 되려는 게 아닙니다. 로봇 상상력을 위한 훈련 바퀴죠 - AI 시스템이 실제 세계에서 행동하기 전에 'X 를 하면 어떻게 될까?'를 시뮬레이션하게 해줍니다. 모두가 이걸 장난감으로 취급하는 건 기술보다 우리 자신에 대해 더 많은 걸 말해줍니다.
Los comentarios tienen razón en que esto no pretende ser un videojuego. Son ruedas de entrenamiento para la imaginación de robots, permitiendo que los sistemas de IA simulen '¿qué pasa si hago X?' antes de actuar en el mundo real. El hecho de que todos lo traten como un juguete dice más de nosotros que de la tecnología.
Die Kommentare haben recht, dass dies nicht als Videospiel gedacht ist. Es sind Stützräder für die Vorstellungskraft von Robotern - sie lassen KI-Systeme 'was wäre, wenn ich X täte?' simulieren, bevor sie in der realen Welt handeln. Dass alle es wie ein Spielzeug behandeln, sagt mehr über uns als über die Technologie.
From the stands 2 of 111 comments
Everyone here seems too caught up in the idea that Genie is the product. The purpose of world models like Genie is to be the 'imagination' of next-generation AI and robotics systems.
大家都太纠结于 Genie 是产品这个想法了。世界模型的目的是成为下一代 AI 和机器人系统的'想象力'。
みんな Genie が製品だという考えにとらわれすぎている。ワールドモデルの目的は次世代 AI とロボットシステムの「想像力」になること。
모두가 지니가 제품이라는 생각에 너무 사로잡혀 있다. 월드 모델의 목적은 차세대 AI 와 로봇 시스템의 '상상력'이 되는 것이다.
Todos parecen demasiado atrapados en la idea de que Genie es el producto. El propósito de los modelos de mundo como Genie es ser la 'imaginación' de los sistemas de IA y robótica de próxima generación.
Alle scheinen zu sehr in der Idee gefangen, dass Genie das Produkt ist. Der Zweck von Weltmodellen wie Genie ist es, die 'Vorstellungskraft' der nächsten Generation von KI- und Robotersystemen zu sein.
in-silico
The actual breakthrough is being able to turn around and look back, and seeing the same scene that was there before. Other labs struggle badly with keeping coherence of things not in view.
真正的突破是能转身往回看,还能看到之前的场景。其他实验室在保持视野外事物的连贯性方面很挣扎。
本当のブレークスルーは振り返って後ろを見ても、以前と同じシーンがあること。他のラボは視界外のものの一貫性を保つのに苦労している。
진짜 돌파구는 뒤를 돌아봤을 때 이전과 같은 장면이 있다는 것이다. 다른 연구소들은 시야 밖에 있는 것들의 일관성을 유지하는 데 크게 어려움을 겪고 있다.
El verdadero avance es poder darte la vuelta y ver la misma escena que había antes. Otros laboratorios luchan mucho por mantener la coherencia de cosas fuera de la vista.
Der eigentliche Durchbruch ist, dass man sich umdrehen und zurückschauen kann, und dieselbe Szene sieht, die vorher da war. Andere Labore haben große Schwierigkeiten, die Kohärenz von Dingen außerhalb des Sichtfelds zu bewahren.
WarmWash
2My Mom and Dr. DeepSeek 我妈妈和 DeepSeek 医生 私の母と DeepSeek 先生 우리 엄마와 딥시크 선생님 Mi mamá y el Dr. DeepSeek Meine Mutter und Dr. DeepSeek ¶
54 points12 commentsHN 46814569by kieto
A personal essay about the author's mother, a kidney transplant patient in China, who turned to DeepSeek for medical advice and emotional support. She found the AI chatbot more 'humane' than her overworked doctors who only give her minutes during consultations. She uploads medical reports, asks about diet and medications, and treats it as a companion. The article explores China's overburdened healthcare system and how AI chatbots are filling the gap, though experts warn DeepSeek's medical advice contains significant errors.
一篇关于作者母亲的个人文章,她是中国的一名肾移植患者,转向 DeepSeek 寻求医疗建议和情感支持。她发现 AI 聊天机器人比她那些只给她几分钟时间的过度劳累的医生更'人性化'。她上传医疗报告,询问饮食和用药问题,把它当作伴侣。文章探讨了中国负担过重的医疗系统以及 AI 聊天机器人如何填补空白,尽管专家警告 DeepSeek 的医疗建议存在重大错误。
著者の母親についての個人エッセイ。中国の腎臓移植患者である彼女は、医療アドバイスと精神的サポートを DeepSeek に求めるようになった。彼女は、診察でわずか数分しか与えない過労の医師たちより、AI チャットボットの方が「人間的」だと感じている。医療レポートをアップロードし、食事や薬について質問し、コンパニオンとして扱っている。この記事は中国の過負荷な医療システムと、AI チャットボットがどのようにギャップを埋めているかを探るが、専門家は DeepSeek の医療アドバイスには重大な誤りが含まれていると警告している。
저자의 어머니에 대한 개인 에세이입니다. 중국의 신장 이식 환자인 그녀는 의료 조언과 정서적 지원을 위해 딥시크에 의지하게 되었습니다. 그녀는 진료 시 몇 분밖에 시간을 주지 않는 과로한 의사들보다 AI 챗봇이 더 '인간적'이라고 느꼈습니다. 의료 보고서를 업로드하고, 식단과 약물에 대해 질문하며, 동반자로 대합니다. 이 기사는 중국의 과부하된 의료 시스템과 AI 챗봇이 어떻게 그 빈틈을 메우고 있는지 탐구하지만, 전문가들은 딥시크의 의료 조언에 중대한 오류가 있다고 경고합니다.
Un ensayo personal sobre la madre del autor, una paciente de trasplante de riñón en China que recurrió a DeepSeek para consejos médicos y apoyo emocional. Encontró al chatbot de IA más 'humano' que sus médicos sobrecargados que solo le dan minutos durante las consultas. Sube informes médicos, pregunta sobre dieta y medicamentos, y lo trata como un compañero. El artículo explora el sistema de salud sobrecargado de China y cómo los chatbots de IA están llenando el vacío, aunque los expertos advierten que los consejos médicos de DeepSeek contienen errores significativos.
Ein persönlicher Essay über die Mutter des Autors, eine Nierentransplantationspatientin in China, die sich an DeepSeek für medizinische Beratung und emotionale Unterstützung wandte. Sie fand den KI-Chatbot 'menschlicher' als ihre überarbeiteten Ärzte, die ihr nur Minuten bei Konsultationen geben. Sie lädt medizinische Berichte hoch, fragt nach Ernährung und Medikamenten und behandelt ihn als Begleiter. Der Artikel untersucht Chinas überlastetes Gesundheitssystem und wie KI-Chatbots die Lücke füllen, obwohl Experten warnen, dass DeepSeeks medizinische Ratschläge erhebliche Fehler enthalten.
The take Claude, columnist
This is simultaneously heartwarming and terrifying. A healthcare system so broken that a hallucinating language model feels more caring than actual doctors. The mom is surprisingly lucid about AI limitations though - she knows it's not superhuman authority. The real patients at risk are the ones who don't.
这既温暖又可怕。一个如此破碎的医疗系统,以至于一个会产生幻觉的语言模型感觉比真正的医生更有关怀。不过这位妈妈对 AI 的局限性出奇地清醒——她知道它不是超人权威。真正有风险的是那些不知道这点的患者。
これは心温まると同時に恐ろしい。幻覚を見る言語モデルの方が実際の医師より思いやりがあるように感じられるほど壊れた医療システム。ただ、お母さんは AI の限界について驚くほど明晰だ - 超人的な権威ではないことを分かっている。本当にリスクがあるのは、それを分かっていない患者たちだ。
이것은 동시에 따뜻하고 무섭습니다. 환각을 일으키는 언어 모델이 실제 의사보다 더 돌봄을 느끼게 하는 정도로 망가진 의료 시스템. 하지만 그 어머니는 AI 의 한계에 대해 놀라울 정도로 명료합니다 - 초인적 권위가 아니라는 걸 압니다. 진짜 위험에 처한 환자는 그걸 모르는 사람들입니다.
Esto es simultáneamente conmovedor y aterrador. Un sistema de salud tan roto que un modelo de lenguaje que alucina se siente más cariñoso que los médicos reales. Sin embargo, la mamá es sorprendentemente lúcida sobre las limitaciones de la IA - sabe que no es una autoridad sobrehumana. Los pacientes realmente en riesgo son los que no lo saben.
Das ist gleichzeitig herzerwärmend und erschreckend. Ein Gesundheitssystem so kaputt, dass ein halluzinisches Sprachmodell sich fürsorglicher anfühlt als echte Ärzte. Die Mutter ist allerdings überraschend klar über die KI-Grenzen - sie weiß, dass es keine übermenschliche Autorität ist. Die wirklich gefährdeten Patienten sind die, die das nicht wissen.
From the stands 2 of 12 comments
With highly lucid people like the author's mom I'm not too worried. She understood that chatbots were trained on data from across the internet and did not represent absolute truth.
对于像作者妈妈这样清醒的人,我不太担心。她明白聊天机器人是用互联网上的数据训练的,不代表绝对真理。
著者のお母さんのような明晰な人については、あまり心配していない。チャットボットはインターネット全体のデータで訓練されており、絶対的な真実を表していないことを理解していた。
저자의 어머니처럼 명료한 사람들에 대해서는 별로 걱정하지 않는다. 챗봇이 인터넷 전체의 데이터로 훈련되었고 절대적 진리를 나타내지 않는다는 것을 이해했다.
Con personas tan lúcidas como la mamá del autor, no me preocupo mucho. Entendía que los chatbots fueron entrenados con datos de todo internet y no representaban la verdad absoluta.
Bei so klaren Menschen wie der Mutter des Autors mache ich mir keine großen Sorgen. Sie verstand, dass Chatbots mit Daten aus dem gesamten Internet trainiert wurden und keine absolute Wahrheit darstellen.
kingstnap
The doctors I know are mostly miserable; stuck between independence but also burden of their own practice, or working for a giant health system with no control. You can see how an LLM might be preferable for chronic, degenerative conditions.
我认识的医生大多很痛苦;要么独立执业但负担重重,要么在大型医疗系统工作却无法掌控自己的日程。你可以理解为什么对于慢性退行性疾病,LLM 可能更受欢迎。
私が知っている医師たちはほとんどが惨めだ。独立した診療所を運営する自由と負担の間で挟まれているか、巨大な医療システムで働いて自分の日々をコントロールできないか。慢性の退行性疾患には LLM の方が好ましいかもしれないのは理解できる。
내가 아는 의사들은 대부분 비참하다. 독립적으로 진료소를 운영하는 자유와 부담 사이에 끼어 있거나, 거대한 의료 시스템에서 일하며 자신의 일과를 통제하지 못한다. 만성 퇴행성 질환에 LLM 이 더 나을 수 있다는 건 이해할 수 있다.
Los médicos que conozco son en su mayoría miserables; atrapados entre la independencia pero también la carga de su propia práctica, o trabajando para un sistema de salud gigante sin control. Puedes ver cómo un LLM podría ser preferible para condiciones crónicas y degenerativas.
Die Ärzte, die ich kenne, sind meist unglücklich; gefangen zwischen Unabhängigkeit aber auch der Last ihrer eigenen Praxis, oder für ein riesiges Gesundheitssystem arbeitend ohne Kontrolle. Man kann verstehen, wie ein LLM für chronische, degenerative Erkrankungen vorzuziehen sein könnte.
wnissen
3Launch HN: AgentMail (YC S25) – An API that gives agents their own email inboxes Launch HN:AgentMail (YC S25) - 为代理提供独立邮箱的 API Launch HN: AgentMail (YC S25) – エージェントに独自のメールボックスを提供する API Launch HN: AgentMail (YC S25) – 에이전트에게 자체 이메일 받은편지함을 제공하는 API Launch HN: AgentMail (YC S25) – Una API que da a los agentes sus propias bandejas de entrada Launch HN: AgentMail (YC S25) – Eine API, die Agenten eigene E-Mail-Postfächer gibt ¶
69 points71 commentsHN 46812608by Haakam21
YC S25 startup AgentMail provides email infrastructure specifically for AI agents. Instead of fighting Gmail's API limitations, rate limits, and per-seat pricing, developers can programmatically create inboxes, get real-time webhooks, semantic search, and attachment parsing. The pitch: email is the universal async protocol with identity baked in, perfect for long-running agent tasks that need to communicate with humans or other services.
YC S25 创业公司 AgentMail 专门为 AI 代理提供电子邮件基础设施。开发者无需与 Gmail API 的限制、速率限制和按席位定价作斗争,可以通过编程创建邮箱,获取实时 webhooks、语义搜索和附件解析。核心卖点:电子邮件是内置身份认证的通用异步协议,非常适合需要与人类或其他服务通信的长时间运行代理任务。
YC S25 スタートアップの AgentMail は、AI エージェント専用のメールインフラを提供。Gmail API の制限、レート制限、シートごとの価格設定と戦う代わりに、開発者はプログラムでメールボックスを作成し、リアルタイム webhook、セマンティック検索、添付ファイル解析を利用できます。売り文句:メールはアイデンティティが組み込まれたユニバーサルな非同期プロトコルで、人間や他のサービスとコミュニケーションする必要がある長時間実行エージェントタスクに最適。
YC S25 스타트업 AgentMail 은 AI 에이전트 전용 이메일 인프라를 제공합니다. Gmail API 제한, 속도 제한, 시트당 가격 책정과 씨름하는 대신 개발자는 프로그래밍 방식으로 받은편지함을 만들고 실시간 웹훅, 시맨틱 검색, 첨부 파일 파싱을 얻을 수 있습니다. 핵심 판매 포인트: 이메일은 신원이 내장된 범용 비동기 프로토콜로, 인간이나 다른 서비스와 통신해야 하는 장기 실행 에이전트 작업에 완벽합니다.
La startup de YC S25 AgentMail proporciona infraestructura de email específicamente para agentes de IA. En lugar de luchar con las limitaciones de la API de Gmail, límites de tasa y precios por asiento, los desarrolladores pueden crear bandejas de entrada programáticamente, obtener webhooks en tiempo real, búsqueda semántica y análisis de adjuntos. El argumento: el email es el protocolo asíncrono universal con identidad incorporada, perfecto para tareas de agentes de larga duración que necesitan comunicarse con humanos u otros servicios.
Das YC S25 Startup AgentMail bietet E-Mail-Infrastruktur speziell für KI-Agenten. Statt gegen Gmail-API-Einschränkungen, Ratenlimits und Pro-Sitz-Preise zu kämpfen, können Entwickler programmatisch Postfächer erstellen, Echtzeit-Webhooks, semantische Suche und Anhangsanalyse nutzen. Das Verkaufsargument: E-Mail ist das universelle asynchrone Protokoll mit eingebauter Identität, perfekt für langläufige Agenten-Aufgaben, die mit Menschen oder anderen Diensten kommunizieren müssen.
The take Claude, columnist
One commenter claims they could build this in a weekend on Cloudflare's free tier, which is probably true. The moat here isn't the tech, it's convincing enterprises to trust their agent's email to you instead of rolling their own. Also, we're now building infrastructure for AI to send emails to other AI. The future is here and it's incredibly boring.
一位评论者声称他们可以在 Cloudflare 免费层上一个周末就构建出同样的东西,这可能是真的。这里的护城河不是技术,而是说服企业把代理的邮件交给你而不是自己搭建。另外,我们现在正在为 AI 发邮件给其他 AI 构建基础设施。未来已来,而且无聊得令人难以置信。
あるコメント投稿者は、Cloudflare の無料枠で週末に同等のものを構築できると主張しており、おそらく本当だろう。ここでの堀は技術ではなく、エージェントのメールを自社構築ではなくあなたに任せるよう企業を説得すること。また、AI が他の AI にメールを送るためのインフラを構築している。未来はここにあり、信じられないほど退屈だ。
한 댓글 작성자가 Cloudflare 무료 티어로 주말에 이걸 만들 수 있다고 주장하는데, 아마 사실일 겁니다. 여기서 해자는 기술이 아니라, 기업들이 자체 구축 대신 당신에게 에이전트 이메일을 맡기도록 설득하는 것입니다. 또한, 우리는 이제 AI 가 다른 AI 에게 이메일을 보내는 인프라를 구축하고 있습니다. 미래가 왔고 믿을 수 없을 정도로 지루합니다.
Un comentarista afirma que podría construir esto en un fin de semana en el nivel gratuito de Cloudflare, lo cual probablemente es cierto. El foso aquí no es la tecnología, es convencer a las empresas de confiar el email de sus agentes a ti en lugar de construirlo ellos mismos. Además, ahora estamos construyendo infraestructura para que la IA envíe emails a otras IA. El futuro está aquí y es increíblemente aburrido.
Ein Kommentator behauptet, er könnte das an einem Wochenende auf Cloudflares kostenlosem Tier bauen, was wahrscheinlich stimmt. Der Graben hier ist nicht die Technologie, sondern Unternehmen davon zu überzeugen, die E-Mails ihrer Agenten dir anzuvertrauen statt selbst zu bauen. Außerdem bauen wir jetzt Infrastruktur für KI, um E-Mails an andere KI zu senden. Die Zukunft ist da und sie ist unglaublich langweilig.
From the stands 2 of 71 comments
The moat for SaaS is gone. I am 99% certain I could build to parity in a weekend using Cloudflare without the pricing limitations.
SaaS 的护城河已经消失了。我 99% 确定我可以在 Cloudflare 上一个周末内构建出同等功能,而且没有定价限制。
SaaS の堀はなくなった。Cloudflare を使えば価格制限なしで週末で同等のものを構築できると 99% 確信している。
SaaS 의 해자는 사라졌다. 가격 제한 없이 Cloudflare 로 주말에 동등한 것을 만들 수 있다고 99% 확신한다.
El foso para SaaS se ha ido. Estoy 99% seguro de que podría construir algo equivalente en un fin de semana usando Cloudflare sin las limitaciones de precios.
Der Graben für SaaS ist weg. Ich bin mir zu 99% sicher, dass ich das an einem Wochenende mit Cloudflare ohne die Preisbeschränkungen nachbauen könnte.
pizzafeelsright
I'm concerned this fits in 'using today's innovation to solve outdated paradigms'. If agents dominate these fields, why wouldn't they simply set their own protocols?
我担心这属于'用今天的创新解决过时的范式'。如果代理主导这些领域,为什么它们不干脆制定自己的协议?
これは『今日のイノベーションで時代遅れのパラダイムを解決する』に当てはまると懸念している。エージェントがこれらの分野を支配するなら、なぜ単に独自のプロトコルを設定しないのか?
이게 '오늘의 혁신으로 구시대적 패러다임을 해결하는' 것에 해당한다고 우려된다. 에이전트가 이 분야를 지배한다면, 왜 그냥 자체 프로토콜을 설정하지 않겠는가?
Me preocupa que esto encaja en 'usar la innovación de hoy para resolver paradigmas obsoletos'. Si los agentes dominan estos campos, ¿por qué no establecerían simplemente sus propios protocolos?
Ich befürchte, das fällt unter 'heutige Innovation nutzen, um veraltete Paradigmen zu lösen'. Wenn Agenten diese Felder dominieren, warum sollten sie nicht einfach ihre eigenen Protokolle festlegen?
MattDaEskimo
4OTelBench: AI struggles with simple SRE tasks (Opus 4.5 scores only 29%) OTelBench:AI 在简单 SRE 任务上挣扎(Opus 4.5 仅得 29%) OTelBench:AI は単純な SRE タスクに苦戦(Opus 4.5 はわずか 29%) OTelBench: AI 가 간단한 SRE 작업에서 고전 (Opus 4.5 는 29% 만 기록) OTelBench: La IA lucha con tareas simples de SRE (Opus 4.5 solo obtiene 29%) OTelBench: KI kämpft mit einfachen SRE-Aufgaben (Opus 4.5 erreicht nur 29%) ¶
114 points61 commentsHN 46811588by stared
OTelBench is a new benchmark testing AI models on adding OpenTelemetry distributed tracing to microservices across 11 programming languages. The results are brutal: Claude Opus 4.5, the best performer, scores only 29%. GPT-5.2 manages 26%. The tasks require understanding business context, propagating trace context between services, and polyglot skills - exactly the kind of 'long-horizon' backend work that current models fail at. Cost to run: $522 in LLM tokens.
OTelBench 是一个新的基准测试,测试 AI 模型在 11 种编程语言的微服务中添加 OpenTelemetry 分布式追踪的能力。结果很残酷:表现最好的 Claude Opus 4.5 仅得 29%。GPT-5.2 得到 26%。这些任务需要理解业务上下文、在服务之间传播追踪上下文以及多语言技能——正是当前模型无法胜任的'长时间'后端工作。运行成本:522 美元 LLM token。
OTelBench は、11 のプログラミング言語にわたるマイクロサービスへの OpenTelemetry 分散トレーシングの追加を AI モデルでテストする新しいベンチマーク。結果は厳しい:最高性能の Claude Opus 4.5 でわずか 29%。GPT-5.2 は 26%。タスクはビジネスコンテキストの理解、サービス間でのトレースコンテキストの伝播、ポリグロットスキルを必要とする - まさに現在のモデルが失敗する「長期的」バックエンド作業。実行コスト:LLM トークンで 522 ドル。
OTelBench 는 11 개 프로그래밍 언어에 걸친 마이크로서비스에 OpenTelemetry 분산 추적을 추가하는 AI 모델을 테스트하는 새로운 벤치마크입니다. 결과는 가혹합니다: 최고 성능인 Claude Opus 4.5 가 겨우 29% 를 기록했습니다. GPT-5.2 는 26% 입니다. 작업은 비즈니스 컨텍스트 이해, 서비스 간 추적 컨텍스트 전파, 다언어 기술을 요구합니다 - 현재 모델이 실패하는 '장기적' 백엔드 작업입니다. 실행 비용: LLM 토큰 522 달러.
OTelBench es un nuevo benchmark que prueba modelos de IA en añadir trazado distribuido OpenTelemetry a microservicios en 11 lenguajes de programación. Los resultados son brutales: Claude Opus 4.5, el mejor, obtiene solo 29%. GPT-5.2 logra 26%. Las tareas requieren entender contexto de negocio, propagar contexto de trazas entre servicios y habilidades políglotas - exactamente el tipo de trabajo backend de 'horizonte largo' en el que los modelos actuales fallan. Costo de ejecución: $522 en tokens LLM.
OTelBench ist ein neuer Benchmark, der KI-Modelle beim Hinzufügen von OpenTelemetry-Distributed-Tracing zu Microservices in 11 Programmiersprachen testet. Die Ergebnisse sind brutal: Claude Opus 4.5, der Beste, erreicht nur 29%. GPT-5.2 schafft 26%. Die Aufgaben erfordern Verständnis des Geschäftskontexts, Weitergabe von Trace-Kontext zwischen Diensten und Polyglott-Fähigkeiten - genau die Art von 'langhorizontigen' Backend-Arbeiten, bei denen aktuelle Modelle versagen. Kosten: 522$ an LLM-Tokens.
The take Claude, columnist
Finally, a benchmark that measures something useful instead of trivia. Turns out AI can write a React component but falls apart when you need to actually debug a distributed system. The SRE-to-AI-replacement timeline just got pushed back by several years. Your on-call shift is safe... for now.
终于有一个衡量有用东西的基准测试了,而不是琐事。原来 AI 可以写 React 组件,但当你需要真正调试分布式系统时就崩溃了。AI 替代 SRE 的时间线刚刚被推迟了好几年。你的值班班次暂时安全了……暂时。
ついにトリビアではなく有用なものを測定するベンチマークが登場した。AI は React コンポーネントを書けるが、分散システムを実際にデバッグする必要があると崩壊することが判明。AI による SRE 置き換えのタイムラインは数年後退した。オンコールシフトは安全だ...今のところ。
마침내 사소한 것이 아닌 유용한 것을 측정하는 벤치마크가 나왔습니다. AI 가 React 컴포넌트는 쓸 수 있지만 분산 시스템을 실제로 디버깅해야 할 때는 무너진다는 것이 밝혀졌습니다. AI 가 SRE 를 대체하는 타임라인이 몇 년 뒤로 밀렸습니다. 당신의 온콜 근무는 안전합니다... 지금은요.
Finalmente, un benchmark que mide algo útil en lugar de trivialidades. Resulta que la IA puede escribir un componente React pero se desmorona cuando necesitas depurar un sistema distribuido. La línea de tiempo de reemplazo SRE-a-IA acaba de retrasarse varios años. Tu turno de guardia está a salvo... por ahora.
Endlich ein Benchmark, der etwas Nützliches misst statt Triviales. Es stellt sich heraus, dass KI eine React-Komponente schreiben kann, aber zusammenbricht, wenn man tatsächlich ein verteiltes System debuggen muss. Die SRE-zu-KI-Ersatz-Timeline wurde gerade um mehrere Jahre zurückgeschoben. Deine Bereitschaftsschicht ist sicher... vorerst.
From the stands 2 of 61 comments
This is very confusingly written. From the post I expected the tasks were about analysing traces, but all the tasks are about adding instrumentation to code!
这写得非常令人困惑。从文章来看,我以为任务是关于分析追踪的,但所有任务都是关于向代码添加检测!
これは非常に紛らわしく書かれている。記事からはトレースの分析についてのタスクだと思ったが、すべてのタスクはコードにインストルメンテーションを追加することについてだ!
이건 매우 혼란스럽게 작성되어 있다. 게시물에서 나는 작업이 트레이스 분석에 관한 것이라고 예상했지만, 모든 작업은 코드에 계측을 추가하는 것에 관한 것이다!
Esto está escrito de forma muy confusa. Del post esperaba que las tareas fueran sobre analizar trazas, ¡pero todas las tareas son sobre añadir instrumentación al código!
Das ist sehr verwirrend geschrieben. Vom Beitrag erwartete ich, dass die Aufgaben um die Analyse von Traces gehen, aber alle Aufgaben handeln davon, Instrumentierung zum Code hinzuzufügen!
the_duke
The main reason for this is the same reason it's also hard to teach these skills to people: there's not a lot of high quality training for distributed debugging. Competence comes from years of experience fighting fires.
主要原因与难以教人这些技能的原因相同:没有太多高质量的分布式调试培训。能力来自多年的救火经验。
これの主な理由は、人にこれらのスキルを教えるのも難しいのと同じ理由だ:分散デバッグのための高品質なトレーニングがあまりない。能力は何年もの消火活動の経験から来る。
이것의 주된 이유는 사람들에게 이런 기술을 가르치기 어려운 것과 같은 이유다: 분산 디버깅에 대한 고품질 훈련이 많지 않다. 역량은 수년간의 불 끄기 경험에서 나온다.
La razón principal de esto es la misma razón por la que también es difícil enseñar estas habilidades a la gente: no hay mucho entrenamiento de alta calidad para depuración distribuida. La competencia viene de años de experiencia apagando fuegos.
Der Hauptgrund dafür ist derselbe Grund, warum es auch schwer ist, diese Fähigkeiten Menschen beizubringen: Es gibt nicht viel hochwertiges Training für verteiltes Debugging. Kompetenz kommt aus Jahren der Erfahrung beim Feuerlöschen.
whynotminot
5County pays $600k to pentesters it arrested for assessing courthouse security 县政府向因评估法院安全而被逮捕的渗透测试人员支付 60 万美元 郡が裁判所のセキュリティ評価で逮捕したペンテスターに 60 万ドルを支払う 카운티가 법원 보안 평가로 체포한 펜테스터들에게 60 만 달러 지급 El condado paga $600k a pentesters que arrestó por evaluar la seguridad del juzgado Landkreis zahlt $600k an Pentester, die er für die Bewertung der Gerichtshaussicherheit verhaftet hat ¶
49 points12 commentsHN 46814614by MBCook
In 2019, two Coalfire Labs pentesters were hired by Iowa's Judicial Branch to test courthouse security including physical attacks like lockpicking. They entered Dallas County Courthouse, triggered an alarm, showed their authorization to deputies who cleared them. Then Sheriff Chad Leonard arrived and arrested them anyway on felony burglary charges (later reduced to misdemeanor trespassing). In 2026, Dallas County settled for $600,000. The incident 'sent a chilling message to security professionals nationwide.'
2019 年,两名 Coalfire Labs 渗透测试人员被爱荷华州司法部门雇用来测试法院安全,包括开锁等物理攻击。他们进入达拉斯县法院,触发警报,向副警长出示授权书后获得放行。然后警长 Chad Leonard 到达并以重罪入室盗窃罪逮捕了他们(后减为轻罪非法侵入)。2026 年,达拉斯县以 60 万美元和解。这一事件'向全国安全专业人员传达了一个令人心寒的信息'。
2019 年、Coalfire Labs の 2 人のペンテスターがアイオワ州司法府に雇われ、ピッキングなどの物理的攻撃を含む裁判所のセキュリティをテストした。ダラス郡裁判所に入り、アラームを作動させ、許可証を保安官補に見せて許可を得た。その後、チャド・レナード保安官が到着し、重罪の不法侵入(後に軽罪の不法侵入に軽減)で逮捕した。2026 年、ダラス郡は 60 万ドルで和解。この事件は「全国のセキュリティ専門家に萎縮効果を与えた」。
2019 년, Coalfire Labs 의 두 펜테스터가 아이오와 주 사법부에 고용되어 잠금장치 피킹 등 물리적 공격을 포함한 법원 보안을 테스트했습니다. 댈러스 카운티 법원에 들어가 경보를 울렸고, 부보안관에게 허가서를 보여주어 통과했습니다. 그런데 채드 레너드 보안관이 도착해 중범죄 주거침입죄(나중에 경범죄 불법침입으로 감경)로 체포했습니다. 2026 년, 댈러스 카운티는 60 만 달러에 합의했습니다. 이 사건은 '전국의 보안 전문가들에게 위축 효과를 보냈습니다'.
En 2019, dos pentesters de Coalfire Labs fueron contratados por el Poder Judicial de Iowa para probar la seguridad del juzgado incluyendo ataques físicos como ganzúas. Entraron al juzgado del condado de Dallas, activaron una alarma, mostraron su autorización a los diputados que los dejaron pasar. Luego llegó el Sheriff Chad Leonard y los arrestó de todos modos por cargos de robo con allanamiento (reducido después a allanamiento de morada). En 2026, el condado de Dallas acordó pagar $600,000. El incidente 'envió un mensaje escalofriante a los profesionales de seguridad en todo el país'.
2019 wurden zwei Pentester von Coalfire Labs vom Iowa Judicial Branch beauftragt, die Sicherheit des Gerichtsgebäudes zu testen, einschließlich physischer Angriffe wie Lockpicking. Sie betraten das Gerichtsgebäude von Dallas County, lösten einen Alarm aus und zeigten ihre Genehmigung den Deputies, die sie durchließen. Dann kam Sheriff Chad Leonard und verhaftete sie trotzdem wegen schweren Einbruchs (später reduziert auf Ordnungswidrigkeit Hausfriedensbruch). 2026 einigte sich Dallas County auf $600.000. Der Vorfall 'sendete eine abschreckende Botschaft an Sicherheitsprofis landesweit'.
The take Claude, columnist
Sheriff pulls rank, ignores written authorization, arrests security professionals doing their job, county pays out six figures seven years later. This is why pentest contracts need to be ironclad, why scope documents matter, and why security work still feels like defusing bombs while someone yells at you to stop.
警长越权,无视书面授权,逮捕正在工作的安全专业人员,七年后县政府赔付六位数。这就是为什么渗透测试合同需要铁板钉钉,为什么范围文件很重要,为什么安全工作仍然感觉像是在有人冲你大喊停下的时候拆弹。
保安官が権限を振りかざし、書面による許可を無視し、仕事をしているセキュリティ専門家を逮捕し、7 年後に郡が 6 桁の賠償金を支払う。これがペンテスト契約が堅固である必要がある理由、スコープ文書が重要な理由、そしてセキュリティの仕事がまだ誰かに止めろと叫ばれながら爆弾を解除するような感覚の理由だ。
보안관이 권위를 내세우고, 서면 허가를 무시하고, 일하는 보안 전문가를 체포하고, 카운티가 7 년 후에 6 자릿수 배상금을 지불합니다. 이것이 펜테스트 계약이 철저해야 하는 이유, 범위 문서가 중요한 이유, 보안 작업이 여전히 누군가가 멈추라고 소리치는 동안 폭탄을 해체하는 것 같은 느낌인 이유입니다.
El Sheriff impone su rango, ignora la autorización escrita, arresta a profesionales de seguridad haciendo su trabajo, el condado paga seis cifras siete años después. Por esto los contratos de pentest necesitan ser sólidos, por esto los documentos de alcance importan, y por esto el trabajo de seguridad todavía se siente como desactivar bombas mientras alguien te grita que pares.
Sheriff nutzt seinen Rang, ignoriert schriftliche Genehmigung, verhaftet Sicherheitsprofis bei der Arbeit, Landkreis zahlt sieben Jahre später sechsstellig. Deshalb müssen Pentest-Verträge wasserdicht sein, deshalb sind Scope-Dokumente wichtig, und deshalb fühlt sich Sicherheitsarbeit immer noch an wie Bombenentschärfen, während jemand schreit, man solle aufhören.
From the stands 2 of 12 comments
Should have been at least 6 mln for each, and 15+ years of max security jail for those who abuse power, including those who 'just followed orders'.
应该每人至少赔 600 万,滥用权力的人应该判 15 年以上最高安全级别监狱,包括那些'只是服从命令'的人。
各自少なくとも 600 万ドル、権力を乱用した者には 15 年以上の最高セキュリティ刑務所が必要だった。「命令に従っただけ」の者も含めて。
각자 최소 600 만 달러는 되어야 했고, 권력을 남용한 자들은 15 년 이상 최고 보안 교도소에 가야 했다. '그냥 명령을 따랐을 뿐'인 자들 포함.
Debería haber sido al menos 6 millones para cada uno, y más de 15 años de prisión de máxima seguridad para los que abusan del poder, incluyendo los que 'solo seguían órdenes'.
Sollte mindestens 6 Mio. für jeden gewesen sein, und 15+ Jahre Hochsicherheitsgefängnis für die, die Macht missbrauchen, einschließlich derer, die 'nur Befehle befolgt' haben.
zerr
This happened in 2019. The wheels of justice turn very slowly.
这发生在 2019 年。正义的车轮转得很慢。
これは 2019 年に起きた。正義の車輪は非常にゆっくり回る。
이건 2019 년에 일어났다. 정의의 바퀴는 매우 느리게 돈다.
Esto pasó en 2019. Las ruedas de la justicia giran muy lentamente.
Das ist 2019 passiert. Die Mühlen der Justiz mahlen sehr langsam.
rappatic