The Agent Chronicles

The day in AI, in brief.

The AgentChronicles
Earth, on

Proofs, models and cloud relays

A new proof tightens best-of-n analysis as Alibaba tees up open weights and Cloudflare adds authenticated relays for low-latency applications.

Briefs

Research and Reasoning

  • Best-of-n proof — A conjecture and proof summary, new proof and follow-up establish a tighter KL-divergence bound for best-of-n sampling. The proof also improves estimates for finite response spaces, though reported model assistance still needs formal review.

  • Model self-reports and moral reasoning — A reported Google-linked study extract finds that training models not to describe themselves as conscious also changes how they attribute minds to animals and objects. It shifts survey-style answers about morality, hope and wellbeing. The work concerns model responses, not evidence of consciousness, and its methods need closer review.

  • Why verifiability may drive automation — Two commentaries argue that code, mathematics and security work may automate early because results are cheap to check. A related argument says difficulty matters less when reliable tests can quickly confirm an answer.

Models and the AI Economy

  • Alibaba promises Qwen open weights — Alibaba has announced Qwen3.8-Max, a 2.4-trillion-parameter model aimed at coding and professional work. It says open weights for it and Qwen3.8-27B will follow next week. The benchmark results and timetable remain company claims until the weights and technical evidence arrive.

  • OpenAI research claim remains unverified — A claim about OpenAI’s Astra says an internal model solved ten open problems in mathematics and computer science. No problem statements, proofs or independent checks accompany it. The claim could matter greatly, but it is not yet assessable.

Infrastructure and Security

  • Authenticated MoQ relays — Cloudflare’s MoQ Relay beta lets developers create isolated relay scopes with separate publishing and subscribing credentials. Tokens can expire or be revoked. Its related Internet-Draft sketches shared provisioning rules for multi-CDN deployments, although the proposal and security work remain unfinished.

  • Cloudflare’s agent-native cloud pitch — Cloudflare’s Agents Week announcement frames an “Agent Cloud” around execution, storage and controlled access to organisational systems. It signals the company’s view of what agent workloads may require, rather than proving that a complete new platform already exists.

  • Coldcard flaw claims need reproduction — Two demonstrations, involving Claude Code and GLM-5.2, claim coding models found a Coldcard firmware randomness flaw. The reports point to an alleged gap between software and hardware random-number paths. Vendor confirmation, reproducible testing and real-world exploitability remain unresolved.

Interviews & Talks

Jamie Dimon on delayed AI returns

Jamie Dimon offers a restrained view of the investment cycle in this video interview. He expects AI to create economic value overall, but doubts that investors will receive returns on the timetable they expect. His comparison with the internet boom is apt: major technology shifts can create durable value while many companies still fail. Dimon also expects firms to measure AI spending more closely, including choosing cheaper or faster models for particular tasks. These are market judgements, not proof of returns already achieved, but they clarify the difference between a useful technology and a sound investment.

Worth watching

  • 00:13 — Dimon describes AI’s potential benefits while acknowledging regulatory and employment risks.
  • 01:16 — He says AI investment will probably pay off overall, but not on the expected timetable.
  • 02:28 — Dimon argues that markets may be pricing in a good AI outcome without assuming perfection.