Back to News Listing

Digital Colliers Daily Briefing — June 25, 2026

Digital Colliers Daily Briefing — June 25, 2026
Digital Colliers Jun 25, 2026 8 min read

Digital Colliers Daily Briefing — June 25, 2026

The compute layer is reorganizing in public. Three developments yesterday — OpenAI's first custom inference silicon, an escalating IP fight between Anthropic and Alibaba, and Qualcomm's move on Modular — together describe an industry where vertical integration, geopolitical friction, and a credible challenge to the Nvidia/CUDA stack are converging on the same quarter. Below: what shipped, what it changes, and what to watch.

1. OpenAI's Jalapeño lands, and the inference stack splinters further

Vintage technician inspecting a silicon wafer with tweezers.

OpenAI and Broadcom unveiled Jalapeño, OpenAI's first custom inference ASIC, with engineering samples already running production workloads in lab — including a model identified as GPT-5.3-Codex-Spark. OpenAI says the chip went from initial design to tape-out in nine months, a cycle the company describes as the fastest ever achieved for high-performance ASICs, and credits its own models with accelerating parts of the design process. Broadcom CEO Hock Tan confirmed gigawatt-scale deployment beginning in 2026, with Microsoft named among the data center partners. Celestica is handling board, rack, and system integration; Broadcom's Tomahawk networking silicon ties the platform together.

OpenAI is not publishing final performance numbers yet, but says early testing shows "substantially better" performance-per-watt than current state-of-the-art accelerators. Community estimates flagged by Latent Space — a near-reticle die, roughly 216GB of HBM3E, 7.1–7.4 TB/s of bandwidth, and around 10 PFLOPS at FP4 — remain unofficial but place Jalapeño firmly in TPU-class territory.

Why it matters. OpenAI is now designing every layer it consumes: models, kernels, serving systems, networking, and silicon. As TechCrunch notes, training is likely to remain on Nvidia hardware for now, but inference is where unit economics decide whether agentic products are viable at scale. A 10–30% improvement in inference cost-per-token, compounded across ChatGPT, Codex, and API traffic, changes the company's gross margin profile materially.

Who is affected. Nvidia retains training dominance but loses a marginal buyer at the inference tier. Broadcom cements its position as the merchant-silicon partner of choice for hyperscaler-class custom chips, joining its existing Google TPU relationship. Microsoft gets a second-source inference path inside its own data centers. AMD, which had positioned MI-series parts as the inference alternative, now faces a customer-designed competitor with a captive workload.

What to watch. OpenAI says a detailed technical report is coming "in the coming months" — the realized-utilization numbers will determine whether Jalapeño is a credible Nvidia alternative or a cost-optimized adjunct. Watch for the first public benchmarks against H200 and B200 on Codex-class workloads, and for whether OpenAI offers Jalapeño capacity to API customers as a differentiated SKU.

Sources:

2. Anthropic accuses Alibaba of distilling Claude, and Alibaba shares hit a 16-month low

Vintage switchboard operator patching cables at a console.

Anthropic publicly accused Alibaba-linked operators of running roughly 25,000 fraudulent accounts that conducted an estimated 28.8 million Claude exchanges, with the alleged purpose of distilling frontier capabilities into Qwen-class systems. According to Bloomberg's Jeanny Yu, Alibaba shares fell roughly 5% in Hong Kong on the news, taking the stock to a 16-month low and a year-to-date decline of 33%. Xiaomi and Baidu fell more than 3% in sympathy.

The dispute lands on top of an already complicated regulatory situation. Anthropic's Fable 5 and Mythos models have been offline for foreign users since June 12, after the Commerce Department's Bureau of Industry and Security imposed an export-control order following NSA findings that Fable 5 guardrails could be bypassed to access Mythos-tier capabilities. According to Wired's reporting, Trump administration officials have grown frustrated with CEO Dario Amodei and are now routing meetings through cofounder Tom Brown and policy chief Sarah Heck. A bipartisan letter from Representatives Liccardo, Obernolte, Franklin, and Lieu has demanded that Commerce Secretary Howard Lutnick clarify redeployment criteria by June 26 — tomorrow.

Why it matters. Adversarial distillation — long a rumored risk in the model-access debate — now has a named accuser, a named target, and a public market reaction. If Anthropic's accounting holds up, it shifts the policy conversation from theoretical export controls on weights to enforcement against API access at scale. As Reid Hoffman told Fortune, the inconsistency in how Washington has treated Anthropic versus OpenAI is itself becoming an investor-risk category, which he characterized as "autocratic willy-nilly."

Who is affected. Alibaba faces both reputational damage and, potentially, leverage in any future trade or sanctions package; Qwen's competitive ascent — GLM-5.2 and Qwen-AgentWorld dominated open-model conversation this same week — now carries an asterisk. Anthropic gains political capital with Washington but remains in a holding pattern on Fable 5 redeployment, with concrete commercial cost. US frontier labs more broadly will revisit account-verification, rate-limiting, and KYC controls on API access.

What to watch. The Commerce Department's response to the June 26 congressional deadline; whether Anthropic publishes technical evidence of the distillation pattern; and whether other labs disclose similar account-cluster activity. Legion's separate legal challenge — arguing that hosted model access is not the legal equivalent of exporting weights — could also reshape the doctrinal basis for the entire export-control regime.

Sources:

3. Qualcomm buys Modular, importing a CUDA challenger and a compiler team

Vintage programmer with punch cards beside a mainframe tape drive.

Qualcomm announced its acquisition of Modular, the Chris Lattner-led company behind the Mojo language and the MAX inference platform. Per statements from Lattner and Modular relayed through Latent Space, Mojo's open-sourcing roadmap remains intact post-acquisition. Terms were not disclosed in the reporting available.

The deal arrives as Qualcomm is publicly pushing into data-center AI silicon for the first time, beyond its mobile and edge franchise. Modular brings the kernel-level compiler stack, the Python-compatible language layer, and — most importantly — a team that has spent four years building a portable runtime explicitly designed to run on non-Nvidia hardware without paying the CUDA tax.

Why it matters. Until now, the most credible CUDA alternatives have been hyperscaler-internal: Google's XLA/Pallas stack for TPUs, AWS Neuron for Trainium, and Microsoft's emerging Maia toolchain. Modular was the leading independent attempt to build a merchant-grade portable inference compiler. Folding it into Qualcomm — which has its own silicon roadmap and customer relationships across edge and increasingly cloud — turns it into a vertically integrated alternative stack rather than a neutral abstraction layer.

Who is affected. Nvidia, whose CUDA moat is the single most-discussed structural advantage in the industry, faces a compiler competitor with a real silicon backer. AMD's ROCm strategy, which has struggled to attract independent tooling, loses a potential ally. Customers of Modular's MAX platform — including teams using it to serve open models on heterogeneous hardware — will need clarity on continued multi-vendor support. Coming the same day as Jalapeño, the announcement reinforces the broader pattern Latent Space flagged: "hyperscaler-style inference silicon is now table stakes," and the software layer underneath it is fragmenting accordingly.

What to watch. Whether Qualcomm preserves Modular's hardware-neutrality in practice, or steers MAX toward Qualcomm silicon over time. Mojo's open-source cadence post-close will be the leading indicator. Also watch for customer reactions from companies that adopted MAX precisely because it was vendor-agnostic, and for whether Qualcomm uses the acquisition as the software anchor for a credible data-center accelerator launch.

Sources:


Taken together, Wednesday's news describes the same structural shift from three angles: who designs the chips, who controls the software that runs on them, and who is allowed to access the models trained on them. OpenAI is answering the first question by building its own silicon; Qualcomm is answering the second by buying the compiler team most likely to break CUDA's grip; and Washington, Anthropic, and Alibaba are litigating the third in public, with a Commerce Department deadline arriving tomorrow. The merchant-GPU era is not ending, but the assumption that it was the only era is.

Related Posts