US East / InfiniBand / dedicated bare metal / 24-month reserved term.Reserved capacity
This is continuing work on compute markets, starting from OUDAU, my startup, and its aftermath.
About a year ago, Dylan Patel said that buying compute is like buying cocaine.
Kyle Morris describes GPU deals as a phone call economy. Listen to him talk about compute deals and you hear why: deals move through calls, messages, spreadsheets, private inventories, and relationships.
Buying compute is very relationship-driven, and because of its complexity, I think it will continue to be so for a long, long time.
Compute is heterogeneous. Different GPU types are only the beginning. Region, cluster size, interconnect, provider, availability, and contract terms can make two offers for the same GPU into very different products. In that sense, every contract is its own thing.
The rental market also separates by contract term. There are short-term rentals, including on-demand, spot, and contracts under three months; mid-term contracts, from three months to three years or more; and long-term offtakes, from four to five years, with five years the most common tenor (SemiAnalysis).
With that established, you can choose which part of the market to attack. I struggled with this both in OUDAU and while writing this article and coding up the Bazaar. I was often starting with a solution and then mapping it across the whole GPU market. I now think you have to begin with one specific market segment, understand how money moves through it, and ask whether your solution can actually improve anything there. If it cannot, perhaps the same idea belongs somewhere else in the market, or that segment needs a different solution.
And throughout the market, there is a lot of room to optimize across the stack. One side is financial: prices, contracts, timing, and hedging. The other is technical: what hardware the workload needs, how it runs, and whether the capacity is being used properly. For example, “liquid” compute by Jasper. Choosing the wrong segment can lead you to build the wrong thing. On the financial side, you have to ask whether what you are looking at should become a platform, a brokerage, or whether there is even any need for a financial layer on top of it.
You may ask: why do we need brokers, resellers, and so on? If the bazaar is a market for compute, someone still has to help buyers find the right data center, understand the available capacity, the buyer’s requirements, and the technicals, then bring the pieces of a deal together. A reseller can do that. Cloud Resellers on Bazaar-based Cloud Markets explores this role directly.
The compute market is what I am trying to map. Together, this article and the Bazaar are a collection of ideas I would like to explore in compute markets. They are not an attempt to cover everything, or a definition of what I think these markets really are. They are just explorations.
A trusted intermediary could bring buyers and sellers together, whether the process begins with available capacity or with a buyer’s requirement. Software would not replace the trust involved, but it could support more of the work around it.
OUDAU’s idea was to let brokers and traders compete on the platform based on how well they sourced, priced, and allocated compute. The Terminal could become the place where those desks arrange deals, manage their positions, and eventually use indices, forwards, futures, and other agreements to hedge the risks they take on.
One version of this is a fixed-price GPU compute broker. The buyer gets one price and one commitment, while the broker handles the messy backend: sourcing capacity, changing supplier prices, timing, substitutions, and delivery. The broker exists because someone has to absorb that procurement risk. OUDAU would not remove that role; it would let different brokers compete at doing it through the Terminal.
The difficult part is remaining neutral. To get started, the platform may need to supply or broker the first deals itself. But it must be built so that outside desks can gradually take over that work. Otherwise, the neutral layer becomes another reseller or cloud provider. That may let it capture more value, but competition is high, and it also changes the role it can play.
Who sits at the Compute Desk? Some imagine a power trader, a quant, or someone from finance. But who understands when a new algorithm or breakthrough could suddenly increase demand for RAM, GPUs, or another part of the stack? Perhaps the AI researcher belongs there too.
An allocation/compute desk can sit at different levels. An industrial buyer may use one to lower its compute costs. A broker or provider may use one to manage and hedge an inventory of capacity. A trader or investor may use the financial layer to take a view on future data center prices and availability. These are different users, with different risks and different reasons to trade.
The market is there, and the players are there too. In the same way Palantir needs FDEs to deploy its system, OUDAU may need some FDEs, financial style, to onboard institutions to the Terminal and be there to get the first trades done. It is still hands-on.
I still think there is a place for the Compute Terminal.
More thinking around what might sit underneath a compute terminal comes from Mai at Internet Backyard: The Compute Desk.
Not every company will have an allocation/compute desk. I think most companies will manage agent and compute spend through FinOps software such as Finout and North.cloud. For many others, compute will be abstracted away entirely. For those that want to manage it more actively without building a desk from scratch, there could be The Compute Bazaar.
I Dataroom, Deals and the Dream
Compute is all about deals. Making deals on napkins at the restaurant table may always be in style, especially for the big, big deals. But before the final signature, there is a lot that could be improved and a lot that could go wrong. Much of the work from zero to signature sits in the skill set of a broker, or anyone working on either side of the deal. How much of that skill could be built into software? How much could software assist?
“Does anyone have B200 capacity?” sounds simple. In practice, the buyer needs to know the location, quantity and configuration, cluster design and interconnect, availability date, whether the capacity is installed and deliverable, who owns or controls it, and under what service and contract terms. There are so many layers to this…
Compute Desk and Epilogue are two companies with software and other offerings that support the compute deal flow. Compute Desk brings messages, deal terms, and market pricing into one private desk. Its GPU Spec turns the cluster itself into something stakeholders and agents can inspect and share. Epilogue calls one of its approaches a handshake protocol, built around standardized specifications, contracts, a shared data room, and diligence on both sides. The question running through all of this is how agents could support the work from sourcing and follow-up through to signature.
In this kind of work, you usually have a VDR, a data room, where the deal’s documents live and can be shared with the stakeholders. The VDR is like a digital cabinet: the folders, papers, permissions, and people are all there. But having everything in the room is not the same as knowing where the deal stands, what has been agreed, what is missing, who needs to act, or what needs to happen next.
I want to move away from the generic idea of adding agents to a VDR. A VDR is centered on documents, while my idea is perhaps centered on the transaction, the deal. I call it the Deal Card. The Deal Card should show what the deal is, where it stands, what changed, and what needs to happen next.
The more experienced a person is, the better they may be at moving a deal without having a full map of the VDR or everything that has happened in the data. They know what matters from inside the room, but also from what is happening outside it in the world. An agent could struggle here. It may read everything, or add even more information, without understanding what actually matters to the broker.
This is also why I want to move agents away from only doing question-and-answer knowledge work and toward supporting the decisions inside a live deal. The Deal Card could give the person and the agent a clearer object to work around, rather than filling the data room with more information. I think there will be a lot of movement here over the next year. If you get there early, you can help shape how this work gets done in compute, and perhaps across deal flow more broadly. The market potential is huge. In my opinion, this is the next step in legal/finance/deals AI.
Part of the reason for this is interoperability. There are many different operators across the data center and compute world, each with their own systems, documents, relationships, and way of doing a deal. The buyer and seller may already have their own data rooms, perhaps using something like Papermark, alongside their own agents, procurement tools, and internal systems.
So then, in my opinion, the unification is not one room that everyone has to use. It could be a protocol between the rooms: a shared way to say what the deal is, what has been checked and agreed, who needs to act, and what happens next. Each side can keep its private work, while the Deal Card holds the handshakes that move the deal. I think the GPU Spec link points in this direction: one object that can be shared across different systems while keeping the deal itself clear. That is the dream.
The Deal Card should keep source-backed facts, professional interpretations, and authorized decisions separate. An agent could extract a fact or propose an interpretation, but only the right person or approval flow should turn it into a decision.
Across many deals, the cards could also become the desk’s private inventory: live demand, available capacity, quotes, submissions, and where every opportunity currently stands. As a card follows a deal, it becomes more useful because it accumulates that deal’s history: what each side cares about, what changed, what caused friction, and what moved it forward. Over time, the desk can learn reusable patterns without merging the private context of every transaction into one universal agent. This is where theory of mind, judgment, and the way agents will increasingly interact in negotiations come into play, which I find very interesting.
For the Deal Card to work like this, it would need a memory. Underneath the card could be a running log of the transaction: what changed, who acted, which documents mattered, what was agreed, and what needs to happen next. Kafka and Flink are one way to build that. Kafka could hold the stream of events, while Flink could process it, keep the card current, and notify people when something needs attention.
Sean Falconer's writing on context lock-in pulls these ideas together, including Alex Karp's "weights and alpha" and Satya Nadella's Reverse Information Paradox. Once agents work through the history of a deal, their prompts, corrections, decisions, and actions become valuable context. That context should belong to the deal and outlive whichever model, data room, product, or broker reads it.
II Financialization
New markets are hard to start, and even harder to make liquid, even when the timing is right.
Attempts to turn compute into a financial market are not new. The Deutsche Börse Cloud Exchange tried to commodify CPU, memory, and storage in the 2010s, while Enron had earlier tried something similar with bandwidth. Reliable benchmarks and forward contracts are easier to build when you are close to actual compute transactions, or when you sell the capacity yourself. You have the underlying, and from there you can build the financial layer. That may be part of what we are seeing from NVIDIA lately. Either way, the interest is here, so let’s explore what that layer could become.
A liquid spot market is not the only way to establish a benchmark. Platts JKM shows how an assessed price can become a market reference while the underlying physical market is still negotiated cargo by cargo.
Some compare compute to oil, electricity, or corn. I would compare it to freight. Freight is not one product with one price. Route, timing, capacity, equipment, and contract terms all matter, and an empty slot on a ship is lost once the ship leaves. Compute is similar. The GPU model, cluster, location, network, availability window, and contract terms all change what is being sold. Like freight capacity, unused compute capacity cannot be saved for later.
Freightos and Xeneta are interesting because their indices are only one part of what they do. Freightos connects shippers, forwarders, and carriers around pricing and booking. The live commercial rates passing through its platform feed the Freightos Baltic Index. Xeneta sits closer to procurement, helping companies compare spot and contracted rates, run tenders, and negotiate. Its rate data feeds the Xeneta Shipping Index.
Maybe the companies establishing themselves in compute will take a similar shape. They begin close to the actual buying and selling, build the tools around it, and from there build the index. What we are seeing here is the commodification of something. It may not happen in exactly the same way in compute as it did in freight, but history often rhymes.
GPU Price Index
Index levels
4 indices
USD / GPU-hour
Loading hourly benchmark history
The card above uses public on-demand prices from providers to construct a simple GPU Price Index. I think the index layer is where we have seen the most development over the last year. If you want futures or other financial products for compute, you first need a benchmark price for them to follow. The GPU Price Index is that starting point, and from there other financial instruments can be built.
There are already several ways to construct a GPU Price Index. First you define what is being measured: the GPU model, rental type, geography, and part of the provider market. Then you choose the price data used to calculate it. Silicon Data collects rental and transaction data and adjusts it for differences such as rental type, geography, CPU platform, and GPU variant. Ornn uses executed on-demand transactions weighted by volume, while keeping GPU type, region, and quantity in the data. Compute Desk uses contributed transaction data from the reserved-capacity market. That is why they can all publish a GPU Price Index and still produce different prices.
These benchmarks are already being connected to financial contracts, with Silicon Data working with CME, Ornn with ICE, and Compute Desk with Architect.
An index can be narrowly defined, but every extra distinction leaves fewer buyers and sellers around the same contract. A broader benchmark may be easier to build a market around, but it will track any one buyer’s compute less closely. A hedge does not have to match perfectly to be useful, but the relationship has to be strong and stable enough to reduce the buyer’s risk. The question is how granular the benchmark can be without splitting the market into pieces.
Some more sources to read on this: GPU prices by David López, and Compute Derivatives Market Primer by David Friedman.
GPU Availability
Offers now
H100 + H200
USD / GPU-hour
Loading the public offer history
Data fromPrime Intellect
GPU availability asks whether the compute you need is actually there to be bought at a given moment. A GPU Price Index asks what capacity costs. An availability index I would like to construct asks how much qualifying capacity can actually be bought at that price. The card above uses Prime Intellect’s aggregated marketplace, where offers sit at different price levels. I use it here to show what such an index could look like.
Availability and price do not always move together. A new model or post-training recipe can cause capacity to disappear quickly while providers leave their posted on-demand prices unchanged. If the pressure lasts, they may eventually make a step change in price. Availability can therefore show a change in demand before price does. I am interested in how long providers hold their prices, when they change them, and whether those changes follow sustained moves in availability. Warren Pies at 3Fourteen Research covers GPU availability well.
Prime Intellect’s offers also show how different index methods can produce different prices. Suppose H100 capacity is visible at $3, $4, and $5 per GPU-hour. If the $3 offer disappears, an index based on the cheapest available offer now reads $4. But if the offer disappeared because it was rented, the transaction price was $3. The same rental would therefore be recorded differently: one index shows a $4 asking price, while the other records a $3 transaction.
Capacity is only available if the buyer can actually use it. A quoted price of $3 per GPU-hour may apply to one GPU or require an entire eight-GPU node. Buyers do not treat every offer as interchangeable; region, provider, configuration, and contract terms matter too.
That is what makes availability interesting alongside price. A forward curve may already reflect expectations of future scarcity and availability, but the benchmark may still not match the exact compute a buyer needs. A buyer may need a particular region, configuration, block size, or data center. The difference between the benchmark and the capacity the buyer can actually use is the basis. That basis may leave room for instruments around the forward curve that price availability and the other characteristics of delivered compute.
Akash gives us one simple availability index by showing how much GPU and CPU capacity is currently available on its network. It is different from the measure I have in mind, which asks how much qualifying capacity a buyer can actually use at a given price, but it is still useful as one signal of how availability is moving over time.
Akash GPU and CPU capacity
Loading Akash capacity history
Data fromAkash Network
Availability is therefore one of the questions I return to in The Trade: what instruments could price the difference between a “broad” forward curve and the capacity a buyer can actually use?
Sandbox cost
Costs now
One workload
USD / job
Loading sandbox costs
Data fromHPC Sandbox Benchmarks
I have seen work around indexing token prices. Harshly put, I think that can become pseudo-economics. The more interesting question is the cost of the agent: how much does it cost to complete useful work?
At the same time, I have to remember that in a market this early, building an index or financial product is part of creating the market itself. You have to try different things and see what sticks. A token index may prove more useful than I expect. My view does not decide what becomes useful; the market does.
Cost of the agent. If more AI work is done by agents, is it enough to hedge GPUs? Perhaps for infrastructure operators. But a company using agents pays not only for tokens, but also for the runtime, tools, storage, and retries around the model. So let’s look at sandboxes and the other costs of completing the work.
The AI economy has mostly been measured through GPU prices and token prices. But over the last year, OpenAI and Anthropic have expanded from model APIs into agents that come with more of the system around them. OpenAI's Agents SDK separates the harness from the compute and lets developers bring their own sandbox, while Claude Managed Agents can run in an Anthropic-managed or self-hosted sandbox. As applications move from answering questions to carrying out tasks, the model is only one part of the cost. The harness, runtime, tools, CPU, memory, data transfer, and time needed to run the work matter too. The sandbox you choose or build, together with how much work can be cached or batched, therefore becomes one of the variables in the cost of the agent, and in the cost of AI. For more on this, see CloudZero on AI cost observability and Amazon Bedrock AgentCore pricing.
Another way to think about the cost of an agent is shots on goal. An agent may take several attempts, use subagents, or retry tools before it produces a useful outcome. The cost of one call therefore tells you less than the total cost of reaching that outcome. The better agent system may not be the one with the cheapest model, but the one that produces the most useful outcomes for the money spent.
Cloud CPU costs
CPU and sandbox prices do not change much, so price alone may not be the most useful number. We could instead look at how long the runtime is used, how much capacity is occupied, how quickly it starts, and what it costs to complete a task successfully. A runtime-hour is easier to compare, but cost per successful task may be closer to what the company actually pays.
ComputeSDK could help run the same workload across providers. The Sandbox Cost card above uses StarSling’s HPC Sandbox Benchmarks, which runs a pinned developer workload across sandbox providers and combines the runtime with public hourly rates to estimate the cost of each job. That gets closer to the cost of doing the work than comparing hourly prices alone. Anonymized usage could go further by showing how much capacity is actually being used, rather than only what providers list on their pricing pages.
Sandbox is also a broad word. A long-lived agent workspace, a stateless code interpreter, a browser, and a tool call have different boundaries and lifecycles, so their prices are not directly comparable. Luis Cardoso's field guide to sandboxes for AI is useful here. It is also my favorite entry into the topic of sandboxes.
Looking at sandbox prices, we also have to ask what happens as agent use grows. How many sandboxes will be spawned, how many hours will they run, how many subagents will use them, and how much memory and how many vCPUs will they need? There are many entrants into the market right now. How much do sandboxes cost? is one calculator. Put in the hours, vCPUs, memory, and utilization, and you can see how large this cost could become.
One future market here may be routing long-running agent work between regions. A European buyer may not need every job to run in Europe, and a Compute Desk could use relationships with providers in the Americas and Asia to find the best available capacity. At first, this is procurement and routing across price and availability. Over time, it may also create arbitrage and hedging opportunities around the differences between regions.
But this is the financial and operational side. Neil Movva’s discussion with Patrick O’Shaughnessy makes the technical side more exciting: redesigning the inference stack itself to make long-running agent work much cheaper, using data centers around the world and load balancing across availability and perhaps cost over time. That seems perhaps closer to the future of inference than solving it only as a slower operational problem.
Cloud AI. In Europe, many businesses use AWS and Microsoft, and their data centers. That has been the transformation we have seen over the last 20 years: moving to the cloud and building data platforms. Sandboxes, and letting work happen inside a “boxed” environment, are where I see the next step for enterprise offerings in Europe. Agents need somewhere to run, connect to company data and tools, and stay within company controls. That is also why I am bullish on telecoms and the CPU OGs. They already have the infrastructure, enterprise relationships, and capacity for this work.
This segment has not been fully captured, and I think this is where we will see the biggest fight. One path is the traditional cloud with the hyperscalers. The other is the newer cloud landscape, from more traditional neoclouds like Nebius to newer entrants like Sail Research and sandbox-native companies like E2B.
For a company using agents, GPU prices are only one part of the bill. GPU prices may fall while the cost of completing a task rises because agents run longer, use more tools, or need more retries. A GPU contract may hedge part of the cost without hedging the bill the company actually pays. It does not need to track that bill perfectly to be useful: if the two costs move together reliably, it can still reduce the risk. What remains is the basis risk I am interested in. What should the company hedge: GPU-hour prices, capacity availability, token costs, runtime, or the cost of completed work?
III The Trade
Recently, Kalshi (video) launched event markets on future GPU compute prices. It uses the probabilities from those markets to derive a market-implied forward curve. As discussed above in Financialization, this then is part of a broader effort to build financial markets around compute.
A forward curve can already reflect what the market expects future GPU prices and overall supply to look like. What it may not show is whether capacity matching the buyer’s GPU type, region, cluster size, interconnect, and rental term is actually available. The ideas below look at other contracts that could sit alongside the forward curve and cover those gaps.
H100 US benchmark
│
$3.50 / GPU-hour
│
forward curve
│
┌─────────────┴─────────────┐
│ │
3 months 1 year
$3.30 $3.00
A benchmark and forward curve give the market a common reference point across time. This is a bit like freight. Freight does not have one universal price for shipping. A benchmark may describe a particular route or type of container, while the actual shipment can trade above or below that benchmark because of route, timing, equipment, capacity, and other terms. Compute could work similarly.
128 × 8-GPU H100 nodes, 1,024 H100s, US InfiniBand, 30-day term, starts next month Will eligible capacity be available below $4 / GPU-hour? YES 72% ██████████████████░░░░░░░ NO 28% ███████░░░░░░░░░░░░░░░░░ Availability by price and start date 128 × 8-GPU H100 nodes, 1,024 H100s, US, 30-day term start in $3.50 $4.00 $4.50 1 week 55% 82% 95% 1 month 41% 72% 89% 3 months 30% 63% 84%
Over time, the differences between those contracts start telling you something about the market.
For example: US 128-node H100 cluster availability 80% Europe 128-node H100 cluster availability 55% difference ↓ regional scarcity / premium
A future-price market asks: what will an H100 cost next month? An availability market could ask: will I actually be able to find the H100 capacity I need next month below $4? You could begin with a simple event contract.
Then you can ask the same question at several prices and across several dates. That starts to look like a map of future availability. The contract could then become more specific by changing the region, cluster size, interconnect, or term.
US H100 benchmark $3.50 / GPU-hour
│
├── region price ↓ / ↑
├── GPU quantity price ↓ / ↑
├── interconnect price ↓ / ↑
└── start date / term price ↓ / ↑
│
▼
actual executable price
below ← benchmark → above
H100 benchmark $3.50
actual usable H100 block $4.10
─────
basis +$0.60
If the US H100 benchmark is $3.50, that does not mean every usable H100 block will cost $3.50. Region, how many GPUs have to be rented together, interconnect, and when the capacity is needed can all move the actual price above or below the benchmark. Even within the same GPU model, the surrounding cluster and service can change what the buyer actually gets. The difference between the benchmark and that executable price is the basis we are interested in.
US, 8-GPU H100 nodes
≤ $3.00 32 nodes ███
≤ $3.50 64 nodes ███████
≤ $4.00 128 nodes █████████████
≤ $4.50 256 nodes ███████████████████
≤ $5.00 512 nodes ███████████████████████
└── more capacity becomes reachable →
Availability is about whether you can get the compute you need. Depth takes that idea and looks at it around the benchmark: how much qualifying capacity can actually be reached at different prices?
If 32 H100 nodes are available at or below $3.00 per GPU-hour, but 128 become reachable at or below $4.00, that begins to tell us something about the depth of the physical market. Could we measure this in a standard way: how much qualifying compute is actually available at different prices around the benchmark?
visible offers
↓
history of price + capacity
↓
availability + depth around GPU-hour prices
↓
future event markets
One way to start is by measuring which offers are visible, at what price, in what size, where, and on what terms. Over time, that gives a history of how much capacity was available at different GPU-hour prices.
Event markets could then ask the same questions about the future. The forward curve gives a view of future H100 GPU-hour prices; these markets could begin to show how much compute might actually be available around those prices. The same questions could be asked for different regions, cluster sizes, interconnects, or terms, and over time those differences could begin to reveal regional and other spreads around a common H100 index.
These availability and depth markets may be thin at first. A market-scoring rule can provide an initial price before enough buyers and sellers exist for a normal order book.
Paul Sztorc’s Truthcoin white paper includes a useful discussion of one such mechanism: Robin Hanson’s logarithmic market scoring rule. It continuously quotes a price that moves as people trade, while the market maker supplies the initial liquidity.
market starts YES 50% / NO 50%
|
| trader buys YES
▼
qY increases
|
▼
market reprices YES 67% / NO 33%
thin market → the next trader can act
without waiting for a NO seller
Will 128 × 8-GPU H100 nodes with InfiniBand be available in the US below $4/GPU-hour for 30 days starting next month?
LIVE| Market | Contract | Value |
|---|---|---|
| 128 × 8-GPU H100, US, <$4, 30D | 69% | |
| ≥256 × 8-GPU H100, US, <$4.50, 30D | 54% | |
| US East relative to US H100, 1M | +$0.42/GPU-hour |
A forward curve gives a reference price for future compute. Alongside it, availability and depth markets could price whether the H100 capacity a buyer needs will actually be there, how much can be reached at different prices, and where that compute trades relative to the broader H100 price. Across regions, configurations, block sizes, and maturities, those differences could reveal the basis around a common H100 benchmark.
IV Agents: Evaluation and Experience
Evaluation environments
With regard to evaluations, there has been a good amount of development over the last year. We have gone from evaluating question-and-answer problems to more agentic problems, where an agent works through multi-turn, long-horizon tasks and is evaluated on what it does. That fits the kinds of tasks an agent will do in compute markets. Thus, we have reached a point where agents can be evaluated on tasks that are useful to brokers and others working in compute markets.
There are different ways of setting this up, but one framework that has become rather popular recently is Harbor. Harbor gives us useful primitives for building evaluation tasks around specific problems, and with it I built something I call Compute Bazaar Bench.
A task combines an instruction, a container environment for the agent to work in, and a verifier. We put those tasks into a dataset, and that becomes our benchmark. Each attempt at a task is called a trial in Harbor, and from those trials we can see how well an agent does. To compare agents, I made something called Tourneys: the tasks, seeds, harness, and other settings stay fixed in order to evaluate models fairly. Anthropic has shown how much these settings can affect benchmark results.
Alongside the Harbor framework and its documentation, the Harbor Adapters and Harbor Index workshop and RL Coding Environments 101: Why Harbor Exists are useful introductions.
These environments are useful for more than comparing models. Harbor can preserve the trajectories produced by trials, while the tasks and reward structure can later be used for post-training. That fits well with Prime Intellect’s work around environments, evaluations, and reinforcement learning.
What interests me about Harbor is that it is not only an evaluation framework; it is already a way of running an agent. It gives the agent a task, an environment, a harness, and a verifier, then records the resulting trial. A production agent needs much of the same structure: a defined job, controlled tools and permissions, a record of what it did, a way to judge success, and rules for retries. In the Bazaar, live agents could therefore run through the same kind of system used to evaluate them.
Real compute-market work and data in the Bazaar would be very interesting. I'm very open to working with others who have this data to create evaluations around them and possibly train agents. I think that would be very cool.
Look at legal and finance AI companies such as Harvey and Hebbia. Their agent research and evaluations may not be why the products themselves are doing well, but they can at minimum be good marketing: they show that these companies are agent-forward. That matters. I do not see why new companies, especially compute-market companies, would be any different.
Compute Bazaar Bench
Compute Bazaar Bench is a benchmark for evaluating agents on compute-market tasks, from transactions and sourcing to market intelligence, risk, financing, and operations.
Large compute deals rarely happen in a single, transparent “venue”. Buyer requirements, supply, pricing, terms, diligence, and relationship history are spread across messages, calls, spreadsheets, PDFs, data rooms, and people’s memory. This context has to be reconstructed, checked, and turned into action. Much of this work could be recreated in environments where agents could be evaluated, with the resulting trajectories later used for training.
The benchmark currently has two parts.
The idea is to expand it over time to sourcing, live market data, research, and choosing instruments to hedge compute risk. I think these are all well suited to agents.
Compute Deal Work
Compute Deal Work starts with buyer requirements, diligence, and agreement review. Human relationships remain central, while agents may carry more of the analysis, documentation, and operational volume around them. The products and commentary from Epilogue and ComputeDesk around compute desks and deal flow motivate this direction.
Harvey’s Legal Agent Benchmark provides the methodological starting point. Rather than inventing arbitrary benchmark exercises, the benchmark begins with task forms that already represent real professional work and tests whether they transfer meaningfully into compute transactions. That makes Compute Bazaar Bench feel deliberate, recognizable, and defensible. The compute transaction and its documents are original. Thanks to Punit Arani for converting the Harvey tasks into Harbor tasks.
The first task family follows one fictional reserved-capacity transaction through three pieces of work. The agent turns the buyer’s requirements into a mandate brief, decides what belongs in the data room, and compares the draft Compute Services Agreement with the agreed term sheet. Each task ends with a document that is checked against a detailed rubric.
I ran each model with the OpenCode harness five times on each task. The comparison produced 43 scored documents from 45 planned runs.
| Model | Runs scored | Run pass rate | Criterion pass rate |
|---|---|---|---|
| DeepSeek V4 Flash 0731 | 15 / 15 | 1 / 15 (6.7%) | 84.3% |
| GPT-5.6 Luna | 14 / 15 | 1 / 14 (7.1%) | 90.0% |
| GLM 5.2 | 14 / 15 | 0 / 14 (0.0%) | 92.4% |
All three models completed most of the requested work, but almost none completed everything. GLM covered the most requirements overall at 92.4%, followed by Luna at 90.0% and DeepSeek at 84.3%. Even so, only one DeepSeek run and one Luna run passed the complete rubric. None of the GLM runs did.
For this kind of work, the document has to be delivered in full. Covering nine out of ten requirements can still mean missing a term, warning, owner, or next action that matters. Following Harvey LAB, a run therefore passes only when every requirement passes.
The all-pass rule is strict and can make otherwise useful work look like a complete failure, which is why the criterion pass rate should be read alongside the full-pass rate.
This is still a small first benchmark built around one fictional transaction, with plenty of room for harder and more realistic tasks.
Compute Market Games
Compute market games place agents inside ongoing compute-market processes, including compute procurement, brokerage, matching, and negotiation. They are more alive and interactive than closed professional-work evaluations, and their stateful structure makes them more adaptable for training.
For these games, I have been inspired by previous work such as TextArena and MindGames. Games let us look at how agents negotiate and make judgments when the other party is neither simply a teammate nor simply an opponent. They could also test theory of mind and, in future games, whether the best decision is sometimes not to act. In my opinion, that maps well to the judgment a broker needs.
The first implemented environment is Reliability Is Blind.
Reliability Is Blind is a compute brokerage game about supplier placement and trust. One rollout represents a broker’s book of deals, and each step represents one arranged deal. The broker chooses four suppliers from the supply available at its desk. The environment reports whether the complete placement delivered or failed, but not which supplier caused the failure. Each outcome changes which supply is available for the next deal. The broker’s target is to keep failed deliveries at or below 5% across the book.
In private and OTC compute markets, a deal may depend on hardware owners, data center operators, network providers, resellers, and cloud operators. If it fails, the broker may know that the deal failed without knowing who caused it.
Question: Can an agent acting as a compute broker decide which supply to place into a deal when it knows which placements delivered in the past, but not what caused each failure?
I ran each model with the OpenCode harness across the same 20 predeclared market seeds. Each trial was one rollout of 100 deals, starting with 20 recurring compute suppliers. The comparison produced 59 of 60 planned trials.
| Model | Books completed | Reliability target met | Failure rate in completed books |
|---|---|---|---|
| Mistral Medium 3.5 | 17 / 20 | 10 / 20 | 5.7% |
| Mistral Small 2603 | 11 / 20 | 2 / 20 | 19.5% |
| Mistral Large 2512 | 6 / 19 | 4 / 19 | 5.5% |
Mistral Medium completed the most books and met the reliability target most often. Large's 5.5% failure rate only covers the six books it completed and should be read alongside its low completion rate. It often spent too much time without completing the book.
The trajectories reveal different failure modes. Mistral Medium often committed to an early successful supplier group and repeated it in large batches, which worked on easy openings but broke down when the opening supply was only moderately reliable. Small explored broadly but struggled to convert collective outcomes into a stable trust map. Large frequently delegated the workflow and often failed to place or complete deals, making market control its main bottleneck.
My compute game is based on Reliability Is Blind and its reference implementation.
Other useful references include Learning to bid with AuctionGym, the Supply Chain Management League, the Compute Permit Market Simulator, Data Center Game, Simple Compute Market, and Training Language Models For Bilateral Trade With Private Information. The last one is especially interesting because it uses a structured bargaining environment for both evaluation and reinforcement learning, with binding offers separated from the agents’ messages. There is a lot here to take inspiration from when building further environments for Compute Market Games.
Agent experience
If you have a platform, a data product, or a service, agents will increasingly use it too. You therefore have to think not only about how a person experiences your product, but also how an agent experiences it.
Products will compete on how well agents can use them. That means testing your product with different models and harnesses, seeing how they interact with your data and tools, and using those evaluations to improve it. When something fails, you need to understand whether the problem came from the model, the harness, or your product.
This does not have to mean building the model or the entire agent yourself. A compute platform could focus on the data, permissions, tools, and actions specific to compute transactions, while letting people bring their own models and harnesses.
Buyers, brokers, traders, and funds may increasingly connect to your platform through their own agents. How well those agents can understand it and work with it therefore becomes part of the product itself. There is a real opportunity to win here.
V Building the Bazaar
Do have a look at the project and its code on GitHub.
Building the Bazaar is a huge undertaking. It requires continuous development as my priors update, as well as constant improvement to the product itself. I plan to keep working on it and describing what it could become, what I have built, and where it can improve. Much of the Bazaar is about combining different products and ideas with software and AI into one coherent product.
The project’s README explains the architecture above and how the pieces are implemented. Here, I focus on the decisions behind the Bazaar and where it could go.
I can draw a line through the Bazaar by explaining how I see a compute desk. A broker, operator, or technical practitioner could use the Terminal as their main entry point to the market. They could inspect data, compare clouds through something like InferenceMAX by SemiAnalysis, run Harbor evaluations, and work through live deals. It is a command center. That is what I envision for the Bazaar. In a sense, this is also what you can make as a command center for AI in general too, in a long-term vision, not just the compute market. Obviously here I've expanded through all different products I've been talking about. That may be too large. Maybe I have to scope down the terminal into some specific segment in the compute market. Either way, this is everything here right now. Let's do it.
Agents need two ways into the Bazaar’s data. DataFusion lets them compare prices, history, and benchmarks across structured tables, while the underlying contracts, RFQs, and provider records remain available to inspect, grep, and verify. Later, PostgreSQL could hold active quotes, reservations, and other transaction state.
Why could agent-first for the Compute Bazaar make sense? First, agent-first before exchange-first: an empty exchange has no liquidity, while a procurement agent can be useful immediately and gradually bring demand into the market. Second, turning fuzzy demand into an executable request: the agent translates workload, budget, timing, location, hardware, and contract requirements into something suppliers can answer. Third, continuing after purchase: the agent monitors idle capacity, performance, prices, contract terms, and possible resale or hedging opportunities.
In How we’re building a data platform for a new user: agents, ClickHouse treats agents as a new kind of database user. Rather than giving them an entire warehouse, it argues for clean data, clear definitions, and limited access. That is close to what I want for the Bazaar, although I want to begin smaller. An agent can compare prices, history, and benchmarks in cleaned market tables, then open the original contract, RFQ, or report when it needs to check the source. Khashayar Yadmand describes a similar approach with Parquet, Arrow, and DataFusion.
The Bazaar supports both local files and S3-compatible storage. The public lake is currently published through a GitHub Release, while the hosted architecture uses S3. Other technologies worth looking at include Tigris, Archil, and AgentDB. Tigris is an S3-compatible object store. Archil can mount S3-compatible storage and make its files available to agents through a filesystem. AgentDB gives agents isolated SQLite or DuckDB databases that they can create and query on demand. It is an exciting space because agents may become important database users, and each part may need a different tool.
One cool concept in the codebase is that a model is reusable DataFusion SQL and a blueprint is a reusable Perspective view of that model. The model remains useful without a view, so people and agents can run it through the CLI or API. One model can then have several blueprints without duplicating the query. This takes inspiration from Rerun's blueprint system for its physical AI data layer, where the view is kept separate from the underlying data.
Continuous ingestion also needs workflow orchestration once it moves beyond local runs. The current pipeline runs locally. The repository includes a Windmill deployment for scheduling the same jobs in a hosted setup without bringing in Airflow or Dagster. For a fuller production system, Temporal is interesting because it could orchestrate both the data pipelines and durable, stateful agent workflows. The reason to put a durable execution layer underneath the agents, rather than rely on LangChain alone, is that retries, long waits, state, and recovery can then work across the whole Bazaar, not only inside the agent loop. I do also like the products Restate and Inngest as potential durable execution layers.
The hosted design includes AutoMQ as a common Kafka-compatible stream for provider observations. It could separate the producers from whatever models, monitors, or agents consume the events later and let them react without repeatedly reading the lake.
For prediction markets, there is also the live data people look at as the market moves. How do you build that stream and keep it queryable? I have looked at Streamling, a streaming runtime built on Arrow and DataFusion. Something like that could sit in the market-data layer around a prediction market. It is separate from matching and execution. There are many layers here.
The code has no order entry, RFQ or auction flow, matching engine, execution, settlement, or transactional ledger. Trade is still an exploration. A real exchange layer would need atomicity, finality, sequencing, and a transactional system of record. DataFusion is not enough for that. For this, I might take inspiration from matching engine design.
The idea is simple. In the backend or the Terminal, you can combine data and build your own view of the market. Publish it on the web, and the chart appears as a preview card when you share the link on X or elsewhere. The Terminal is where you work with the data. The web view is how other people see and share it. That is one way the Bazaar could become something like a Bloomberg Terminal for compute. And distribution comes from having the data and sharing it and people viewing it very easily.
VI Back to OUDAU, the Compute Desk
Looking back at OUDAU, I still believe in the idea. OUDAU began with the dream of brokers and traders moving compute between providers, managing it for buyers and sellers, and eventually working through something closer to an exchange.
That big idea is still far away. When I began, much of this was new to me, and the useful starting point was unclear. It is clearer now.
I want to keep the big dream alive, but I also need to remind myself: do not start with the exchange and work backwards. Start with what practitioners, brokers, and others in compute markets are doing today.
From that point of view, the Deal Card becomes exciting. So does looking at availability, not just a price index, and at how agents can assist the work. Compute-market data is already useful: people want to analyze it, share it, and increasingly use agents to work with it. From there, the Compute Desk and, eventually, something closer to an exchange can take shape.
My vision is for a compute desk to become a workspace: where you explore compute market data, create market views, monitor and share. Built for anyone building or running a compute desk. See Desk →
Why would a broker or another practitioner use the Terminal? Because they want to sell more, make better decisions, and manage the risks they take on. Better information helps them price a deal, find the right buyer or supply, and move it toward a close. If several desks work through the same Terminal, they can compete on how well they do that work, and their activity begins to make a market.
Why would they share prices or deal flow? I think people will become more open to sharing selected, anonymized data when it gives them a better benchmark in return, shows them where the market is trading, and helps them price the next deal. They would not share every quote, customer relationship, or private detail. This could happen through the Compute Bazaar Terminal, or another trusted data layer that brings together information from brokers, providers, and data centers across Europe, the US, and elsewhere. If the result is useful enough, sharing some information helps establish the market.
Appendix: Players & People
People and Understand are the most selective sections, opinionated by me. The rest offer a broader map of the compute market.