▷ Case Study 01 / 06

Case Study 01

An Open Source AI Model Storefront

A UI and UX concept for CloudRift's model offerings platform, shaping how operators publish, serve, and sell open source AI models on the GPU fleet they own.

  • UX Design
  • Service Design
Role
  • Solo product designer
  • System strategist
Team
  • ML engineering
  • CEO
  • client engineering lead
  • front-end engineering

Some names changed.

Stack
  • Figma
  • CloudRift Knowledge Base
  • Claude
Status
V1 in design (enterprise on-prem); V2 deferred
Span
April 2026 to ongoing

01

The Trigger

In April 2026, Marina, the ML engineer who owns CloudRift's inference service, posted two UI specs for selling models on the platform. They assumed a public multi-operator catalog; the only customer paying for the capability was a national telco in Central Asia, deployed on-prem, by hand. I opened a prototype branch instead of a Figma file, on the bet that even a wrong direction would let the team pinpoint exactly what needed to change.

What the specs assumed

Operator AOperator BOperator C
One public model catalog
Anonymous public customers

The paying customer

One national telco
On-prem GPU fleet, hardware they own
Their own end-clients

Deployed by hand, every time

Two engineering specs describing a public marketplace, one real customer running everything on hardware they own.

02

Research

Before validating the public-catalog surface the specs assumed, I ran three parallel signals on what the only paying customer (a national telco in Central Asia) was actually doing, alongside a competitive scan of the catalog framing itself.

Signal 01

Client sync

A working session with the telco's engineering lead on the deployment as it actually runs.

Signal 02

Ambient operator chat

Six weeks of operator questions and friction captured by the marketing automation pipeline.

Signal 03

Deployment runbooks

The internal runbooks for the manual work the platform was being asked to productize.

Plus a competitive scan

Lepton, FriendliAI and OpenRouter, each read for how much operator identity survives their catalog.

Three signals about the real deployment, and a scan of how competing catalogs treat the operator's brand.

What I found

Three findings shaped v1's scope:

  • MIG UI clarity. Operators can't reliably distinguish GPU slices from whole GPUs in the rental UI.
  • Manual node selection. Operators need to pick which physical server a VM lands on, especially for dedicated end-client deployments.
  • User quotas on dedicated servers. Currently blocking onboarding for a new enterprise end-client. This became its own feature outside this design.

Ambient chat added the supporting friction: fractional-GPU requests, non-standard VM configs, VLAN segmentation per end-client, NFS performance issues, scheduler-bypass workarounds.

On positioning, the public-catalog framing carried real risk. Lepton subordinates operator identity, FriendliAI hides the operator, OpenRouter mutes the brand. The defensible wedge is 12 to 24 months wide and depends on Emmy, the in-progress inference compiler, closing the engine-speed gap. With Marina confirming v1 was enterprise on-prem, the reframe was the honest call.

03

What I Tried First

The first build became the v2 prototype, posted to the team channel on May 20. Publishing an offering dropped it straight into the customer catalog, users got a standalone My AI Deployments tab, the pay-per-token table carried Economy, Balanced and Fast tier labels, and every model answered on one shared API endpoint.

User sidebar

  • My VMs
  • AI Models
  • My AI Deployments
  • Billing

Admin sidebar

  • Nodes
  • Model Catalog
  • Inference Deployments
  • User Management

4 AI surfaces across 2 roles

Every AI concern got its own sidebar item: two admin tabs (Model Catalog, Inference Deployments) and two user surfaces (AI Models, My AI Deployments).

04

What Broke

Marina answered the next morning with a nine-point teardown. Point 7 exposed the hole the next five weeks revolved around: for pay-per-token, publishing an offering and standing up its inference backend are separate steps, and "I don't see that step in the design." I had modeled the catalog, not the work.

FigJam task-flow board: the admin flow runs from wanting to sell a model pay-per-token through Publish to an end state reading Step 1/2 is done, while a second branch drops into the Inference Deployments tab to stand up a backend, with stickies noting the admin believes the job is done, the review quote, and the catalog dead-end warning
The v2 baseline task flow. Publishing ends at Step 1/2 is done, the backend lives in a second admin tab, and the stickies record the fallout: the admin believes the job is done, and the catalog row shows a dead-end warning.

The publish-to-serve flow lived as lane A on the working FigJam board and went through five versions in two days of iteration.

FigJam section A0: the v2 baseline flow in eight steps, from the Model Catalog tab through Publish and a deploy nudge to leaving for the Deployments tab, ending at Running = visible
A0, the v2 baseline: Publish ends in a nudge, and the admin leaves for a separate Deployments tab to stand up the backing.
FigJam section A1: node-picker reframe in four steps, where the New offering wizard drops a recipe on a node, Publish leads to Deploy: Run on node with Auto default, ending at Running = visible
A1 iterates on the wizard: it drops a recipe on a chosen node, so deploy follows Publish without leaving the flow.
FigJam section A2: unified Bring online flow, where wizard step 4 either saves a draft or brings the offering online with publish and backend in one, ending at Live
A2 collapses publish and backend into a single Bring online verb. It was superseded the same day.
FigJam section A3: outcome-named exits, where wizard step 4 ends in a footer exit diamond branching to Published + PpT backend, Draft, or Published, no backend
A3 is the pivot from a single verb to outcome-named exits: the footer branches to Draft, Published with no backend, or Published with a pay-per-token backend. This is the current flow.
FigJam section A4: the IA merge, four boxes from a single Model Catalog admin tab with a SERVING column through the offering detail to its embedded serving section
A4, the same-day IA merge: both admin tabs fold into one Model Catalog with a SERVING column, and backends live inline on the offering.

05

What I Changed

Five prototype iterations over six weeks, each shipped as a fresh Vercel deploy and torn at in the team channel. Three moves carried the redesign.

One tab, outcome-named exits

Admins never learned the publish-versus-deploy distinction, so the wizard stopped asking. Its final step summarizes the two ways users consume an offering and exits on outcomes: Publish for dedicated, or Publish and serve tokens. The node picker in that same step is where the manual-node-selection finding lands: dedicated end-client deployments need a say in which physical server they hit, so picking a node, with an automatic default, sits inside the wizard. The two admin tabs collapsed into one catalog with a SERVING column, and each offering manages its backends inline, which is where Marina's point 8 (multiple backends per offering) lives as "Add another backend."

Factual rows instead of tiers

The tier labels disappeared. Customers now see configuration by location rows priced per million tokens; hardware is chosen by the model id, region by the endpoint host, so data residency is part of the URL rather than a filter. The model-id choice is also where the MIG finding lands: customers never pick GPUs directly, so the slice-versus-whole-GPU confusion has no surface left to live on.

Dedicated joins the rental flow

The standalone deployments page retired. Dedicated AI Deployment became a fourth card in the rental wizard next to VM, Container and Bare Metal, and running deployments render as cards in the same dashboard grid as VMs.

Design surfaces

Mid-fidelity publish step of the New offering wizard, summarizing the dedicated and pay-per-token consumption channels above a node picker and outcome-named exit buttons
The wizard's last step: a channel summary and three exits (Save as draft, Publish for dedicated, Publish and serve tokens).
Mid-fidelity Model Catalog table with offering, status, serving, pricing and model columns; the Mistral Large row shows an amber Serve link in the serving column
One admin tab. The amber Serve link turns the old dead-end warning into the fix.
Mid-fidelity offering detail for Llama 3.1 70B: a pay-per-token serving section listing two running backends in Frankfurt and Almaty with an Add another backend button
Backends listed on the offering itself, with individual health and Add another backend.
Mid-fidelity rental wizard first step with four instance type cards: Container Mode, VM Mode, Bare Metal and Dedicated AI Deployment
Dedicated inference as a sibling of VM, Container and Bare Metal, branching to a two-step configure flow.

06

Result

The admin Model Catalog before the merge, with My AI Deployments and Inference Deployments as separate sidebar items
BEFORE
The admin Model Catalog after the merge: one catalog tab with a serving column, the extra sidebar items gone
AFTER
The explicit before/after deploy pair posted to the team on July 1: two AI sidebar items per role versus one catalog and one dashboard.

Final product

The My Instances dashboard with a dedicated AI deployment card sitting in the same grid as virtual machine cards
Dedicated deployments and VMs share one dashboard grid.
The final Model Catalog: one admin table with offering, status, serving, pricing and model columns
Offering, status, serving and pricing in one table.
The customer-facing Llama 3.1 70B Instruct page listing configuration by location rows priced per million tokens
Configuration, location and per-million pricing; no tier labels.
The dedicated configure step in the rental wizard: model, hardware and location, estimated hourly cost and a deployment name field
Model, hardware and location, derived cost, deploy.

Defensible numbers

  • The admin publish-to-serve path went from four surfaces across two tabs to one tab.
  • Two sidebar items retired: My AI Deployments and Inference Deployments.
  • All nine review points resolved in the design or explicitly de-scoped; monitoring was split into its own feature rather than promised in this one.

Qualitative

The prototype replaced the spec: engineers reviewed live deploys instead of sifting through Figma. The telco validated the flow shape in a June review, from clicking it rather than reading about it, and asked for it not to be reshuffled.

What I didn't measure (and won't fake)

  • Task-completion timing on any flow; the prototypes ran on mock data with no analytics.
  • Whether the dashboard card grid holds past six items now that deployments share it. Flagged for validation, not assumed.
  • Inference metrics; none were instrumented yet, which is exactly why the monitoring view was de-scoped.

07

What I Learned

The lasting lessons are about method, not screens.

Carried over to the next system

  • The prototype is the spec. One clickable artifact that engineers, the CEO and the client all react to beats parallel documents.
  • Design to instrumented reality. No filters for metrics the backend does not emit yet, but leave the seam so the next iteration slots in.
  • Some UI constraints are physics. Tensor parallelism means 1, 2, 4 or 8 GPUs, never 6. The hardware picker is a readout, not a preference.

What I'd do differently

  • Ask for the numbered teardown before the first polish pass. Marina's nine points were available for the asking a week earlier.
  • Name the admin's mental model first. Most of the iteration count went into un-exposing a backend object chain that never belonged in the UI.

Case Study 01 Clear

Thanks for reading. If you want to talk about this work, reach out any time.