AI INFRASTRUCTURE FOR EVERY AI WORKLOAD
haimaker.ai is an AI infrastructure platform that runs every AI workload, from a single API for 200+ models to dedicated GPU clusters, for developers, enterprises, and governments that need AI to run where their data lives.
Founded 2025 · San Francisco · 10,000+ developers
WHAT HAIMAKER DOES
One account, one bill, one OpenAI-compatible API, from serverless tokens to bare-metal GPUs. Every workload below runs on the same platform.
Model API
Serverless inference for 200+ frontier and open-source models from 29 providers through a single endpoint. Integrate once with the OpenAI SDK and switch models by changing a string. Model API
Auto-Router
Set the model to haimaker/auto and each request goes to the optimal model for cost, speed, or capability, with routing rules learned from your own traffic. Users report up to 96% lower token spend. Auto-Router
Dedicated Endpoints
Single-tenant GPU deployments of any supported model, billed by time rather than token. Consistent latency, private capacity, and the option to pin the deployment to a region. Dedicated Endpoints
Batch Inference
Submit a file of requests and get results back asynchronously at a discount to real-time pricing. Built for evals, bulk summarization, and offline enrichment pipelines. Batch Inference
GPU Instances
On-demand GPU virtual machines with SSH access for training, experimentation, and custom serving stacks. Pay for the compute you use with no long-term commitment. GPU Instances
GPU Clusters
Multi-node GPU clusters with high-speed interconnect for distributed training and large-scale inference, sized to the job. GPU Clusters
Managed Kubernetes
GPU-backed Kubernetes clusters with the control plane managed for you, so platform teams can run their own inference and training services without operating the cluster. Managed Kubernetes
Fine-Tuning & Custom Models
Managed LoRA fine-tuning on your data and hosting for your own model weights on dedicated endpoints. On the roadmap; contact us to join early access.
Sovereign AI
The same models and API, served from GPUs physically located in your country, with data residency, audit logging, and compliance attestation for governments and regulated industries. Sovereign AI
WHAT MAKES HAIMAKER DIFFERENT
No platform markup on major models
You pay the provider's per-token rate, not a gateway surcharge on top of it. There are no seat fees and no minimum commitments; deposit $10 and you get $10 in bonus credit to start. Most gateways and inference platforms add a percentage on every request or gate features behind a subscription.
Auto-routing that lowers your bill instead of adding to it
Auto-Router picks the cheapest model that meets the quality bar for each request and learns cost-saving rules from your own traffic. Users report token-cost reductions of up to 96%. Routing adds no per-token markup and no measurable latency.
Sovereign by design, not by exception
In-country deployment is a parameter on the request, not a separate product line. Partner GPU capacity is live in the UAE, with MENA and Southeast Asia onboarding. Most inference platforms run only in US and EU regions and cannot attest to where a request was served.
One platform from API call to bare metal
Serverless tokens, dedicated endpoints, batch jobs, GPU instances, clusters, and Kubernetes all live under one account and one bill. Start on the API and promote a workload to dedicated or sovereign capacity without changing SDKs, vendors, or billing relationships.
Day-one models and transparent provenance
New models are added within days of release, often the same day. Every request logs which provider served it and what it cost, so there are no silent provider swaps behind a model name. The platform targets 99.99% uptime.
WHO USES HAIMAKER
- Developers and AI-native startups already building on the OpenAI SDK who want one key for every model and no vendor lock-in.
- Teams running coding agents such as Claude Code, Codex, Cline, opencode, Kilo Code, and OpenClaw, who route agent traffic through one endpoint and cut token spend with Auto-Router.
- Engineering teams at regulated enterprises in healthcare, finance, and other industries that need inference pinned to a known region with an auditable data-handling posture.
- Governments, telcos, and state-owned enterprises in the Middle East and Southeast Asia that must process data on GPUs inside their own borders.
- Telcos, data centers, and GPU owners with under-used capacity who partner with haimaker to offer an LLM platform to their own customers. How to list on haimaker
THE TEAM BEHIND HAIMAKER
Su Le and Wayne Pan founded haimaker in 2025. Su had spent years at NEOM and SambaNova watching governments and enterprises struggle to run AI on infrastructure they controlled; Wayne had built developer products at LinkedIn and Accord. haimaker joins the two: a developer-grade API on top of infrastructure that can run in any country. The company is headquartered in San Francisco with a team spanning engineering and go-to-market, working with infrastructure partners across the UAE, the wider Middle East, and Southeast Asia.
Su Le
CEO & Co-Founder
LinkedInGlobal AI technology leader and innovator.
Su has spent over a decade advancing AI, chips, and cloud infrastructure on a global stage. At Cisco he led Edge Compute and Industrial IoT; as Chief Strategy and Digital Officer at NEOM, the $500B cognitive city, he shaped nation-scale technology strategy and worked closely with governments worldwide.
Most recently, as Chief Growth Officer at SambaNova Systems, Su drove global expansion, built the partner ecosystem, and scaled AI chip and cloud infrastructure across the Americas and the Middle East.
Wayne Pan
CTO & Co-Founder
LinkedInProduct and engineering leader, serial entrepreneur.
Wayne co-founded a recommendation startup that was acquired by LinkedIn, where he then served as Director of Engineering. He most recently co-founded Accord, an AI-powered Revenue Execution Platform.
He builds developer products that ship fast and scale, combining deep expertise in AI-driven development with a sharp product instinct.
Ahmed Abdulla
Head of GTM & Strategic Alliances
LinkedInGlobal tech executive shaping sovereign AI ventures.
Ahmed builds and scales sovereign AI ventures with telcos and government partners. Over 15 years across telco, banking, tech, and government, he closed multi-million-dollar deals and managed $100M+ pipelines at SambaNova and FTI Delta.
Earlier he co-led FTI Delta’s technology practice and led Data & AI for banks and telcos at McKinsey, operating across the Middle East, Eastern Europe, Africa, Japan, and Southeast Asia. He holds a Ph.D. in Computer Science from Tohoku University, Japan.
HOW HAIMAKER WORKS
Three paths onto the platform, depending on who you are.
SELF-SERVE IN MINUTES
Sign up, create an API key, and point the OpenAI SDK at api.haimaker.ai/v1. No sales call, no contract. Docs live at docs.haimaker.ai; questions go to [email protected] or the community Discord.
SCOPE, PILOT, DEPLOY
Submit the contact form and it lands with the leadership team, not a queue. We scope the workload, region, and compliance requirements, run a pilot on partner capacity, then move to a dedicated or in-country deployment under a custom agreement.
HAND OVER SSH, WE RUN THE STACK
Partners with GPU capacity give haimaker access to their servers. We install and operate the inference stack, billing, and customer-facing APIs; the partner keeps the hardware and the customer relationship. Partner details
KEY FACTS
| Company Name | haimaker, Inc. (haimaker.ai) |
|---|---|
| Type | Privately held AI infrastructure platform |
| Founded | 2025 |
| Founders | Su Le (CEO), Wayne Pan (CTO) |
| Headquarters | 447 Sutter St Ste 405 #122, San Francisco, CA 94108, USA |
| Website | https://haimaker.ai |
| Core Offering | OpenAI-compatible API for 200+ models from 29 providers, with auto-routing; dedicated endpoints; batch inference; GPU instances, clusters, and managed Kubernetes; sovereign in-country AI deployment |
| Pricing | Pay-as-you-go. Per-token for inference with no platform markup on major models; per-time for dedicated endpoints and GPU compute. Free to sign up; $10 bonus credit on a first deposit of $10 or more |
| Contract Terms | No contract or minimum for self-serve API use; cancel anytime. Custom agreements for dedicated, GPU, and sovereign deployments |
| Services | Model API, Auto-Router, Dedicated Endpoints, Batch Inference, GPU Instances, GPU Clusters, Managed Kubernetes, Sovereign AI; Fine-Tuning and Custom Models on the roadmap |
| Communication | [email protected], Discord, docs.haimaker.ai; sales and partnerships via [email protected] or the contact form |
| Infrastructure Partners | GPU partners live in the UAE; MENA and Southeast Asia onboarding |
| Customers Served | 10,000+ developers and teams |
| Competitors | OpenRouter, Together AI, Fireworks AI, RunPod, Lambda, Nebius |
| Social | LinkedIn, X, Discord |
FREQUENTLY ASKED QUESTIONS
What is haimaker.ai?
haimaker.ai is an AI infrastructure platform. Developers get one OpenAI-compatible API for 200+ models from 29 providers; enterprises and governments get dedicated endpoints, GPU instances and clusters, managed Kubernetes, and in-country sovereign deployment. Everything runs under one account and one bill.
Is haimaker.ai OpenAI-compatible?
Yes. Point the OpenAI SDK you already use at https://api.haimaker.ai/v1 with a haimaker API key and your code works unchanged. Switch between GPT, Claude, Gemini, Llama, DeepSeek, Qwen, Kimi, and 200+ other models by changing the model name.
How does haimaker.ai pricing work?
Pay as you go. Inference is billed per token at transparent per-model rates with no platform markup on major models, no seats, and no minimums. Dedicated endpoints and GPU compute are billed by time. New accounts that deposit $10 or more receive $10 in bonus credit.
What is auto-routing?
Set the model to haimaker/auto and each request is routed to the optimal model for the task by cost, speed, or capability. Simple requests go to fast, inexpensive models; complex reasoning goes to frontier models. Users report token-cost reductions of up to 96%.
Does haimaker.ai store my prompts or use them for training?
No. The gateway is stateless: requests are proxied to the serving provider without being stored or used for training. Traffic is encrypted in transit and API keys are hashed at rest. You retain full ownership of your inputs and outputs.
What does "sovereign AI" mean at haimaker.ai?
Inference pinned to a specific country or compliance regime, on GPUs physically located there, with the audit trail to prove it. The same models and the same API as the public platform, plus a compliance parameter on the request. Partner GPU capacity is live in the UAE, with MENA and Southeast Asia onboarding.
Can I rent GPUs or run my own models on haimaker.ai?
Yes. haimaker offers single-tenant dedicated endpoints, GPU instances, multi-node GPU clusters, and managed Kubernetes, provisioned on partner infrastructure. Fine-tuning and custom model hosting are on the roadmap. Contact us to scope capacity and timelines.
How do I get support?
Email [email protected], join the community Discord, or read the docs at docs.haimaker.ai. Enterprise, partner, and sovereign inquiries go through the contact form straight to the leadership team.
Where is haimaker.ai based?
haimaker, Inc. is headquartered in San Francisco, California, and was founded in 2025 by Su Le and Wayne Pan. The team works with infrastructure partners across the UAE, the wider Middle East, and Southeast Asia.
ONE API. EVERY AI MODEL. ANY COUNTRY.
Get an API key in minutes, or talk to the team about dedicated and sovereign capacity.
Free to sign up · Pay as you go · Cancel anytime