haimaker.ai
SOVEREIGN
Model APIDedicated EndpointsBatch InferenceGPU InstancesGPU ClustersManaged Kubernetes
COMPUTE
START BUILDING NOW
SOVEREIGNDEVELOPERModel APIDedicated EndpointsBatch InferenceGPU InstancesGPU ClustersManaged KubernetesCOMPUTESTART BUILDING NOW
  1. Home
  2. /Models
  3. /NVIDIA
NVIDIA logo

NVIDIA Models

haimaker.ai provides API access to all 5 NVIDIA models, with context windows up to 262K tokens, priced from $0.05 to $0.60 per 1M input tokens. OpenAI-compatible endpoint, instant access.

Models
5
Max context
262K
From
$0.05
/1M input
Modes
chat
NewestNVIDIA Nemotron 3.5 Lightning 30B A3B NVFP4
NVIDIA logo

NVIDIA Nemotron 3.5 Lightning 30B A3B NVFP4

NVIDIA

nvidia/nemotron-3.5-lightning

17.8B params262K ctxIn: $0.07/1MOut: $0.20/1MCached in: $0.040/1M
function callingreasoning
NVIDIA logo

NVIDIA Nemotron 3 Ultra 550B A55B BF16

NVIDIA

nvidia/nemotron-3-ultra-550b-a55b

560.5B params262K ctxIn: $0.60/1MOut: $2.40/1MCached in: $0.12/1M
function callingreasoning
NVIDIA logo

Nemotron 3.5 Content Safety

NVIDIA

nvidia/nemotron-3.5-content-safety

4.3B params131K ctxIn: $0.20/1MOut: $0.20/1M
visionreasoning
NVIDIA logo

Nemotron 3 Nano Omni 30B A3B Reasoning BF16

NVIDIA

nvidia/nemotron-3-nano-30b-a3b

33.0B params262K ctxIn: $0.05/1MOut: $0.20/1MCached in: $0.030/1M
function callingreasoning
NVIDIA logo

NVIDIA Nemotron 3 Super 120B A12B NVFP4

NVIDIA

nvidia/nemotron-3-super-120b-a12b

67.2B params262K ctxIn: $0.08/1MOut: $0.45/1M
function callingreasoning
haimaker.ai

447 SUTTER ST STE 405 #122
SAN FRANCISCO, CA 94108

GET IN TOUCH
hello@haimaker.ai
SOVEREIGNDEVELOPERCOMPUTELEADERSHIPCONTACT US
BLOGDEVELOPER DOCSMODELSTERMSPRIVACY POLICY