GLM 5.3 Flash
z-ai/glm-5.3-flashGLM 5.3 Flash (z-ai/glm-5.3-flash) is an AI model from Z.ai with a 1,310,720-token context window and 131,072 max output tokens, priced at $0.07/1M input and $0.25/1M output tokens. Available via the haimaker.ai OpenAI-compatible API.
Overview
Zhipu AI's GLM (General Language Model) with reasoning and function calling capabilities.
Features & Capabilities
| Mode | chat |
| Context Window | 1,310,720 tokens |
| Max Output | 131,072 tokens |
| Function Calling | Supported |
| Vision | Supported |
| Reasoning | Supported |
| Web Search | Not supported |
| Url Context | Not supported |
API Usage
from openai import OpenAI
client = OpenAI(
base_url="https://api.haimaker.ai/v1",
api_key="YOUR_API_KEY",
)
response = client.chat.completions.create(
model="z-ai/glm-5.3-flash",
messages=[
{"role": "user", "content": "Hello, how are you?"}
],
)
print(response.choices[0].message.content)Frequently Asked Questions
What is the context window of GLM 5.3 Flash?
GLM 5.3 Flash (z-ai/glm-5.3-flash) has a 1,310,720-token context window and supports up to 131,072 output tokens per request.
How much does GLM 5.3 Flash cost?
GLM 5.3 Flash is priced at $0.07 per 1M input tokens and $0.25 per 1M output tokens when accessed via the haimaker.ai OpenAI-compatible API.
What features does GLM 5.3 Flash support?
GLM 5.3 Flash supports function calling, vision, reasoning.
How do I use GLM 5.3 Flash via API?
Send requests to https://api.haimaker.ai/v1/chat/completions with model "z-ai/glm-5.3-flash" using any OpenAI-compatible SDK. Authentication uses a Bearer API key from https://app.haimaker.ai.
Use GLM 5.3 Flash with the haimaker API
OpenAI-compatible endpoint. Start building in minutes.