Connect a Gateway
Open in CMSA gateway lets your agents use your own Claude Code or Codex (ChatGPT) account β by subscription or API key β instead of Qlar credits.
You run a small program on a computer or server you control, next to Claude Code and/or Codex. It opens one outbound connection to Qlar's gateway service, receives requests from your agents, runs them through those tools, and sends the answers back. Nothing connects in, so no firewall changes are needed.
Who can do this? Owners and Admins, in an organization where Qlar has turned gateways on. If you don't see Connect a gateway, ask Qlar support to turn gateways on for your organization.
Qlar models, API keys and gateways
An organization can get its models from three places. All three appear side by side under Organization β AI Models, and your agents pick from them the same way.
| Qlar | API key | Gateway | |
|---|---|---|---|
| Who pays the AI provider | Qlar, billed to you in credits | You, on your provider account | You, through your Claude or ChatGPT plan or API key |
| Qlar credits for the model | Yes | None | None |
| Where the key or sign-in lives | Qlar | Stored encrypted in Qlar | On your computer only |
| Needs a computer of yours to stay on | No | No | Yes |
| Requests at a time | Platform capacity | Your provider's limits | 2 at a time per gateway by default, plus your plan's limits |
| Set up in | Nothing to set up | Bring Your Own Key | This page |
Choose a gateway when you already pay for a Claude or ChatGPT plan and want your agents to use it, or when an API key must never leave your own machine. Choose an API key when the model has to be available around the clock without a computer of yours being on.
What a gateway protects, and what it does not
It does protect:
- Your network. The gateway only connects outward over HTTPS. No inbound port, no listening socket.
- Your keys and sign-ins. API keys you give the gateway are stored encrypted on that computer, and the Claude Code and Codex sign-ins stay where they are. None of them is ever sent to Qlar.
- Your computer. Claude Code and Codex run in an empty temporary folder with their own tools switched off: no shell, no files, no web, no MCP servers or plugins. A request that tries to use one is stopped. (Codex's image generation is the only exception, for image requests.)
- Who may connect. The one-time code is useless on its own. A gateway serves nothing until a person approves its key fingerprint in Qlar, and Qlar never stores the code.
It does not protect:
- The conversation itself. What users ask, and the answers, still pass through Qlar and reach Anthropic or OpenAI through your account. That account's own terms and data settings apply.
- Availability. The computer must stay on and the gateway running. While it is off, the agents using its models can't answer, and voice calls on it end.
- Your plan's limits. Your subscription's usage limits, fair-use rules and terms apply to everything your agents send through it. When the plan hits its limit, replies fail until it resets.
- Capacity. A gateway runs 2 requests at a time by default (the gateway's card shows its own limit). More at once are turned away as busy, so a gateway suits an internal team or a pilot better than a busy public website.
- Anyone who can use that computer as the user that runs the gateway can also use its signed-in tools.
Before you start
- A computer or server that can stay on:
- Windows 10 or 11, x64 (Windows on Arm runs it under emulation). Uses PowerShell; no admin rights needed to install.
- macOS on Apple silicon (M1 or later). Intel Macs are not supported.
- Linux x64 or arm64 with glibc (Debian, Ubuntu, Fedora, β¦). Alpine and other musl-based systems are not supported.
- Claude Code and/or Codex installed on that computer. You sign them in during step 4 β with your subscription (your Claude or ChatGPT plan) or an API key.
- Outbound HTTPS (port 443) from that computer. For voice on a ChatGPT subscription, also outbound UDP port 3478 (see Voice on your gateway).
- Owner or Admin role in the organization.
Setting it up
Open the Organization page β Configure on your organization β AI Models tile β Sources tab.

1. Name it and pick its system
Click Connect a gateway. Give the gateway a Name your team will recognise in model lists, for example "Office PC" (you can rename it later), and choose the system of the computer it will run on: Windows, macOS or Linux. Then click Create code.

2. Run the command on that computer
Qlar shows one command, built for the system you chose, with your one-time code already in it. Click Copy command, then on that computer:
| System | Where to paste it |
|---|---|
| Windows | PowerShell (no admin needed) |
| macOS | Terminal |
| Linux | A shell |
The command looks like this β always copy it from the page, because Qlar fills in the gateway service's address and your code:
# Windows
& ([scriptblock]::Create((irm 'https://<gateway service>/install.ps1'))) -Code 'XXXX-XXXX-XXXX' -Force
# macOS and Linux
curl -fsSL 'https://<gateway service>/install.sh' | sh -s -- --code 'XXXX-XXXX-XXXX' --force
It downloads the gateway, checks its checksum, installs it for your user (no admin rights), and connects it with your code. Leave the window open: the gateway keeps running in it.

Good to know:
- The code is single use and expires after 15 minutes. Qlar doesn't keep it, so it can't be shown again. If it expires or you lose it, click Get a new code.
- Setting up from your phone? Tap Send to my computer⦠to send the command to yourself (by message or email) and open it on the computer. Where sharing isn't available, the button copies the command instead.
- Running it on a computer that already has a gateway replaces that gateway. The command always re-enrols the computer.
- Gateway already installed? The page also shows a shorter command,
qlar-inference-gateway enroll --hub β¦ --code β¦ --force. - You can leave the page. Setup continues from the gateway's card under Sources. Cancel deletes a gateway that hasn't been approved yet.

3. Compare the fingerprint and approve
When the command has run, the gateway prints a key fingerprint in the terminal: 16 groups of 4 characters, in 4 rows. Qlar shows the computer's name, its system and the same fingerprint, laid out the same way.
Compare every group, and approve only if they are identical. Click Fingerprint matches β approve.

This is the step that matters. Until you approve, the gateway receives nothing. If the fingerprint is different from your terminal, someone else used your code: click It doesn't match β cancel and start again.
4. Sign in each tool and choose models
Back in the terminal, the gateway asks, for each tool it found, how it signs in:
- Subscription uses the tool's own sign-in: your Claude plan for Claude Code, your ChatGPT plan for Codex.
- API key asks for a key (typed hidden). It is stored encrypted on that computer and never reaches Qlar.
Qlar shows each tool's state β for example Claude Code Β· subscription (Max) or Codex (ChatGPT) Β· not signed in β and the models the computer offers. Tick the ones your agents may use and give each a name: that is what your team sees in model lists. Names use lower-case letters, digits, ., _ or -, and can't contain + (that marks Fast mode). Click Finish.

A tool that isn't signed in shows what to run on that computer β claude (then sign in) for Claude Code, codex login for Codex β as the same user that runs the gateway. After signing in, restart the gateway or wait up to 5 minutes for it to notice. You can come back to add models later with Models on the gateway's card.
Only models that can chat, call tools and return structured answers are offered; Claude Code and Codex both can.
5. Keep it running
The gateway runs in the window you started it from, so it stops when that window closes or the computer restarts. Once setup is done, stop it (Ctrl+C) and install it as a service, once:
| System | Command | Where |
|---|---|---|
| Windows | qlar-inference-gateway install-service | PowerShell opened with Run as administrator. It asks for your Windows password so the task can run as you at boot. |
| macOS | "$HOME/.local/bin/qlar-inference-gateway" install-service | Terminal. Starts at login. The installer prints this exact line. |
| Linux | sudo "$HOME/.local/bin/qlar-inference-gateway" install-service | A shell. Runs as you, restarts on failure. The installer prints this exact line. |
The service runs as your user, so it uses the same Claude Code and Codex sign-ins. That's it: the card shows Connected Β· β Online.
Using gateway models
Gateway models appear under Models in their own group, Your gateways, with where they run: β Office PC while online, β Office PC, offline 3 h ago while not.

From then on you use them like any other model:
- In an agent: as the conversation model, in Allowed Conversation AI Models, and as the Answering AI model of a Natural conversation voice call. Model lists group them under the gateway's name.
- As a feature default under Defaults, for every feature except KB Training.

What to expect:
- No Qlar credits are charged for a gateway model's tokens. Tools the answer uses β web search, knowledge base, plugins β are billed as usual.
- Not for content training or image description. Knowledge base training and describing uploaded images keep using Qlar or API-key models; gateway models aren't offered there.
- Replies fail while the computer is offline. Users see "This assistant can't answer right now. Please try again later." β never the computer's name. When all its request slots are in use they see "This assistant is busy right now. Please try again in a moment." The details, with the gateway's name and the reason, go to the agent's Monitoring Logs. The AI Models page also shows a banner while a gateway with models in use is offline.
- If Qlar turns gateways off for your organization, their models stop being used and features fall back to the Qlar default. Nothing is deleted: everything comes back as it was when gateways are turned on again. You can still remove a gateway meanwhile.
Voice on your gateway
A gateway with Codex can also run the voice of a Natural conversation call, using the realtime voice of the account signed in on that computer. Claude Code has no voice.
On the agent's Conversation β Voice page, choose Natural conversation, then under Voice runs on pick the gateway instead of Qlar. Gateways that can't take calls are shown greyed out with the reason (offline, not connected yet, its tools can't do voice, or Codex isn't signed in).

| What happens | |
|---|---|
| Voices | Assistant's voice lists the voices offered by the account on that gateway, starting with the gateway's default. There is no preview. |
| Calls at a time | By default 2 calls at once per gateway (the Voice page shows the gateway's own limit). Another caller sees that voice calls are busy. |
| Computer offline | The call ends. It never switches to Qlar's voice, so there are no surprise credits. (Unlike a Qlar voice call, it doesn't fall back to Turn-by-turn either.) |
| ChatGPT subscription | Subscription voice needs outbound UDP port 3478 from that computer. API-key voice doesn't. |
| Billing | No Qlar credits per minute. The Answering AI model's usage is billed as usual β and costs no credits either when it is a gateway model too. |
Voice on a gateway isn't available while Privacy Mode is on; calls of an agent set to a gateway end until it is off. If a gateway that runs an agent's voice is removed, that agent's calls stop until you choose where voice runs again. The Talk by Voice guide has more.
Running it
Each gateway's card under Sources shows its state, the computer, its system and version, how many requests are running, and how each tool is signed in. The β― menu holds the rest.
| Task | What to do |
|---|---|
| Check it is alive | The card shows β Online. On the computer, qlar-inference-gateway status shows the connection, the service and each tool's sign-in. |
| Rename it | β― β Rename. Only the name your team sees changes. |
| Add or remove models | Models on the card. Removing a model that is a feature default asks you to change that default first. Agents using a removed model fall back to the organization default. |
| Bring it back online | Bring it back online on the card. Same computer: start it again with qlar-inference-gateway run, or restart the service (Windows: schtasks /Run /TN "Qlar Inference Gateway"; macOS: launchctl kickstart -k gui/$(id -u)/ai.qlar.inference-gateway; Linux: sudo systemctl restart qlar-inference-gateway). New or reset computer: get a new code, as below. |
| Move it to another computer | β― β Move to another computer. You get a new one-time code to run on the new computer, then approve its fingerprint. Its models, the agents using them and their settings stay as they are. The old computer keeps working until you approve the new one, then it stops. |
| Update it | Run the command from step 2 again with --no-enroll in place of --code 'β¦' --force (on Windows, -NoEnroll in place of -Code 'β¦' -Force, after stopping the gateway: its file is locked while it runs). Then restart it. |
| Remove it | β― β Removeβ¦. The dialog lists what changes first (see below). |
Removing a gateway disconnects it for good β it stops on its computer and can't be used again β and:
- its models are removed from the organization, and agents using them fall back to the organization default;
- feature defaults that used them go back to the Qlar default, which uses Qlar credits;
- agents whose voice runs on it stop taking voice calls until another voice is chosen. They don't move to Qlar's voice.
A gateway that was never approved is removed straight away. To remove the program from the computer as well, run qlar-inference-gateway uninstall-service (elevated on Windows), "$HOME/.local/bin/qlar-inference-gateway" uninstall-service on macOS, or sudo "$HOME/.local/bin/qlar-inference-gateway" uninstall-service on Linux, then delete it.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| Card shows Connected Β· β Offline; replies fail | The gateway isn't running (the computer restarted, the window was closed) or can't reach the internet | Bring it back online on the card. Install it as a service (step 5) so it survives restarts; allow outbound HTTPS. |
| Waiting for the command to run | The command hasn't reported in yet | Run it on the computer. If the code is lost or expired, Get a new code. |
| Waiting for your approval | Nobody approved the fingerprint yet | Review fingerprint on the card, compare, approve. This works even if the page that created the code was closed. |
| Code expired | The command wasn't run within 15 minutes | Get a new code and run the new command. |
| The fingerprint in Qlar differs from the terminal | Someone else used your code, or the command ran on another computer | Don't approve. It doesn't match β cancel, then start again. |
| Approving says the key changed | A new key enrolled while you were comparing | Compare the new fingerprint again before approving. |
| Codex (ChatGPT) Β· not signed in (or Claude Code) | The tool isn't signed in for the user that runs the gateway β often the service user | On that computer, as that user: claude and sign in, or codex login. Then restart the gateway or wait up to 5 minutes. To use an API key instead: qlar-inference-gateway configure provider codex --auth api-key (or claude-code). |
| A model shows not offered by this computer | The tool was signed out, uninstalled, or no longer offers that model | Sign it in again or reinstall it; or remove the model and add another. |
| Users see "This assistant is busy right now" | More requests at once than the gateway allows (2 by default), or its voice-call limit reached | Wait, spread the load over another gateway, or use a Qlar or API-key model for busy agents. |
| Replies fail after working for a while | Your plan's usage limit was reached (the log says rate_limited) | Wait for the limit to reset, or switch that tool to an API key. |
| Voice calls end right after connecting, on a ChatGPT subscription | Outbound UDP 3478 is blocked | Allow it, or sign Codex in with an API key. |
| Revoked or No longer registered; the gateway exits | It was removed or revoked, or the gateway service no longer knows it (exit code 3 or 4) | Remove the card, then connect the computer again with a new code. |
| Status unavailable | Qlar couldn't reach the gateway service | Retry in a moment. |
| "A gateway is already running" when running the command | Another copy of the gateway is running on that computer | Stop it (Ctrl+C in its window, or stop the service), then run the command again. |