Connect a Gateway

Open in CMS

A gateway lets your agents use your own Claude Code or Codex (ChatGPT) account β€” by subscription or API key β€” instead of Qlar credits.

You run a small program on a computer or server you control, next to Claude Code and/or Codex. It opens one outbound connection to Qlar's gateway service, receives requests from your agents, runs them through those tools, and sends the answers back. Nothing connects in, so no firewall changes are needed.

Who can do this? Owners and Admins, in an organization where Qlar has turned gateways on. If you don't see Connect a gateway, ask Qlar support to turn gateways on for your organization.


Qlar models, API keys and gateways

An organization can get its models from three places. All three appear side by side under Organization β†’ AI Models, and your agents pick from them the same way.

QlarAPI keyGateway
Who pays the AI providerQlar, billed to you in creditsYou, on your provider accountYou, through your Claude or ChatGPT plan or API key
Qlar credits for the modelYesNoneNone
Where the key or sign-in livesQlarStored encrypted in QlarOn your computer only
Needs a computer of yours to stay onNoNoYes
Requests at a timePlatform capacityYour provider's limits2 at a time per gateway by default, plus your plan's limits
Set up inNothing to set upBring Your Own KeyThis page

Choose a gateway when you already pay for a Claude or ChatGPT plan and want your agents to use it, or when an API key must never leave your own machine. Choose an API key when the model has to be available around the clock without a computer of yours being on.


What a gateway protects, and what it does not

It does protect:

  • Your network. The gateway only connects outward over HTTPS. No inbound port, no listening socket.
  • Your keys and sign-ins. API keys you give the gateway are stored encrypted on that computer, and the Claude Code and Codex sign-ins stay where they are. None of them is ever sent to Qlar.
  • Your computer. Claude Code and Codex run in an empty temporary folder with their own tools switched off: no shell, no files, no web, no MCP servers or plugins. A request that tries to use one is stopped. (Codex's image generation is the only exception, for image requests.)
  • Who may connect. The one-time code is useless on its own. A gateway serves nothing until a person approves its key fingerprint in Qlar, and Qlar never stores the code.

It does not protect:

  • The conversation itself. What users ask, and the answers, still pass through Qlar and reach Anthropic or OpenAI through your account. That account's own terms and data settings apply.
  • Availability. The computer must stay on and the gateway running. While it is off, the agents using its models can't answer, and voice calls on it end.
  • Your plan's limits. Your subscription's usage limits, fair-use rules and terms apply to everything your agents send through it. When the plan hits its limit, replies fail until it resets.
  • Capacity. A gateway runs 2 requests at a time by default (the gateway's card shows its own limit). More at once are turned away as busy, so a gateway suits an internal team or a pilot better than a busy public website.
  • Anyone who can use that computer as the user that runs the gateway can also use its signed-in tools.

Before you start

  • A computer or server that can stay on:
    • Windows 10 or 11, x64 (Windows on Arm runs it under emulation). Uses PowerShell; no admin rights needed to install.
    • macOS on Apple silicon (M1 or later). Intel Macs are not supported.
    • Linux x64 or arm64 with glibc (Debian, Ubuntu, Fedora, …). Alpine and other musl-based systems are not supported.
  • Claude Code and/or Codex installed on that computer. You sign them in during step 4 β€” with your subscription (your Claude or ChatGPT plan) or an API key.
  • Outbound HTTPS (port 443) from that computer. For voice on a ChatGPT subscription, also outbound UDP port 3478 (see Voice on your gateway).
  • Owner or Admin role in the organization.

Setting it up

Open the Organization page β†’ Configure on your organization β†’ AI Models tile β†’ Sources tab.

Sources tab of AI Models with the Gateways section: Office PC connected and online with one model, Laptop offline with a Bring it back online button, and Mac mini waiting for approval of a new computer with a Review fingerprint button

1. Name it and pick its system

Click Connect a gateway. Give the gateway a Name your team will recognise in model lists, for example "Office PC" (you can rename it later), and choose the system of the computer it will run on: Windows, macOS or Linux. Then click Create code.

Connect a gateway, step 1 of 4: the Name field filled with Front desk PC, Where will it run? with Windows, macOS and Linux, and the Create code button

2. Run the command on that computer

Qlar shows one command, built for the system you chose, with your one-time code already in it. Click Copy command, then on that computer:

SystemWhere to paste it
WindowsPowerShell (no admin needed)
macOSTerminal
LinuxA shell

The command looks like this β€” always copy it from the page, because Qlar fills in the gateway service's address and your code:

# Windows
& ([scriptblock]::Create((irm 'https://<gateway service>/install.ps1'))) -Code 'XXXX-XXXX-XXXX' -Force
# macOS and Linux
curl -fsSL 'https://<gateway service>/install.sh' | sh -s -- --code 'XXXX-XXXX-XXXX' --force

It downloads the gateway, checks its checksum, installs it for your user (no admin rights), and connects it with your code. Leave the window open: the gateway keeps running in it.

Connect a gateway, step 2 of 4: the Windows command in a dark box, a full-width Copy command button, the one-time code, the enrol command for a computer that already has the gateway, and a countdown while waiting for the gateway to report in

Good to know:

  • The code is single use and expires after 15 minutes. Qlar doesn't keep it, so it can't be shown again. If it expires or you lose it, click Get a new code.
  • Setting up from your phone? Tap Send to my computer… to send the command to yourself (by message or email) and open it on the computer. Where sharing isn't available, the button copies the command instead.
  • Running it on a computer that already has a gateway replaces that gateway. The command always re-enrols the computer.
  • Gateway already installed? The page also shows a shorter command, qlar-inference-gateway enroll --hub … --code … --force.
  • You can leave the page. Setup continues from the gateway's card under Sources. Cancel deletes a gateway that hasn't been approved yet.
Step 2 on a phone: the command wraps inside its box, with Send to my computer… above Copy command

3. Compare the fingerprint and approve

When the command has run, the gateway prints a key fingerprint in the terminal: 16 groups of 4 characters, in 4 rows. Qlar shows the computer's name, its system and the same fingerprint, laid out the same way.

Compare every group, and approve only if they are identical. Click Fingerprint matches β€” approve.

Connect a gateway, step 3 of 4: A gateway has enrolled β€” confirm it is yours, with the computer DESKTOP-7Q, Windows x64, gateway v0.1.3, the key fingerprint as a grid of 16 groups, and the buttons Fingerprint matches β€” approve and It doesn't match β€” cancel

This is the step that matters. Until you approve, the gateway receives nothing. If the fingerprint is different from your terminal, someone else used your code: click It doesn't match β€” cancel and start again.

4. Sign in each tool and choose models

Back in the terminal, the gateway asks, for each tool it found, how it signs in:

  • Subscription uses the tool's own sign-in: your Claude plan for Claude Code, your ChatGPT plan for Codex.
  • API key asks for a key (typed hidden). It is stored encrypted on that computer and never reaches Qlar.

Qlar shows each tool's state β€” for example Claude Code Β· subscription (Max) or Codex (ChatGPT) Β· not signed in β€” and the models the computer offers. Tick the ones your agents may use and give each a name: that is what your team sees in model lists. Names use lower-case letters, digits, ., _ or -, and can't contain + (that marks Fast mode). Click Finish.

Connect a gateway, step 4 of 4: Front desk PC is approved and online, Claude Code signed in with a subscription, Codex not signed in with the command to run, and a checklist of models with a name for each

A tool that isn't signed in shows what to run on that computer β€” claude (then sign in) for Claude Code, codex login for Codex β€” as the same user that runs the gateway. After signing in, restart the gateway or wait up to 5 minutes for it to notice. You can come back to add models later with Models on the gateway's card.

Only models that can chat, call tools and return structured answers are offered; Claude Code and Codex both can.

5. Keep it running

The gateway runs in the window you started it from, so it stops when that window closes or the computer restarts. Once setup is done, stop it (Ctrl+C) and install it as a service, once:

SystemCommandWhere
Windowsqlar-inference-gateway install-servicePowerShell opened with Run as administrator. It asks for your Windows password so the task can run as you at boot.
macOS"$HOME/.local/bin/qlar-inference-gateway" install-serviceTerminal. Starts at login. The installer prints this exact line.
Linuxsudo "$HOME/.local/bin/qlar-inference-gateway" install-serviceA shell. Runs as you, restarts on failure. The installer prints this exact line.

The service runs as your user, so it uses the same Claude Code and Codex sign-ins. That's it: the card shows Connected Β· ● Online.


Using gateway models

Gateway models appear under Models in their own group, Your gateways, with where they run: ● Office PC while online, β—‹ Office PC, offline 3 h ago while not.

Models tab with the filter chips All, Qlar, API key and Gateway, and the groups Qlar, Your API keys and Your gateways; the gateway models show Office PC online and Laptop offline

From then on you use them like any other model:

  • In an agent: as the conversation model, in Allowed Conversation AI Models, and as the Answering AI model of a Natural conversation voice call. Model lists group them under the gateway's name.
  • As a feature default under Defaults, for every feature except KB Training.
Defaults tab with Summarization set to sonnet-office on Office PC and the other features on the platform default

What to expect:

  • No Qlar credits are charged for a gateway model's tokens. Tools the answer uses β€” web search, knowledge base, plugins β€” are billed as usual.
  • Not for content training or image description. Knowledge base training and describing uploaded images keep using Qlar or API-key models; gateway models aren't offered there.
  • Replies fail while the computer is offline. Users see "This assistant can't answer right now. Please try again later." β€” never the computer's name. When all its request slots are in use they see "This assistant is busy right now. Please try again in a moment." The details, with the gateway's name and the reason, go to the agent's Monitoring Logs. The AI Models page also shows a banner while a gateway with models in use is offline.
  • If Qlar turns gateways off for your organization, their models stop being used and features fall back to the Qlar default. Nothing is deleted: everything comes back as it was when gateways are turned on again. You can still remove a gateway meanwhile.

Voice on your gateway

A gateway with Codex can also run the voice of a Natural conversation call, using the realtime voice of the account signed in on that computer. Claude Code has no voice.

On the agent's Conversation β†’ Voice page, choose Natural conversation, then under Voice runs on pick the gateway instead of Qlar. Gateways that can't take calls are shown greyed out with the reason (offline, not connected yet, its tools can't do voice, or Codex isn't signed in).

Conversation style card with Natural conversation selected and Voice runs on set to Office PC (online), Lab server offline and disabled, and a warning listing the gateway's limits
What happens
VoicesAssistant's voice lists the voices offered by the account on that gateway, starting with the gateway's default. There is no preview.
Calls at a timeBy default 2 calls at once per gateway (the Voice page shows the gateway's own limit). Another caller sees that voice calls are busy.
Computer offlineThe call ends. It never switches to Qlar's voice, so there are no surprise credits. (Unlike a Qlar voice call, it doesn't fall back to Turn-by-turn either.)
ChatGPT subscriptionSubscription voice needs outbound UDP port 3478 from that computer. API-key voice doesn't.
BillingNo Qlar credits per minute. The Answering AI model's usage is billed as usual β€” and costs no credits either when it is a gateway model too.

Voice on a gateway isn't available while Privacy Mode is on; calls of an agent set to a gateway end until it is off. If a gateway that runs an agent's voice is removed, that agent's calls stop until you choose where voice runs again. The Talk by Voice guide has more.


Running it

Each gateway's card under Sources shows its state, the computer, its system and version, how many requests are running, and how each tool is signed in. The β‹― menu holds the rest.

TaskWhat to do
Check it is aliveThe card shows ● Online. On the computer, qlar-inference-gateway status shows the connection, the service and each tool's sign-in.
Rename itβ‹― β†’ Rename. Only the name your team sees changes.
Add or remove modelsModels on the card. Removing a model that is a feature default asks you to change that default first. Agents using a removed model fall back to the organization default.
Bring it back onlineBring it back online on the card. Same computer: start it again with qlar-inference-gateway run, or restart the service (Windows: schtasks /Run /TN "Qlar Inference Gateway"; macOS: launchctl kickstart -k gui/$(id -u)/ai.qlar.inference-gateway; Linux: sudo systemctl restart qlar-inference-gateway). New or reset computer: get a new code, as below.
Move it to another computerβ‹― β†’ Move to another computer. You get a new one-time code to run on the new computer, then approve its fingerprint. Its models, the agents using them and their settings stay as they are. The old computer keeps working until you approve the new one, then it stops.
Update itRun the command from step 2 again with --no-enroll in place of --code '…' --force (on Windows, -NoEnroll in place of -Code '…' -Force, after stopping the gateway: its file is locked while it runs). Then restart it.
Remove itβ‹― β†’ Remove…. The dialog lists what changes first (see below).

Removing a gateway disconnects it for good β€” it stops on its computer and can't be used again β€” and:

  • its models are removed from the organization, and agents using them fall back to the organization default;
  • feature defaults that used them go back to the Qlar default, which uses Qlar credits;
  • agents whose voice runs on it stop taking voice calls until another voice is chosen. They don't move to Qlar's voice.

A gateway that was never approved is removed straight away. To remove the program from the computer as well, run qlar-inference-gateway uninstall-service (elevated on Windows), "$HOME/.local/bin/qlar-inference-gateway" uninstall-service on macOS, or sudo "$HOME/.local/bin/qlar-inference-gateway" uninstall-service on Linux, then delete it.


Troubleshooting

SymptomCauseFix
Card shows Connected Β· β—‹ Offline; replies failThe gateway isn't running (the computer restarted, the window was closed) or can't reach the internetBring it back online on the card. Install it as a service (step 5) so it survives restarts; allow outbound HTTPS.
Waiting for the command to runThe command hasn't reported in yetRun it on the computer. If the code is lost or expired, Get a new code.
Waiting for your approvalNobody approved the fingerprint yetReview fingerprint on the card, compare, approve. This works even if the page that created the code was closed.
Code expiredThe command wasn't run within 15 minutesGet a new code and run the new command.
The fingerprint in Qlar differs from the terminalSomeone else used your code, or the command ran on another computerDon't approve. It doesn't match β€” cancel, then start again.
Approving says the key changedA new key enrolled while you were comparingCompare the new fingerprint again before approving.
Codex (ChatGPT) Β· not signed in (or Claude Code)The tool isn't signed in for the user that runs the gateway β€” often the service userOn that computer, as that user: claude and sign in, or codex login. Then restart the gateway or wait up to 5 minutes. To use an API key instead: qlar-inference-gateway configure provider codex --auth api-key (or claude-code).
A model shows not offered by this computerThe tool was signed out, uninstalled, or no longer offers that modelSign it in again or reinstall it; or remove the model and add another.
Users see "This assistant is busy right now"More requests at once than the gateway allows (2 by default), or its voice-call limit reachedWait, spread the load over another gateway, or use a Qlar or API-key model for busy agents.
Replies fail after working for a whileYour plan's usage limit was reached (the log says rate_limited)Wait for the limit to reset, or switch that tool to an API key.
Voice calls end right after connecting, on a ChatGPT subscriptionOutbound UDP 3478 is blockedAllow it, or sign Codex in with an API key.
Revoked or No longer registered; the gateway exitsIt was removed or revoked, or the gateway service no longer knows it (exit code 3 or 4)Remove the card, then connect the computer again with a new code.
Status unavailableQlar couldn't reach the gateway serviceRetry in a moment.
"A gateway is already running" when running the commandAnother copy of the gateway is running on that computerStop it (Ctrl+C in its window, or stop the service), then run the command again.