Model & Context

Open in CMS

The Model & Context page contains settings that influence which AI model the agent uses for conversation, what users may attach to their messages, and how much conversation history it "remembers."

Navigate to Conversation β†’ Model & Context in the left sidebar. The CMS lists this page under Conversation; in these docs it belongs to the Fine-Tune Answer Quality stage, because the defaults work for most agents and you usually change them only to solve a specific problem.


Before you start


Model Capability

Conversation AI Model

The primary AI model used for all user-facing conversation. This model handles every message the user sends and generates the visible response.

It is also the model that answers during Natural conversation voice calls, unless the Voice page picks a different Answering AI model.

Choose based on the trade-off between response quality, speed, and cost:

ConsiderationGuidance
High accuracy requiredUse the most capable model available
High volume / cost-sensitiveUse a smaller, faster model
Mixed needsEnable User to Change AI Model below

Model levels and + models

Qlar's built-in models come in six levels, from fastest and cheapest to deepest-thinking. The picker shows each model's description and its estimated credit ratio β€” how much it costs relative to the others.

ModelCredit ratioBest for
Light1.5xFAQ bots, auto-replies, simple formatting; answers without reasoning first
Smart3xEveryday customer service, bookings, basic sales
Analyst35xStructured workflows: forms, API orchestration, rule-based automation
Insight45xKnowledge-base Q&A and support that needs context
Thinker90xComplex, multi-step reasoning: data analysis, compliance checks
Sage230xStrategic, high-stakes decisions

Every level also has a + version β€” Light+, Smart+, Analyst+, Insight+, Thinker+, Sage+. A + model is the same model, thinking exactly the same way, running on Fast mode: its answers come about 1.5–2x faster, at twice the credit ratio (Light+ 3x, Smart+ 6x, Analyst+ 70x, Insight+ 90x, Thinker+ 180x, Sage+ 460x). Choose a + model when waiting matters β€” busy chats, voice calls β€” not for better answers. Under heavy load an answer can occasionally run at normal speed; it is then billed at the normal rate. Background work such as summaries and knowledge-base training never uses Fast mode, so it costs the same with either version.

Enable User to Change AI Model

Allows end users to switch between different AI models during their conversation. When enabled, a model picker appears in the chat interface.

Allowed Conversation AI Models

Visible only when Enable User to Change AI Model is on. Select which models users are allowed to choose from. Users cannot pick a model outside this list. If Allow Image Upload is on, include at least one model that accepts images.

The first model in the list becomes the default β€” it is the model the agent uses at the start of every conversation before the user makes a selection.

Enforce Thinking

When enabled, the AI is required to reason through its approach before generating the final response. This improves the quality of answers to complex or multi-step questions.

Use this for agents dealing with analysis, troubleshooting, or any task requiring careful reasoning.

Trade-off: Thinking adds latency. Responses take slightly longer to appear.


Attachments

These settings decide what users can attach β€” in text chat, and during an immersive voice call. They sit on this page because whether an image can be read depends on the model chosen above.

Allow Image Upload

Lets users attach images (JPG, PNG, WEBP) to their messages. In text chat, the Conversation AI Model must accept image input; if it doesn't β€” or, with model switching on, if none of the allowed models do β€” the page shows a warning and users won't be able to send images in chat. Turn-by-turn voice calls are not affected: that style always reads images. In Natural conversation, images go to the Voice page's Answering AI model, so whether they can be sent in a call depends on that model β€” see Voice β†’ Attachments.

Allow Document/File Upload

Lets users attach documents (PDF, TXT, MD). The document's text is extracted independently of the AI model (scanned PDFs go through OCR), so this isn't affected by model capability.

Both switches apply to chat and to immersive voice calls alike. Turn both off and neither the chat input nor the voice call offers an attach option.

Document Upload Restrictions

A free-text field where you describe what users must not upload. Every attached file β€” image or document, in chat or in a voice call β€” is checked against this text when it is uploaded, and a file that violates it is refused with the reason, before it ever reaches the conversation.

Example restriction:

Prevent upload of documents containing personally identifiable information such as national ID numbers, passport numbers, or documents with photographs of individuals. Block files with sensitive financial data such as credit card numbers or bank account details.

Leave this field blank to allow all supported files.


Context Window

These settings control how much of the conversation history the agent "remembers" when composing each response.

Summarize Previous Answer

When enabled, the agent summarizes its own previous responses before including them in the context window. This reduces the number of tokens used to represent prior turns, allowing longer conversations to fit within the model's context limit.

Use this for agents that handle extended multi-turn conversations.

Trade-off: Summaries may lose fine-grained detail from earlier answers. If precision in referencing past responses is important, keep this off.

Keep Function Tools Answer

When enabled, the output of any external tools or function calls (e.g., API responses, database query results) is retained in the conversation context for follow-up turns.

Use this when users are likely to ask follow-up questions about data retrieved by a tool (e.g., "Can you sort that by date?" after a database query).

User Messages to Carry Over

The number of recent user messages to include in the context window when starting a new conversation segment or after thread separation.

ValueBehaviour
0No history is carried into the new segment
10 (default)The last 10 user messages and their responses are included
Higher valuesMore continuity but higher token usage

Set this based on how much conversational continuity matters for your agent and how long typical conversations are.

Max Response Token

The longest answer the agent may write, in tokens. 0 (the default) follows the model's own maximum. A positive value hard-stops the response at that many tokens β€” useful to cap cost, at the risk of cutting a long answer short.


Saving Changes

Click Save after adjusting any setting. Publish your agent to propagate the changes to end users.


Check the result

After you save, click the reload button in the preview panel on the right side of the page to start a fresh conversation, then repeat the test questions you wrote down. Compare speed and answer quality with before. If you allowed attachments, try sending an image or document and check that the agent can read it.

The preview runs your draft. End users get the new model and context settings only after you publish. See Preview and Publish Changes.


Next steps

  • Retrieval: tune how the agent searches your knowledge (next step).
  • Chat: language, dictation, formatting, and other text-conversation settings.
  • Talk by Voice: the voice call, including attachments during an immersive call.

Looking for the default timezone? It is set in Profile.