Model & Context
Open in CMSThe Model & Context page contains settings that influence which AI model the agent uses for conversation, what users may attach to their messages, and how much conversation history it "remembers."
Navigate to Conversation β Model & Context in the left sidebar. The CMS lists this page under Conversation; in these docs it belongs to the Fine-Tune Answer Quality stage, because the defaults work for most agents and you usually change them only to solve a specific problem.
Before you start
- Finish the basics first: Give Your Agent an Identity, Teach Your Agent, and Shape the Conversation.
- Write down the problem you want to solve (for example "answers are too slow", "users need to send photos", or "the agent forgets what was said earlier") and a few test questions that show it.
Model Capability
Conversation AI Model
The primary AI model used for all user-facing conversation. This model handles every message the user sends and generates the visible response.
It is also the model that answers during Natural conversation voice calls, unless the Voice page picks a different Answering AI model.
Choose based on the trade-off between response quality, speed, and cost:
| Consideration | Guidance |
|---|---|
| High accuracy required | Use the most capable model available |
| High volume / cost-sensitive | Use a smaller, faster model |
| Mixed needs | Enable User to Change AI Model below |
Model levels and + models
Qlar's built-in models come in six levels, from fastest and cheapest to deepest-thinking. The picker shows each model's description and its estimated credit ratio β how much it costs relative to the others.
| Model | Credit ratio | Best for |
|---|---|---|
| Light | 1.5x | FAQ bots, auto-replies, simple formatting; answers without reasoning first |
| Smart | 3x | Everyday customer service, bookings, basic sales |
| Analyst | 35x | Structured workflows: forms, API orchestration, rule-based automation |
| Insight | 45x | Knowledge-base Q&A and support that needs context |
| Thinker | 90x | Complex, multi-step reasoning: data analysis, compliance checks |
| Sage | 230x | Strategic, high-stakes decisions |
Every level also has a + version β Light+, Smart+, Analyst+, Insight+, Thinker+, Sage+. A + model is the same model, thinking exactly the same way, running on Fast mode: its answers come about 1.5β2x faster, at twice the credit ratio (Light+ 3x, Smart+ 6x, Analyst+ 70x, Insight+ 90x, Thinker+ 180x, Sage+ 460x). Choose a + model when waiting matters β busy chats, voice calls β not for better answers. Under heavy load an answer can occasionally run at normal speed; it is then billed at the normal rate. Background work such as summaries and knowledge-base training never uses Fast mode, so it costs the same with either version.
Enable User to Change AI Model
Allows end users to switch between different AI models during their conversation. When enabled, a model picker appears in the chat interface.
Allowed Conversation AI Models
Visible only when Enable User to Change AI Model is on. Select which models users are allowed to choose from. Users cannot pick a model outside this list. If Allow Image Upload is on, include at least one model that accepts images.
The first model in the list becomes the default β it is the model the agent uses at the start of every conversation before the user makes a selection.
Enforce Thinking
When enabled, the AI is required to reason through its approach before generating the final response. This improves the quality of answers to complex or multi-step questions.
Use this for agents dealing with analysis, troubleshooting, or any task requiring careful reasoning.
Trade-off: Thinking adds latency. Responses take slightly longer to appear.
Attachments
These settings decide what users can attach β in text chat, and during an immersive voice call. They sit on this page because whether an image can be read depends on the model chosen above.
Allow Image Upload
Lets users attach images (JPG, PNG, WEBP) to their messages. In text chat, the Conversation AI Model must accept image input; if it doesn't β or, with model switching on, if none of the allowed models do β the page shows a warning and users won't be able to send images in chat. Turn-by-turn voice calls are not affected: that style always reads images. In Natural conversation, images go to the Voice page's Answering AI model, so whether they can be sent in a call depends on that model β see Voice β Attachments.
Allow Document/File Upload
Lets users attach documents (PDF, TXT, MD). The document's text is extracted independently of the AI model (scanned PDFs go through OCR), so this isn't affected by model capability.
Both switches apply to chat and to immersive voice calls alike. Turn both off and neither the chat input nor the voice call offers an attach option.
Document Upload Restrictions
A free-text field where you describe what users must not upload. Every attached file β image or document, in chat or in a voice call β is checked against this text when it is uploaded, and a file that violates it is refused with the reason, before it ever reaches the conversation.
Example restriction:
Prevent upload of documents containing personally identifiable information such as national ID numbers, passport numbers, or documents with photographs of individuals. Block files with sensitive financial data such as credit card numbers or bank account details.
Leave this field blank to allow all supported files.
Context Window
These settings control how much of the conversation history the agent "remembers" when composing each response.
Summarize Previous Answer
When enabled, the agent summarizes its own previous responses before including them in the context window. This reduces the number of tokens used to represent prior turns, allowing longer conversations to fit within the model's context limit.
Use this for agents that handle extended multi-turn conversations.
Trade-off: Summaries may lose fine-grained detail from earlier answers. If precision in referencing past responses is important, keep this off.
Keep Function Tools Answer
When enabled, the output of any external tools or function calls (e.g., API responses, database query results) is retained in the conversation context for follow-up turns.
Use this when users are likely to ask follow-up questions about data retrieved by a tool (e.g., "Can you sort that by date?" after a database query).
User Messages to Carry Over
The number of recent user messages to include in the context window when starting a new conversation segment or after thread separation.
| Value | Behaviour |
|---|---|
0 | No history is carried into the new segment |
10 (default) | The last 10 user messages and their responses are included |
| Higher values | More continuity but higher token usage |
Set this based on how much conversational continuity matters for your agent and how long typical conversations are.
Max Response Token
The longest answer the agent may write, in tokens. 0 (the default) follows the model's own maximum. A positive value hard-stops the response at that many tokens β useful to cap cost, at the risk of cutting a long answer short.
Saving Changes
Click Save after adjusting any setting. Publish your agent to propagate the changes to end users.
Check the result
After you save, click the reload button in the preview panel on the right side of the page to start a fresh conversation, then repeat the test questions you wrote down. Compare speed and answer quality with before. If you allowed attachments, try sending an image or document and check that the agent can read it.
The preview runs your draft. End users get the new model and context settings only after you publish. See Preview and Publish Changes.
Next steps
- Retrieval: tune how the agent searches your knowledge (next step).
- Chat: language, dictation, formatting, and other text-conversation settings.
- Talk by Voice: the voice call, including attachments during an immersive call.
Looking for the default timezone? It is set in Profile.