Text Generation¶
Text Generation mode provides the main chat workflow for interacting with the configured local Leeroy model.
๐งญ Purpose¶
Text Generation mode lets users submit prompts, apply system instructions, tune generation controls, stream responses, and persist chat history in SQLite. It is the primary conversational interface for drafting, analysis, summarization, explanation, and general local LLM work.
๐งฑ Workflow Position¶
User Prompt
โ
โโโ System Instructions
โโโ Chat History
โโโ Optional Semantic Context
โโโ Optional Basic Documents
โ
โผ
Prompt Builder
โ
โผ
Local llama.cpp Model
โ
โผ
Streaming Chat Response
โ
โผ
SQLite Chat History
๐ฅ๏ธ Opening Text Generation Mode¶
- Start Leeroy.
- In the sidebar, select:
- Confirm that the main page displays the Text Generation heading and chat interface.
๐ง System Instructions¶
Use the System Instructions expander to define behavior before submitting prompts.
Examples:
System instructions are included in the prompt before the user message and can also be populated from reusable templates.
โ๏ธ Response Controls¶
| Control | Purpose | Practical Guidance |
|---|---|---|
| Temperature | Controls randomness. | Use lower values for factual or deterministic output. |
| Top-P | Controls nucleus sampling. | Keep near default unless generation quality needs tuning. |
| Top-K | Limits token candidates. | Lower values can make responses more conservative. |
| Use Grounding | Indicates whether grounding behavior should be used. | Enable when using contextual material. |
๐๏ธ Probability Controls¶
| Control | Purpose |
|---|---|
| Repeat Window | Sets the recent token window used for repetition checks. |
| Repeat Penalty | Discourages repeated phrases and loops. |
| Presence Penalty | Penalizes tokens that have already appeared. |
| Frequency Penalty | Penalizes tokens according to repeated frequency. |
๐๏ธ Context Controls¶
| Control | Purpose |
|---|---|
| Context Window | Sets how much context the local model can process. |
| CPU Threads | Controls CPU resources used by local inference. |
| Max Tokens | Limits the generated response length. |
| Random Seed | Supports reproducible output when fixed. |
๐ฌ Chat Workflow¶
- Enter a prompt in the chat input.
- Leeroy saves the user message.
- Leeroy builds a model prompt from system instructions, optional context, and chat history.
- The local model streams a response.
- Leeroy saves the assistant response.
- The conversation remains available through the current session and persisted SQLite history.
๐งช Example Prompts¶
๐งน Clearing Chat History¶
Use the Clear Chat button when you want to reset the current conversation.
Clearing chat history removes rows from the local chat_history table but does not remove prompt
templates, embeddings, imported data tables, or exception logs.
โ Recommended Sequence¶
- Set system instructions for the desired role or output style.
- Use conservative generation settings for analytical work.
- Submit the prompt.
- Review the streamed response.
- Adjust instructions rather than over-tuning sampling controls.
- Clear chat history when moving to a different task context.
๐ Related API Pages¶
| API Page | Purpose |
|---|---|
| App API | Source documentation for prompt building, chat persistence, and local model execution. |
| Configuration API | Runtime constants for model path, context defaults, and UI help text. |