App
app ¶
Assembly: Leeroy
Filename: app.py
Author: Terry D. Eppler
Created: 05-31-2024
Last Modified By: Terry D. Eppler
Last Modified On: 05-01-2025
Leeroy is a data analysis tool integrating various Generative GPT, Text-Processing, and
Machine-Learning algorithms for federal analysts.
Copyright © 2022 Terry Eppler
Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NON-INFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
You can contact me at: terryeppler@gmail.com or eppler.terry@epa.gov
write_error ¶
Write an application error record without interrupting fallback handling.
Purpose
Wraps a caught exception in the project Error object and writes it through the
SQLite-backed Logger. The helper is used by existing fallback handlers that must
preserve their original return behavior while still recording diagnostic metadata.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
error
|
Exception
|
Caught exception instance to wrap and persist. |
required |
cause
|
str
|
Logical workflow component associated with the failure. |
required |
method
|
str
|
Stable function or method signature associated with the failure. |
required |
Source code in app.py
local_llm_available ¶
Determine whether the configured local model file is available.
Purpose
Checks the resolved model-path state created during module import and returns whether the optional local GGUF file exists in the current execution environment. This helper allows the Streamlit UI and local model loader to degrade safely when the model file is not present.
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Return value produced by the operation. |
Source code in app.py
throw_if ¶
Throw if.
Purpose
Validates that a required argument contains a usable value before the surrounding workflow continues. This guard centralizes early validation so provider wrappers and UI routines fail with consistent, readable error messages.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Name value used by the operation. |
required |
value
|
object
|
Value value used by the operation. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
This function performs its work through side effects and does not return a value. |
Source code in app.py
image_to_base64 ¶
Convert an image file to a Base64 text string.
Purpose
Reads a local image file and converts the raw bytes into a Base64-encoded string that can be embedded in Streamlit or Markdown output. The helper supports UI image rendering workflows that need inline image data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str
|
Filesystem path to the image file. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
cosine_sim ¶
Calculate cosine similarity between two vectors.
Purpose
Computes normalized dot-product similarity for semantic retrieval, document chunk ranking, and fallback vector matching when sqlite-vec is unavailable. The helper returns a safe zero score when either vector has zero magnitude.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
a
|
ndarray
|
First vector. |
required |
b
|
ndarray
|
Second vector. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
float |
float
|
Return value produced by the operation. |
Source code in app.py
initialize_database ¶
Create required application database tables.
Purpose
Ensures the local SQLite storage directory exists and creates the tables required for chat history, semantic embeddings, and categorized prompt templates.
Raises:
| Type | Description |
|---|---|
Error
|
Raised when SQLite initialization fails after writing diagnostic metadata. |
Source code in app.py
normalize_text ¶
Normalize text for matching and comparison workflows.
Purpose
Standardizes free text by lowercasing content, removing punctuation except sentence delimiters, normalizing sentence-boundary spacing, and collapsing repeated whitespace. The helper supports prompt, document, and search workflows that need consistent text comparison behavior.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Source text to process. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
chunk_text ¶
Split text into overlapping chunks.
Purpose
Creates overlapping text windows used by semantic indexing, retrieval-augmented generation, and document Q&A workflows. The overlap preserves local context across chunk boundaries so retrieved excerpts remain coherent.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Source text to process. |
required |
size
|
int
|
Maximum character length for each chunk. |
1200
|
overlap
|
int
|
Number of characters shared between adjacent chunks. |
200
|
Returns:
| Type | Description |
|---|---|
List[str]
|
list[str]: Return value produced by the operation. |
Source code in app.py
convert_xml ¶
Convert XML-like prompt sections into Markdown.
Purpose
Transforms prompt text containing XML-like opening and closing tags into Markdown section blocks. The function treats tags as lightweight section delimiters instead of strict XML, allowing prompt templates to be rendered more clearly in the UI.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Source text to process. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
markdown_converter ¶
Convert between Markdown headings and XML-like heading tags.
Purpose
Auto-detects whether the supplied text contains simple hN heading tags or Markdown heading syntax. HTML-like headings are converted to Markdown headings; otherwise Markdown headings are converted to matching hN tags. Non-string or empty inputs return an empty string.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
Any
|
Source text to process. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
inject_response_css ¶
Inject chat-response CSS into the Streamlit page.
Purpose
Adds inline CSS that styles chat-message paragraphs, headings, and links inside Streamlit chat responses. The style layer keeps generated responses visually consistent with the Leeroy dark-mode interface and shared blue accent color.
Source code in app.py
style_subheaders ¶
Inject shared subheader CSS into the Streamlit page.
Purpose
Applies the Leeroy blue accent color to selected Markdown and chat subheaders in the main Streamlit UI. The helper centralizes visual styling used by sidebar and mode- rendering sections.
Source code in app.py
save_message ¶
Persist a chat message to SQLite.
Purpose
Writes a single chat-history record to the configured application database. The function supports durable conversation history for the Text Generation and Document Q&A modes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
role
|
str
|
Chat role associated with the message. |
required |
content
|
str
|
Message content to store. |
required |
Source code in app.py
load_history ¶
Load persisted chat history from SQLite.
Purpose
Reads stored chat messages in insertion order so Streamlit session state can be reconstructed when the application starts or reruns.
Returns:
| Type | Description |
|---|---|
List[Tuple[str, str]]
|
list[tuple[str, str]]: Return value produced by the operation. |
Source code in app.py
clear_history ¶
Delete persisted chat history.
Purpose
Removes all rows from the SQLite chat_history table so the Streamlit chat UI can be reset without affecting prompt templates, embeddings, or imported data.
Source code in app.py
supports_prompt_category ¶
Determine whether a prompt category fits the local text model.
Purpose
Excludes categories that clearly require unsupported image, audio, video, speech, or hosted-tool capabilities while retaining text, code, retrieval, and document workflows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
category
|
str
|
Stored prompt category name. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
True when the category is suitable for Leeroy's local Llama 3.2 model. |
Source code in app.py
fetch_prompt_categories ¶
Retrieve prompt categories supported by the local text model.
Purpose
Reads distinct category names from the configured Prompts table and returns the alphabetically ordered categories that fit Leeroy's text-only local model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
db_path
|
str
|
SQLite database path. |
required |
Returns:
| Type | Description |
|---|---|
List[str]
|
list[str]: Supported prompt categories. |
Source code in app.py
fetch_prompt_choices ¶
Retrieve prompt identifiers and display labels for a category.
Purpose
Loads category-filtered prompt choices using the primary key as the selector value and Caption plus Name as the human-readable label.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
db_path
|
str
|
SQLite database path. |
required |
category
|
str
|
Selected prompt category. |
required |
Returns:
| Type | Description |
|---|---|
List[Tuple[int, str]]
|
list[tuple[int, str]]: Prompt identifiers paired with display labels. |
Source code in app.py
fetch_prompts ¶
Load prompt metadata for prompt administration.
Purpose
Reads prompt metadata from the configured SQLite Prompts table, orders rows by most recent identifier first, and inserts a selection flag used by the Prompt Engineering data editor.
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: Return value produced by the operation. |
Source code in app.py
fetch_prompt_by_id ¶
Load a prompt record by primary key.
Purpose
Retrieves the full prompt record for a selected ID value and returns a dictionary keyed by SQLite column names for prompt-editing workflows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pid
|
int
|
Prompt primary key value. |
required |
Returns:
| Type | Description |
|---|---|
Dict[str, Any] | None
|
dict[str, Any] | None: Return value produced by the operation. |
Source code in app.py
fetch_prompt_by_name ¶
Load a prompt record by caption.
Purpose
Retrieves the full prompt record for a selected prompt caption and returns a dictionary keyed by SQLite column names for template-loading workflows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Prompt caption or object name used by the operation. |
required |
Returns:
| Type | Description |
|---|---|
Dict[str, Any] | None
|
dict[str, Any] | None: Return value produced by the operation. |
Source code in app.py
insert_prompt ¶
Insert a prompt record.
Purpose
Writes a new prompt-template record to the configured SQLite Prompts table using the fields collected by the Prompt Engineering UI.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Dict[str, Any]
|
Prompt field dictionary. |
required |
Source code in app.py
update_prompt ¶
Update an existing prompt record.
Purpose
Updates an existing prompt-template record in the configured SQLite Prompts table using the selected ID value and edited field values.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pid
|
int
|
Prompt primary key value. |
required |
data
|
Dict[str, Any]
|
Prompt field dictionary. |
required |
Source code in app.py
delete_prompt ¶
Delete a prompt record.
Purpose
Removes a prompt-template record from the configured SQLite Prompts table using the selected ID value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pid
|
int
|
Prompt primary key value. |
required |
Source code in app.py
clear_active_prompt_metadata ¶
Clear metadata for the active system-instruction template.
Purpose
Resets the shared prompt identity fields without changing conversation, document, retrieval, or inference state.
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
This function performs its work through side effects and does not return a value. |
Source code in app.py
render_system_instructions ¶
Render a categorized System Instructions expander.
Purpose
Displays the shared editable system text with mode-specific category and ID-backed template selectors. The selector exposes only categories supported by Leeroy's local text model while keeping selection state independent between application modes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
category_key
|
str
|
Session-state key for the mode-specific category selector. |
required |
prompt_key
|
str
|
Session-state key for the mode-specific prompt selector. |
required |
clear_key
|
str
|
Widget key for the mode-specific clear button. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
This function performs its work through side effects and does not return a value. |
Source code in app.py
build_prompt ¶
Build a llama.cpp-compatible chat prompt.
Purpose
Combines system instructions, optional semantic retrieval context, basic document context, and in-memory chat history into the chat-template prompt consumed by the local llama.cpp model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_input
|
str
|
User message or constructed prompt text for the current generation turn. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
run_llm_turn ¶
run_llm_turn(
user_input: str,
temperature: float,
top_p: float,
repeat_penalty: float,
max_tokens: int,
stream: bool,
output: Any | None = None,
) -> str
Run one local LLM generation turn.
Purpose
Builds the shared prompt, loads the optional local llama.cpp model, and either streams generated tokens into a Streamlit placeholder or returns the completed response text for downstream workflows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_input
|
str
|
User message or constructed prompt text for the current generation turn. |
required |
temperature
|
float
|
Sampling temperature passed to the local model. |
required |
top_p
|
float
|
Nucleus sampling probability passed to the local model. |
required |
repeat_penalty
|
float
|
Repeat penalty passed to the local model. |
required |
max_tokens
|
int
|
Maximum number of generated tokens. |
required |
stream
|
bool
|
Whether to stream generated text into the UI. |
required |
output
|
Any | None
|
Optional Streamlit placeholder used for streaming output. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
create_connection ¶
Create a SQLite connection.
Purpose
Opens a connection to the configured application database used by chat history, prompt records, semantic embeddings, and data-management tables.
Returns:
| Type | Description |
|---|---|
Connection
|
sqlite3.Connection: Return value produced by the operation. |
Source code in app.py
list_tables ¶
List user-visible SQLite tables.
Purpose
Reads table names from sqlite_master and returns them in sorted order for data- management browsing, administration, and query workflows.
Returns:
| Type | Description |
|---|---|
List[str]
|
list[str]: Return value produced by the operation. |
Source code in app.py
create_schema ¶
Read a SQLite table schema.
Purpose
Retrieves PRAGMA table_info metadata for the selected table so data-management views can display columns, types, nullability, defaults, and primary-key indicators.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
SQLite table name. |
required |
Returns:
| Type | Description |
|---|---|
List[Tuple]
|
list[tuple]: Return value produced by the operation. |
Source code in app.py
read_table ¶
Read rows from a SQLite table.
Purpose
Builds a SELECT query for the requested table and optional pagination arguments, then returns the result as a pandas DataFrame for browsing and analysis.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
SQLite table name. |
required |
limit
|
int
|
Optional maximum number of rows to return. |
None
|
offset
|
int
|
Number of rows to skip before returning results. |
0
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: Return value produced by the operation. |
Source code in app.py
drop_table ¶
Drop a SQLite table when requested.
Purpose
Removes a selected table from the application database when the data-management administrator explicitly requests deletion.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
SQLite table name. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
Raised when the requested table name is empty. |
Source code in app.py
rename_table ¶
Rename an existing SQLite table.
Purpose
Attempts a native SQLite ALTER TABLE rename and falls back to a schema-preserving rebuild when needed. The fallback preserves column data and index SQL where possible.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
old_name
|
str
|
Existing table or column name. |
required |
new_name
|
str
|
Replacement table or column name. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
Raised when the source table definition is missing or malformed. |
Source code in app.py
1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 | |
rename_column ¶
Rename a column within a SQLite table.
Purpose
Attempts a native SQLite column rename and falls back to a schema-preserving table rebuild that keeps column order, data, defaults, nullability, and index SQL where possible.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table_name
|
str
|
SQLite table name. |
required |
old_name
|
str
|
Existing table or column name. |
required |
new_name
|
str
|
Replacement table or column name. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
Raised when the table definition is missing or the selected column does not exist. |
Source code in app.py
1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 | |
create_index ¶
Create a safe SQLite index.
Purpose
Validates the requested table and column against the active database schema, creates a sanitized index name, and creates the index using quoted identifiers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
SQLite table name. |
required |
column
|
str
|
SQLite column name. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
Raised when the table or column is not present in the active schema. |
Source code in app.py
apply_filters ¶
Apply an interactive DataFrame filter.
Purpose
Renders Streamlit filter controls and applies the selected comparison or containment operation to the provided DataFrame for data-management exploration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
DataFrame used by the operation. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: Return value produced by the operation. |
Source code in app.py
create_aggregation ¶
Render an interactive aggregation result.
Purpose
Renders Streamlit controls for selecting a numeric column and aggregation function, computes the selected aggregate, and displays the result as a metric.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
DataFrame used by the operation. |
required |
Source code in app.py
create_visualization ¶
Render an interactive DataFrame visualization.
Purpose
Renders Streamlit chart controls and uses Plotly Express to display histograms, bars, lines, scatter plots, boxes, pies, or correlation heatmaps from the provided DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
DataFrame used by the operation. |
required |
Source code in app.py
convert_dataframe ¶
Create a SQLite table from DataFrame columns.
Purpose
Maps pandas column dtypes to SQLite storage types and creates a table definition using normalized column names for imported tabular data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table_name
|
str
|
SQLite table name. |
required |
df
|
DataFrame
|
DataFrame used by the operation. |
required |
Source code in app.py
insert_data ¶
Insert DataFrame rows into SQLite.
Purpose
Normalizes DataFrame column names to SQLite-friendly identifiers and bulk-inserts row values into the specified table.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table_name
|
str
|
SQLite table name. |
required |
df
|
DataFrame
|
DataFrame used by the operation. |
required |
Source code in app.py
get_sqlite_type ¶
Map a pandas dtype to a SQLite type.
Purpose
Converts pandas dtype information into the SQLite storage class used by import and table-creation workflows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dtype
|
object
|
Pandas dtype or dtype-like object. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
create_custom_table ¶
Create a custom SQLite table.
Purpose
Validates table and column identifiers, translates user-provided column definitions into SQLite DDL, and creates the requested table when it does not exist.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table_name
|
str
|
SQLite table name. |
required |
columns
|
list
|
Column-definition dictionaries used to create the table. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
Raised when table or column definitions are invalid. |
Source code in app.py
is_safe_query ¶
Validate whether a SQL query is read-only.
Purpose
Checks query text before execution in the SQL console, allowing read-oriented statements and blocking destructive or mutating operations and multiple statements.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
query
|
str
|
SQL or search query text. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Return value produced by the operation. |
Source code in app.py
create_identifier ¶
Create a safe SQLite identifier.
Purpose
Normalizes arbitrary text into a SQLite-safe identifier by replacing invalid characters, ensuring a valid leading character, and rejecting empty results.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Prompt caption or object name used by the operation. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Raises:
| Type | Description |
|---|---|
ValueError
|
Raised when the input cannot be converted into a valid identifier. |
Source code in app.py
get_indexes ¶
List indexes for a SQLite table.
Purpose
Reads PRAGMA index_list metadata for the selected table so the data-management administration UI can display available indexes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
SQLite table name. |
required |
Returns:
| Type | Description |
|---|---|
List[Tuple]
|
list[tuple]: Return value produced by the operation. |
Source code in app.py
add_column ¶
Add a column to a SQLite table.
Purpose
Sanitizes the requested column name and executes an ALTER TABLE statement that appends the column with the requested SQLite type.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
SQLite table name. |
required |
column
|
str
|
SQLite column name. |
required |
col_type
|
str
|
SQLite column type. |
required |
Source code in app.py
create_profile_table ¶
Create a profile summary for a SQLite table.
Purpose
Reads the selected table into a DataFrame and computes per-column type, null percentage, distinct percentage, and numeric summary statistics for exploration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
SQLite table name. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: Return value produced by the operation. |
Source code in app.py
drop_column ¶
Drop a column from a SQLite table.
Purpose
Rebuilds the selected table without the specified column while preserving remaining column definitions, data, and indexes that do not depend on the dropped column.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
SQLite table name. |
required |
column
|
str
|
SQLite column name. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
Raised when the table, column, or resulting schema would be invalid. |
Source code in app.py
1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 | |
extract_text_from_bytes ¶
Extract text from uploaded document bytes.
Purpose
Attempts PDF extraction with PyMuPDF and falls back to defensive text decoding when PDF parsing fails. The helper supports document preview, summarization, and retrieval workflows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
file_bytes
|
bytes
|
Uploaded file bytes or PDF byte stream. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
route_document_query ¶
Route a document question through the chat pipeline.
Purpose
Builds a retrieval-grounded Document Q&A input from the user prompt and submits it to the shared local LLM generation workflow using current Streamlit runtime controls.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
prompt
|
str
|
Document question or prompt text. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
summarize_active_document ¶
Summarize the active document context.
Purpose
Constructs a structured summary prompt and routes the request through the Document Q&A workflow. The shared generation pipeline applies current system instructions once.
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
compute_fingerprint ¶
Compute a stable active-document fingerprint.
Purpose
Hashes active document names, byte lengths, and byte content hashes to support cache invalidation when uploaded Document Q&A inputs change.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
active_docs
|
List[str]
|
Active document names selected in session state. |
required |
doc_bytes
|
Dict[str, bytes]
|
Mapping of document names to uploaded bytes. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
extract_text ¶
Extract text from PDF bytes.
Purpose
Uses PyMuPDF to read text from each page of a PDF byte stream and returns the combined text for chunking and retrieval workflows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
file_bytes
|
bytes
|
Uploaded file bytes or PDF byte stream. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
load_sqlite_vec ¶
Load the sqlite-vec extension into a connection.
Purpose
Attempts to register sqlite-vec support on the provided SQLite connection so Document Q&A can use vector-table retrieval when available.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
conn
|
Connection
|
SQLite connection. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Return value produced by the operation. |
Source code in app.py
ensure_vec_schema ¶
Ensure the Document Q&A vector schema exists.
Purpose
Loads sqlite-vec and creates the docqna_vec virtual table for document embeddings when the extension is available in the runtime environment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim
|
int
|
Embedding dimension used by the vector table. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Return value produced by the operation. |
Source code in app.py
rebuild_index ¶
Build or refresh the Document Q&A vector index.
Purpose
Compares the active document fingerprint with cached state, extracts active document text, chunks content, creates embeddings, and stores vectors either in sqlite-vec or in the in-memory fallback rows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
embedder
|
SentenceTransformer
|
SentenceTransformer used to create embeddings. |
required |
Source code in app.py
2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 | |
retrieve_chunks ¶
Retrieve document chunks relevant to a query.
Purpose
Embeds the user query, refreshes the document index when needed, retrieves nearest chunks from sqlite-vec when available, and falls back to cosine similarity over cached vectors.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
query
|
str
|
SQL or search query text. |
required |
k
|
int
|
Number of retrieved chunks to include. |
6
|
Returns:
| Type | Description |
|---|---|
List[Tuple[str, str, float]]
|
list[tuple[str, str, float]]: Return value produced by the operation. |
Source code in app.py
2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 | |
build_docqna_input ¶
Build a retrieval-grounded Document Q&A prompt.
Purpose
Retrieves relevant document excerpts and combines them with the user question to form the grounded input submitted to the shared LLM pipeline. System instructions remain in the dedicated system-message path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_query
|
str
|
User question for the Document Q&A workflow. |
required |
k
|
int
|
Number of retrieved chunks to include. |
6
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Return value produced by the operation. |
Source code in app.py
load_embedder ¶
Load the sentence-transformer embedder.
Purpose
Creates the cached sentence-transformers model used for semantic indexing, Document Q&A retrieval, and fallback vector scoring.
Returns:
| Name | Type | Description |
|---|---|---|
SentenceTransformer |
SentenceTransformer
|
Return value produced by the operation. |
Source code in app.py
local_llm_enabled ¶
Determine whether local LLM support is enabled.
Purpose
Checks configuration and model-file availability to decide whether the local llama.cpp model should be loaded in the current runtime environment.
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Return value produced by the operation. |
Source code in app.py
load_llm ¶
Load the optional local llama.cpp model.
Purpose
Lazily imports llama_cpp and creates a cached Llama instance using the configured model path, context window, CPU thread count, and batch settings when local support is enabled.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ctx
|
int
|
Optional context window override. |
required |
threads
|
int
|
Optional CPU thread-count override. |
required |
Returns:
| Type | Description |
|---|---|
Any | None
|
Any | None: Return value produced by the operation. |
Source code in app.py
get_llm ¶
Return the available local llama.cpp model.
Purpose
Applies default context and thread settings when overrides are not supplied and returns the cached local model instance when enabled and loadable.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ctx
|
int | None
|
Optional context window override. |
None
|
threads
|
int | None
|
Optional CPU thread-count override. |
None
|
Returns:
| Type | Description |
|---|---|
Any | None
|
Any | None: Return value produced by the operation. |
Source code in app.py
get_embedder ¶
Return the cached embedding model.
Purpose
Returns the cached sentence-transformer model when it can be loaded and preserves a None fallback when embedding support is unavailable.
Returns:
| Type | Description |
|---|---|
SentenceTransformer | None
|
SentenceTransformer | None: Return value produced by the operation. |
Source code in app.py
get_conn ¶
Create a Prompt Engineering database connection.
Returns:
| Type | Description |
|---|---|
Connection
|
sqlite3.Connection: Open connection to the configured application database. |
reset_selection ¶
Reset the Prompt Engineering record selection and editor fields.
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
This function performs its work through side effects and does not return a value. |
Source code in app.py
load_prompt ¶
Load one prompt into the Prompt Engineering editor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pid
|
int
|
Prompt primary key value. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
This function performs its work through side effects and does not return a value. |
Source code in app.py
save_prompt_record ¶
Create or update the prompt currently shown in the editor.
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
This function performs its work through side effects and does not return a value. |
Source code in app.py
delete_prompt_record ¶
Delete the selected prompt record.
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
This function performs its work through side effects and does not return a value. |