Based on the retrieved codebase memory, here's how query routing works: Routing Logic: 1. Explicit Override: The `_should_use_cloud` function first checks for an explicit `use_cloud` parameter that can force cloud mode (DeepSeek) regardless of other conditions #main.py:986-1005. 2. Keyword Detection: If the user query contains keywords like "architecture", "explain", "how does", or "overview", it routes to DeepSeek's cloud mode #config.py:43-44. 3. Word Count Threshold: Queries exceeding 40 words automatically route to cloud/DeepSeek via `_should_use_cloud`'s word count check #main.py:992-994, with this threshold defined in config as `CLOUD_ESCALATION_WORD_COUNT = 40` #config.py:37. 4. Default Mode: Falls back to whatever is set as `DEFAULT_MEMORY_MODE`, which defaults to "hybrid" unless explicitly changed via the `MEMORY_MODE` environment variable #config.py:31-32, with this function returning based on whether that mode equals "cloud" #main.py:996-998. Important Note: The retrieved context does not contain explicit information about Ollama's role in the routing decision or how local model queries are constructed when cloud mode is NOT selected.