2a9111c58c
server/llm/ollama_cloud.py wraps the Ollama Cloud direct API (https://ollama.com/api/chat + bearer, stream=True) behind LLMProvider. chat() streams LLMStreamChunk (is_first flag for TTFT measurement); chat_full() accumulates for the debrief / branch classifier (offline). Two models: gemma4:cloud (roleplay_model) + deepseek-v4-flash:cloud (debrief_model, no_think mode for latency, D-020). Resolves R6 — the adapter confirms the direct API + bearer path; a live first-token confirmation is pending the R3 probe with a real key. Graceful no-key degradation (no chunks, no crash). 6 unit tests pass (mocked httpx streaming response + env model selection + chat_full accumulation). ---ci--- phase: 1 milestone: v0.1 plan: 02 task: 02-03 status: execute persona: backend-engineer requirements: covered: [REQ-LLM-01, REQ-LLM-02] ---/ci---