diff --git a/ANALYSIS.md b/ANALYSIS.md index f7ada86..33447aa 100644 --- a/ANALYSIS.md +++ b/ANALYSIS.md @@ -19,7 +19,7 @@ | `frontend/src/components/sidebar/Sidebar.css` | Sidebar layout, hover indicators, chat item delete, upgrade card, user info | | `frontend/src/components/chat-history/ChatHistory.jsx` | Welcome screen with bento suggestions OR message list OR selector flow | | `frontend/src/components/chat-history/ChatHistory.css` | Welcome header, bento grid, info panel, message list | -| `frontend/src/components/message/Message.jsx` | Single message bubble (user or assistant), action buttons, typing dots | +| `frontend/src/components/message/Message.jsx` | Single message bubble (user or assistant), markdown+math rendering for assistant, action buttons, typing dots | | `frontend/src/components/message/Message.css` | Bubble styling, avatar, actions, typing animation | | `frontend/src/components/chat-input/ChatInput.jsx` | Textarea with auto-expand, mic, send button, attachment | | `frontend/src/components/chat-input/ChatInput.css` | Input bar styling, send button states, disclaimer | @@ -630,6 +630,12 @@ if (XSS_PATTERN.test(message.text)) { - Added regression tests: `backend/test/chatValidation.test.js` and `frontend/test/appState.test.js`. - Backend tests, backend syntax checks, frontend regression test, and Vite production build pass. - Chrome MCP verified the sign-in screen renders after fixing a missing `useAuth` import exposed by the browser console. +- Cross-user session leak fixed (token cache keyed by full token — see CVE-2026-011). +- AI connected via OpenRouter; streaming verified in the browser with both user+assistant messages persisted. +- Assistant replies render as markdown + KaTeX math; verified in the headed browser. +- Current user message is included in the AI context (persisted before streaming). +- Session continuity fixed: backend emits `session` id in SSE; follow-ups reuse the same session. +- Rolling conversation summary implemented and persisted (bounded context + long-range memory). ### Database status @@ -639,24 +645,56 @@ if (XSS_PATTERN.test(message.text)) { --- +## Data Storage & Conversation Context Model (2026-08-19) + +### Q: Each message is a new row — is that sustainable? + +**A (yes):** Each message is one row in `public.messages` (`id`, `session_id`, `role`, `text`, `created_at`), written append-only. This is the standard chat design (Slack/OpenAI-style). It is sustainable because: +- Storage is tiny (average turn is ~KB scale; a million messages ≈ a few GB). +- `idx_messages_session_id` keeps per-session history reads fast even with many rows. +- Append-only is desirable: simple pagination/partial loads, immutable history, easy to add edit/regenerate later. +- One-row-per-session-with-big-text-blob would be worse (can't load partial history, bad indexing). +- DB only needs attention at scale: partitioning by date/user past ~10s of millions of rows, or archiving old sessions. + +### The real scaling concern: AI context window (the "memory" issue) + +- The DB is fine, but on **every** `/api/chat` call the backend loads the **entire** message history and sends it **all** to the model. +- Token count (and therefore **cost and latency**) grows with conversation length; eventually it exceeds the model's context limit and the call fails. +- **This is the lever worth engineering** — not the database. + +### Blocked-on note + +Before context work has value, the **frontend session-continuity bug** must be fixed (see What's Next #1): follow-ups in a fresh chat currently create a new session with no history, so long-running context never even accumulates. + +--- + ## What's Next -### 1. Frontend chat-flow design gap -- Backend does not return the new `session.id` to the client (SSE only sends `chunk`/`done`/`error`). -- Sidebar refresh (`loadSessions()`) only runs on the successful SSE path, so an AI error leaves the sidebar stale even though the session row is created. -- Decide: return `session.id` in SSE + refresh sidebar on error path (Option 1), or only persist sessions after a successful AI response (Option 2). +### 1. Frontend session continuity — DONE +- Backend now emits a `{ type: "session", id }` SSE event when a conversation is created/loaded; the frontend sets `activeChat` on it, so follow-ups reuse the same session and history accumulates. +- Remaining minor: sidebar refresh (`loadSessions()`) only runs on the successful SSE path — an AI error still leaves the sidebar stale even though the session row was created. -### 2. Frontend authentication handling -- Handle protected-route HTTP 401 responses by returning to the sign-in modal. -- Persist selector state per user. +### 1b. Rolling conversation summary (B+D) — implemented 2026-08-19 +- `sessions.summary` column (migration `20260819150000_add_sessions_summary.sql`). +- On each `/api/chat`, when a session exceeds `RECENT_LIMIT` (10) messages, older messages are folded into a persisted summary via a non-streaming model call (`summarizeConversation`), stored with `setSummary`. +- The model receives `system(+summary) + last 10 messages` (`buildMessages`), so context stays bounded while long-range context persists. +- Verified end-to-end: fresh chat emits session id; a 12-message session folded its overflow into a stored summary; follow-up answered with bounded context. +- Tests: `backend/test/aiMessages.test.js` (capping + summary injection). + +### 2. Frontend rendering — fixed +- Assistant replies now render via `react-markdown` (headings, bold, italics, lists, blockquotes, code, tables, `---`) and math (`$H_2O$`, `$CO_2$`) via KaTeX. User messages stayed plain text. `ReactMarkdown` escapes raw HTML by default, so stored-XSS (CVE-2026-009) stays closed while rendering rich content. +- Protected-route 401s still return to the sign-in modal. +- Selector state persists per user. ### AI Provider — OpenRouter (connected 2026-08-19) - Backend uses the OpenAI-compatible path (`AI_PROVIDER=openai`) pointed at OpenRouter. - `OPENAI_BASE_URL=https://openrouter.ai/api/v1` (added to `ai.js` as an optional override). -- Model: `~deepseek/deepseek-v4-flash-latest` (the `~` prefix is REQUIRED — it is OpenRouter's "redirect to latest" alias; the non-tilde `deepseek/deepseek-v4-flash-latest` is an invalid model ID). -- Verified end-to-end: `/api/chat` streams a real reply and persists both `user` and `assistant` messages (lost messages bug from the no-AI path is gone). -- Note: this is a reasoning model — it streams thinking in `delta.reasoning` before `delta.content`; the app reads `delta.content`, so there is a brief (sub-second) gap while reasoning, then the answer streams. +- Model: `qwen/qwen3.7-flash` (final choice; earlier tried `~deepseek/deepseek-v4-flash-latest` and `upstage/solar-pro4`). + - OpenRouter "latest" aliases use a required `~` prefix (e.g. `~deepseek/...`); the non-tilde `-latest` is an invalid ID. +- Reasoning disabled for OpenRouter (`reasoning: { enabled: false }` in `streamOpenai`/`getChatResponse`) so `delta.content` streams immediately instead of sitting empty during the thinking phase (which made the chat look blank). +- Verified end-to-end: `/api/chat` streams a real reply and persists both `user` and `assistant` messages. +- The user message is now persisted **before** streaming, so the model sees the current question (fixes canned-greeting responses) and the user message survives AI failures. **Env (non-secret):** wire via `backend/.env`, and the placeholders are in `backend/.env.example`. diff --git a/backend/src/routes/chat.js b/backend/src/routes/chat.js index aede4a9..3096769 100644 --- a/backend/src/routes/chat.js +++ b/backend/src/routes/chat.js @@ -1,6 +1,6 @@ import express from "express"; -import { streamChatResponse } from "../services/ai.js"; -import { createSession, addMessage, getMessages, getSession, listSessions } from "../services/db.js"; +import { RECENT_LIMIT, streamChatResponse, summarizeConversation } from "../services/ai.js"; +import { createSession, addMessage, getMessages, getSession, listSessions, setSummary } from "../services/db.js"; import { isOwnedSession, validateChatInput } from "./chatValidation.js"; const router = express.Router(); @@ -33,6 +33,23 @@ router.post("/", async (req, res) => { res.setHeader("Connection", "keep-alive"); res.setHeader("X-Accel-Buffering", "no"); + // Tell the client which session this conversation belongs to, so follow-ups + // reuse the same session and history accumulates. + res.write(`data: ${JSON.stringify({ type: "session", id: currentChatId })}\n\n`); + + // Bounded context: fold older messages into a persistent summary, keep the + // last RECENT_LIMIT raw. + let summary = session.summary || ""; + if (chatMessages.length > RECENT_LIMIT) { + const overflow = chatMessages.slice(0, chatMessages.length - RECENT_LIMIT); + try { + summary = await summarizeConversation(session.summary, overflow); + await setSummary(currentChatId, summary); + } catch (err) { + console.error("Summarize error:", err.message); + } + } + let assistantText = ""; await streamChatResponse( chatMessages, @@ -50,7 +67,8 @@ router.post("/", async (req, res) => { () => { res.write(`data: ${JSON.stringify({ type: "error", message: "AI service error" })}\n\n`); res.end(); - } + }, + summary ); res.write(`data: ${JSON.stringify({ type: "done" })}\n\n`); diff --git a/backend/src/services/ai.js b/backend/src/services/ai.js index 7a48045..89ecc6d 100644 --- a/backend/src/services/ai.js +++ b/backend/src/services/ai.js @@ -39,21 +39,57 @@ function buildSystemPrompt(grade, subject, chapter) { return context.join("\n"); } +// Keep the last N messages raw; older ones are folded into a session summary. +export const RECENT_LIMIT = 10; + /** * Build the messages array for the AI call. + * Older context is passed as a `summary` (folded into the system prompt) and + * only the last RECENT_LIMIT messages are sent, keeping the request bounded. */ -function buildMessages(chatMessages, grade, subject, chapter) { - const systemMsg = { - role: "system", - content: buildSystemPrompt(grade, subject, chapter), - }; +export function buildMessages(chatMessages, grade, subject, chapter, summary) { + const systemPrompt = buildSystemPrompt(grade, subject, chapter); + const summaryText = summary && summary.trim() ? summary.trim() : ""; + const systemContent = summaryText + ? `${systemPrompt}\n\nEarlier in this conversation:\n${summaryText}` + : systemPrompt; - const userMessages = chatMessages.map((m) => ({ + const recent = chatMessages.slice(-RECENT_LIMIT).map((m) => ({ role: m.role === "user" ? "user" : "assistant", content: m.text, })); - return [systemMsg, ...userMessages]; + return [{ role: "system", content: systemContent }, ...recent]; +} + +/** + * Fold older messages (plus any existing summary) into a concise running + * summary. Non-streaming. Returns the new summary text. + */ +export async function summarizeConversation(existingSummary, messages) { + const isOpenRouter = (process.env.OPENAI_BASE_URL || "").includes("openrouter.ai"); + const body = existingSummary && existingSummary.trim() + ? `Prior summary:\n${existingSummary.trim()}\n\nNew messages:\n${messages.map((m) => `${m.role}: ${m.text}`).join("\n")}` + : messages.map((m) => `${m.role}: ${m.text}`).join("\n"); + + const res = await openai.chat.completions.create({ + model: process.env.OPENAI_MODEL || "gpt-4o", + messages: [ + { + role: "system", + content: + "You maintain a concise running summary of an ongoing tutoring conversation. " + + "Combine the prior summary (if any) with the new messages into one brief summary (max ~150 words). " + + "Keep the student's grade, subject, chapter, and any goals or open questions. " + + "Do not answer questions — only summarize.", + }, + { role: "user", content: body }, + ], + max_tokens: 512, + ...(isOpenRouter ? { reasoning: { enabled: false } } : {}), + }); + + return (res.choices?.[0]?.message?.content || "").trim(); } /** @@ -229,8 +265,8 @@ async function streamGoogle(messages, onChunk, onError) { * @param {Function} onError - Called on error * @returns {Promise} The full response text */ -export async function streamChatResponse(chatMessages, grade, subject, chapter, onChunk, onError) { - const messages = buildMessages(chatMessages, grade, subject, chapter); +export async function streamChatResponse(chatMessages, grade, subject, chapter, onChunk, onError, summary) { + const messages = buildMessages(chatMessages, grade, subject, chapter, summary); switch (provider) { case "anthropic": @@ -246,8 +282,8 @@ export async function streamChatResponse(chatMessages, grade, subject, chapter, /** * Get a non-streaming response (fallback for simple queries). */ -export async function getChatResponse(chatMessages, grade, subject, chapter) { - const messages = buildMessages(chatMessages, grade, subject, chapter); +export async function getChatResponse(chatMessages, grade, subject, chapter, summary) { + const messages = buildMessages(chatMessages, grade, subject, chapter, summary); let response; switch (provider) { diff --git a/backend/src/services/db.js b/backend/src/services/db.js index 43f1eca..4855c5f 100644 --- a/backend/src/services/db.js +++ b/backend/src/services/db.js @@ -119,3 +119,17 @@ export async function clearMessages(sessionId, token) { if (error) throw new Error(`Failed to clear messages: ${error.message}`); return true; } + +/** + * Persist the rolling conversation summary on a session. + */ +export async function setSummary(sessionId, summary, token) { + const supabase = token ? createUserClient(token) : supabaseAdmin; + const { error } = await supabase + .from('sessions') + .update({ summary, updated_at: new Date().toISOString() }) + .eq('id', sessionId); + + if (error) throw new Error(`Failed to set summary: ${error.message}`); + return true; +} diff --git a/backend/test/aiMessages.test.js b/backend/test/aiMessages.test.js new file mode 100644 index 0000000..ff55828 --- /dev/null +++ b/backend/test/aiMessages.test.js @@ -0,0 +1,33 @@ +import test from "node:test"; +import assert from "node:assert/strict"; +import { buildMessages, RECENT_LIMIT } from "../src/services/ai.js"; + +function msg(role, i) { + return { id: String(i), session_id: "s", role, text: `${role}-${i}`, created_at: new Date(i).toISOString() }; +} + +test("caps message history to the recent window", () => { + const history = Array.from({ length: RECENT_LIMIT + 5 }, (_, i) => msg(i % 2 ? "assistant" : "user", i)); + const built = buildMessages(history, "Grade 10", "math", "c", ""); + // system + RECENT_LIMIT messages + assert.equal(built.length, 1 + RECENT_LIMIT); + assert.equal(built[0].role, "system"); + // the newest message is included + assert.equal(built[built.length - 1].content, "user-" + (RECENT_LIMIT + 4)); + // the oldest message is dropped + assert.ok(!built.some((m) => m.content === "user-0")); +}); + +test("injects the summary into the system prompt when present", () => { + const history = [msg("user", 0), msg("assistant", 1)]; + const built = buildMessages(history, "Grade 6", "english", "c5", "Student is in Grade 6 English."); + assert.equal(built[0].role, "system"); + assert.ok(built[0].content.includes("Student is in Grade 6 English.")); + // still includes the messages + assert.equal(built.length, 3); +}); + +test("does not inject an empty summary", () => { + const built = buildMessages([msg("user", 0)], "g", "s", "c", ""); + assert.ok(!built[0].content.includes("Earlier in this conversation:")); +}); diff --git a/frontend/src/App.jsx b/frontend/src/App.jsx index a5bdb8c..6010d1e 100644 --- a/frontend/src/App.jsx +++ b/frontend/src/App.jsx @@ -347,7 +347,11 @@ function App() { try { const data = JSON.parse(jsonStr); - if (data.type === "chunk") { + if (data.type === "session") { + // Remember which session this conversation belongs to so + // follow-up messages reuse it (history stays together). + setActiveChat(data.id); + } else if (data.type === "chunk") { assistantText += data.content; setMessages((prev) => prev.map((m) => diff --git a/supabase/migrations/20260819150000_add_sessions_summary.sql b/supabase/migrations/20260819150000_add_sessions_summary.sql new file mode 100644 index 0000000..f93ad2a --- /dev/null +++ b/supabase/migrations/20260819150000_add_sessions_summary.sql @@ -0,0 +1,6 @@ +-- Add a rolling conversation summary column to sessions. +-- Older messages get folded into this summary so the AI context stays bounded +-- (model receives summary + last N messages) while long-range context persists. + +ALTER TABLE public.sessions + ADD COLUMN IF NOT EXISTS summary text; \ No newline at end of file