Commit Graph

4 Commits

Author SHA1 Message Date
JasonFraser a25909a459 Add per-seat trip watchlist, fix cross-URL caching bug, and improve results table UX
The biggest fix: change-detection and the fast-path cache lookup were
scoped only by competitor_id, not by which specific trip page was
scraped. Since trip-intent matching lets a PM scrape a different URL per
run, get_last_hash_for_url() compared each new scrape against whatever
page happened to be hashed most recently for that competitor — almost
always a different page — so change_detected never settled back to
False and the true no-scrape fast path was unreachable in practice.
Scoped scrape_runs/products lookups by (competitor_id, url) instead, and
threaded a cleaned confirmed_url consistently through needs_refresh(),
detect_change(), and analyse_from_cache(). Verified end-to-end with a
regression test reproducing the exact reported scenario.

New: a per-seat, tier-capped trip watchlist (1 trip/seat at Analyse).
The nightly cron now re-scrapes only the specific trips a PM chose to
track instead of every competitor's homepage (a page already flagged
low-signal by is_generic_homepage()) — bounded, predictable volume, and
every resulting alert is about a trip someone actually cares about.
Reuses the existing confirmed_url/per-url caching machinery with zero
changes needed to jobs.py's job loop. Also fixes a real gap:
auth_service.py resolved the caller's active seat only to validate the
JWT and then discarded it — every route before this only ever needed
tenant_id. get_seat_from_request() now returns it.

Also fixed the Firecrawl /map timeout (15s → 30s; a large real site's
sitemap fetch took ~21s and was silently read as "no match found"), a
job left stuck at "Stopping…" forever after a worker process was killed
mid-run (status endpoint now self-heals via the underlying RQ job
state), and several results-table issues: hardcoded "G Adventures" in
the extraction prompt and a literal "G" in the comments column label
(replaced with the real tenant name), non-clickable trip links, a
cache-source badge that was silently clipped on most columns, a missing
view into the client's own scraped trip data, and bolding for
Strengths/Weaknesses/Differences/Pricing in the comments text.

370+ new/updated backend tests; full suite green.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-04 22:03:54 -04:00
JasonFraser ad1d7f1815 Fix critical scrape-targeting bug and harden proxy, matching, and dashboard reliability
The highest-impact fix: run_research_job() only ever read entry['id'] from
confirmed competitors, silently discarding the Firecrawl-matched or
manually-swapped URL — every scrape fell back to the competitor's bare
root domain regardless of what was confirmed. This explains most of the
"unpredictable/wrong data" behavior seen against the old app.py. Threaded
confirmed_url through jobs.py -> analysis_service.py so the actually
confirmed page is what gets scraped.

Other fixes bundled in this pass:
- Proxy layer: Webshare rotating-gateway country code lowercased (was
  silently failing auth), Bright Data port corrected to 33335 (current
  cert), dedicated-IP candidate rotation added before falling back to the
  shared gateway/Bright Data escalation.
- Trip-intent matching: score_and_rank_urls() now hard-gates on
  destination keyword and excludes bare root-domain candidates, instead
  of letting duration/category signals alone push a wrong-country or
  homepage URL over the confidence threshold.
- force_refresh cache-bypass toggle added (opt-in, collapsed "Advanced
  options"), plus a "Skip this competitor" affordance for unmatched URLs.
- De-hardcoded "G Adventures" (the hackathon reference client) out of the
  extraction prompt — now uses the real tenant's company name.
- RunHistory/Apps-Script "latest results" queries filtered to
  triggered_by = "manual" so per-competitor content-hash bookkeeping rows
  no longer drown out real batch runs.
- Dashboard fixes: stuck-state bug on revisiting the Analyse page,
  AccountPage/BattlecardsPage restacked to single-column layout, Alerts
  moved to a header bell dropdown, invisible Export/Run Analysis icons
  (fill -> stroke), tour popup width mismatch, horizontal-scroll bug on
  cards (missing min-width: 0 in flex layout).
- New PocketBase migrations for scrape_runs (content_hash fields,
  tab_type, created_by_email) and products (departures, start_days).

335 backend tests passing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-04 17:33:05 -04:00
JasonFraser e0de2df6bf Phase 5 + 6: automation, content-hash change detection, semantic trip matching
Phase 5 — Automation:
- Daily scrape cron, daily digest email, and Monday brief n8n workflows
- price_history_repository.py + price diffing/is_new detection wired
  into analysis_service's refresh path
- Numeric unread-count badge on the topbar bell (was a plain dot)

Change detection architecture correction:
- CLAUDE.md's Firecrawl Monitor design doesn't match the real
  self-hosted product: no /v1/monitor in v1, the real /v2/monitor is
  cron-scheduled and credit-metered (Supabase-gated), and self-hosted
  Firecrawl is a 3-service stack (api/worker/playwright-service), not
  the single container originally spec'd
- Replaced with content-hash comparison computed during the scrape
  itself (core/scraper.py's compute_content_hash/normalise_html,
  analysis_service's detect_change()) — no external dependency, and
  the refresh path now skips LLM extraction entirely when a scrape
  finds no content diff
- Fixed a real bug in the first pass of this migration: change_detected
  was set and then immediately reset within the same function call,
  making it unobservable to the CompetitorsPage badge, the Apps Script
  sidebar, and the n8n cron filter. It now stays visible until the
  next scrape confirms nothing further changed.
- Added a composite index on scrape_runs(competitor_id, content_hash,
  started_at) backing the per-competitor recency lookup
- Docker Compose's firecrawl service replaced with the real
  firecrawl-api/firecrawl-worker/firecrawl-playwright stack, gated
  behind an opt-in profile
- Added CORS (absent from the original spec entirely) so the
  browser-based dashboard can reach the API cross-origin

Phase 6 — Semantic trip matching:
- client_trips_repository.py — this collection existed in schema but
  nothing ever populated it; comparable-trip matching could never
  produce a real match without it
- embedding_service.save_client_trip() upserts a client_trips record
  from every /research submission's trip-intent description, with
  stale-embedding invalidation when destination/duration change
- /internal/embed-products now embeds client_trips as well as
  competitor products; /internal/match-comparable reads both sides
  from PocketBase instead of expecting client_trips in the request
  body, so it's callable unattended by n8n
- comparable_matches upserts on (client_product, competitor_product)
  instead of always creating — first_matched/last_matched only make
  sense if repeat weekly matches update in place, and a PM's dismissed
  match now survives re-matching
- embedding_and_matching_cron.json — Sunday night n8n workflow

249 tests passing, 96% coverage.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-02 11:23:23 -04:00
JasonFraser f7aaafa533 Phase 1: data foundation — schema, repositories, services, routes, jobs
Implements CLAUDE.md Build Order Phase 1 end to end:
- PocketBase migration covering all 15 MVP-scope collections
- Repository layer (11 repos) with real PocketBase-backed + in-memory
  mock implementations for every collection
- Service layer (licence, trip-finder, battlecard, embedding, brief,
  alert, analysis, onboarding, billing, auth) — Phase 2 scraping/
  extraction dependencies are lazy-imported so this layer is fully
  testable ahead of the app.py refactor
- Flask API (api.py) covering every MVP route from the spec, with
  X-API-Key auth, Flask-Limiter rate limits, and structured JSON errors
- RQ job queue (jobs.py) with per-competitor failure isolation and
  duration-based Gotify alerting
- 142 tests, 95% coverage, 100% on licence validation / seat
  management / Stripe webhook dispatch per CLAUDE.md's testing floor

/sheet/push and /sheet/tabs intentionally return structured 501s until
the Apps Script Web App deployment exists (Phase 4).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-01 12:21:04 -04:00