The biggest fix: change-detection and the fast-path cache lookup were
scoped only by competitor_id, not by which specific trip page was
scraped. Since trip-intent matching lets a PM scrape a different URL per
run, get_last_hash_for_url() compared each new scrape against whatever
page happened to be hashed most recently for that competitor — almost
always a different page — so change_detected never settled back to
False and the true no-scrape fast path was unreachable in practice.
Scoped scrape_runs/products lookups by (competitor_id, url) instead, and
threaded a cleaned confirmed_url consistently through needs_refresh(),
detect_change(), and analyse_from_cache(). Verified end-to-end with a
regression test reproducing the exact reported scenario.
New: a per-seat, tier-capped trip watchlist (1 trip/seat at Analyse).
The nightly cron now re-scrapes only the specific trips a PM chose to
track instead of every competitor's homepage (a page already flagged
low-signal by is_generic_homepage()) — bounded, predictable volume, and
every resulting alert is about a trip someone actually cares about.
Reuses the existing confirmed_url/per-url caching machinery with zero
changes needed to jobs.py's job loop. Also fixes a real gap:
auth_service.py resolved the caller's active seat only to validate the
JWT and then discarded it — every route before this only ever needed
tenant_id. get_seat_from_request() now returns it.
Also fixed the Firecrawl /map timeout (15s → 30s; a large real site's
sitemap fetch took ~21s and was silently read as "no match found"), a
job left stuck at "Stopping…" forever after a worker process was killed
mid-run (status endpoint now self-heals via the underlying RQ job
state), and several results-table issues: hardcoded "G Adventures" in
the extraction prompt and a literal "G" in the comments column label
(replaced with the real tenant name), non-clickable trip links, a
cache-source badge that was silently clipped on most columns, a missing
view into the client's own scraped trip data, and bolding for
Strengths/Weaknesses/Differences/Pricing in the comments text.
370+ new/updated backend tests; full suite green.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The highest-impact fix: run_research_job() only ever read entry['id'] from
confirmed competitors, silently discarding the Firecrawl-matched or
manually-swapped URL — every scrape fell back to the competitor's bare
root domain regardless of what was confirmed. This explains most of the
"unpredictable/wrong data" behavior seen against the old app.py. Threaded
confirmed_url through jobs.py -> analysis_service.py so the actually
confirmed page is what gets scraped.
Other fixes bundled in this pass:
- Proxy layer: Webshare rotating-gateway country code lowercased (was
silently failing auth), Bright Data port corrected to 33335 (current
cert), dedicated-IP candidate rotation added before falling back to the
shared gateway/Bright Data escalation.
- Trip-intent matching: score_and_rank_urls() now hard-gates on
destination keyword and excludes bare root-domain candidates, instead
of letting duration/category signals alone push a wrong-country or
homepage URL over the confidence threshold.
- force_refresh cache-bypass toggle added (opt-in, collapsed "Advanced
options"), plus a "Skip this competitor" affordance for unmatched URLs.
- De-hardcoded "G Adventures" (the hackathon reference client) out of the
extraction prompt — now uses the real tenant's company name.
- RunHistory/Apps-Script "latest results" queries filtered to
triggered_by = "manual" so per-competitor content-hash bookkeeping rows
no longer drown out real batch runs.
- Dashboard fixes: stuck-state bug on revisiting the Analyse page,
AccountPage/BattlecardsPage restacked to single-column layout, Alerts
moved to a header bell dropdown, invisible Export/Run Analysis icons
(fill -> stroke), tour popup width mismatch, horizontal-scroll bug on
cards (missing min-width: 0 in flex layout).
- New PocketBase migrations for scrape_runs (content_hash fields,
tab_type, created_by_email) and products (departures, start_days).
335 backend tests passing.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Phase 5 — Automation:
- Daily scrape cron, daily digest email, and Monday brief n8n workflows
- price_history_repository.py + price diffing/is_new detection wired
into analysis_service's refresh path
- Numeric unread-count badge on the topbar bell (was a plain dot)
Change detection architecture correction:
- CLAUDE.md's Firecrawl Monitor design doesn't match the real
self-hosted product: no /v1/monitor in v1, the real /v2/monitor is
cron-scheduled and credit-metered (Supabase-gated), and self-hosted
Firecrawl is a 3-service stack (api/worker/playwright-service), not
the single container originally spec'd
- Replaced with content-hash comparison computed during the scrape
itself (core/scraper.py's compute_content_hash/normalise_html,
analysis_service's detect_change()) — no external dependency, and
the refresh path now skips LLM extraction entirely when a scrape
finds no content diff
- Fixed a real bug in the first pass of this migration: change_detected
was set and then immediately reset within the same function call,
making it unobservable to the CompetitorsPage badge, the Apps Script
sidebar, and the n8n cron filter. It now stays visible until the
next scrape confirms nothing further changed.
- Added a composite index on scrape_runs(competitor_id, content_hash,
started_at) backing the per-competitor recency lookup
- Docker Compose's firecrawl service replaced with the real
firecrawl-api/firecrawl-worker/firecrawl-playwright stack, gated
behind an opt-in profile
- Added CORS (absent from the original spec entirely) so the
browser-based dashboard can reach the API cross-origin
Phase 6 — Semantic trip matching:
- client_trips_repository.py — this collection existed in schema but
nothing ever populated it; comparable-trip matching could never
produce a real match without it
- embedding_service.save_client_trip() upserts a client_trips record
from every /research submission's trip-intent description, with
stale-embedding invalidation when destination/duration change
- /internal/embed-products now embeds client_trips as well as
competitor products; /internal/match-comparable reads both sides
from PocketBase instead of expecting client_trips in the request
body, so it's callable unattended by n8n
- comparable_matches upserts on (client_product, competitor_product)
instead of always creating — first_matched/last_matched only make
sense if repeat weekly matches update in place, and a PM's dismissed
match now survives re-matching
- embedding_and_matching_cron.json — Sunday night n8n workflow
249 tests passing, 96% coverage.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>