← Back to projects
Gemma 4YOLO11OpenCVPythonFastAPINext.jsTypeScriptOllamaVercel

Cruz Watch

Real-time hazard detection for Santa Cruz, with an AI Agent that decides what's worth escalating

Cruz Watch architecture diagram

Inspiration

On June 10, 2026, two students from UC Berkeley were swept into the ocean by dangerous surf and rising tides along the Santa Cruz County coast near Panther/Bonny Doon beach, and died. Last year a classmate from my computational models class at UCSC drowned while cliff jumping. I also know surfers who have come close to drowning.

Every one of these happened at a spot with no lifeguard and nobody watching, and by the time someone on shore noticed and called for help, the window to save them was already closing. Santa Cruz has a short list of spots where this keeps happening. I built Cruz Watch because those spots should have something watching them.

What It Does

Cruz Watch landing page

  • Watches camera feeds at high-risk Santa Cruz spots, both coastal and downtown.
  • A lightweight YOLO detector runs on every frame and fires triggers: a person dwelling in a danger zone, a person going horizontal in water or on the ground, and motion anomalies.
  • On a trigger, Gemma 4 runs locally on device and receives the structured event plus that site's context — never raw video.
  • Gemma decides severity, decides whether to escalate, writes the dispatch report, and calls the (simulated) emergency dispatch endpoint.
  • The same trigger means different things at different sites, and Gemma responds differently: a wader at Seabright got LOW and no escalation, while a face-down floater got CRITICAL and an immediate marine rescue dispatch.
  • A live dashboard shows a video wall of 10 cams, a sidebar of Gemma-written incident reports, and per-camera detail with the agent's streamed reasoning.

The dashboard, the beach cam detail, and the urban cam detail:

Photo 1 of 3
1 / 3

How I Built It

  • YOLO11n + ByteTrack in a Python/FastAPI pipeline, with three reusable trigger primitives. Each site is just a JSON config with zones, thresholds, and context for the agent.
  • Gemma 4 (gemma4:e2b through Ollama) on my laptop GPU, streaming JSON-constrained output, with a hard timeout and a template fallback so a stall can never hang the pipeline.
  • Footage from real Santa Cruz cameras: the Steamer Lane surf cam archive, the WebCOOS Walton Lighthouse and Wharf cameras, and the harbor webcam recorded live.
  • Drowning and collapse demos use staged footage only: a lifeguard training drill and a research fall dataset, labeled as staged on every frame.
  • A Next.js dashboard on Vercel. Detector output is precomputed and stored, and every panel carries a provenance badge saying what is real and what is simulated.

Challenges I Ran Into

  • My 15GB laptop froze twice running the detector and the LLM together, so I serialized all GPU work and moved to a precompute architecture.
  • The motion anomaly trigger fired 98 times in 180 seconds on surf footage, because every wave ride is a speed spike. I disarmed it and documented why.
  • Cameras fought back: auth-walled streams I could only capture by polling a snapshot endpoint once per second, and PTZ cams that switch views and silently break zone polygons.
  • Finding ethical distress footage was hard: no real victims, no minors, staged drills only.
  • When I told Gemma a feed was a staged drill it correctly refused to escalate. That taught me the disclosure belongs in the UI, and the agent brief is policy.

What I Learned

  • Cheap CV on every frame, plus an LLM only on triggers, is what makes agentic AI realistic on edge hardware.
  • Context does the real work: identical detector events produce different severities and different responders purely from the site brief.
  • Honesty is a feature. Labeling what is real, archived, staged, and simulated made the whole project more credible, not less.
  • An agent's prompt is not a description, it is policy.

What's Next

  • The real edge unit: Raspberry Pi / Jetson Orin, RGB + thermal camera, solar power — roughly $500 to $800 per site for about 16 target locations.
  • Partnerships for live access to the cameras I used as archives.
  • Real open-water distress detection research, multilingual alerts, and a pilot with the county.

The proof-of-concept edge unit — what runs on the device, and the parts it is built from:

Photo 1 of 2
1 / 2

Built With

Gemma 4, Ollama, Python, FastAPI, YOLO11, ByteTrack, OpenCV, PyTorch, Next.js, TypeScript, Tailwind, FFmpeg, and Vercel.