Skip to content
Work

Backend · Scraping·2026

Site Scope

Turns a university's admissions pages into a structured dossier: deadlines, test requirements, fees, funding and contacts, with every field traceable back to the page it came from.

Stack
TypeScriptNode.jsPlaywrightVite
Links
Source
Admissions dossierDeep scan · 5 pages
Deadline
src ↗
English test
src ↗
Application fee
src ↗
Funding
src ↗
Contacts
src ↗
.md.csv

The problem

Researching graduate programmes means reading dozens of admissions sites that all hide the same twelve facts in different places. I wanted the facts, not the browsing.

How it works

  • Renders pages in Playwright Chromium and simulates scrolling to surface lazy-loaded content, falling back to plain HTTP when a browser isn't available.
  • Strips navigation, headers, footers, sidebars, forms and scripts before analysing text, so the extractor reads content, not chrome.
  • Classifies links into application portals, official documents, faculty contacts, financial aid and internal subpages.
  • Quick Scan reads the current page; Deep Scan follows up to five relevant subpages.

A deliberate choice

Extraction is deterministic and source-grounded rather than an LLM summary. When a deadline matters, you want to know exactly where it came from. The normalised output is shaped so a model can sit on top later without touching the scraping layer.

Next case studyBranch Performance Dashboards

Work with me

Hiring for a backend, mobile or cloud role, or have something you want built? Tell me what you have in mind.

Send me a note