Backend · Scraping·2026
Site Scope
Turns a university's admissions pages into a structured dossier: deadlines, test requirements, fees, funding and contacts, with every field traceable back to the page it came from.
- Stack
- TypeScriptNode.jsPlaywrightVite
- Links
- Source
Admissions dossierDeep scan · 5 pages
- Deadline
- src ↗
- English test
- src ↗
- Application fee
- src ↗
- Funding
- src ↗
- Contacts
- src ↗
.md.csv
The problem
Researching graduate programmes means reading dozens of admissions sites that all hide the same twelve facts in different places. I wanted the facts, not the browsing.
How it works
- Renders pages in Playwright Chromium and simulates scrolling to surface lazy-loaded content, falling back to plain HTTP when a browser isn't available.
- Strips navigation, headers, footers, sidebars, forms and scripts before analysing text, so the extractor reads content, not chrome.
- Classifies links into application portals, official documents, faculty contacts, financial aid and internal subpages.
- Quick Scan reads the current page; Deep Scan follows up to five relevant subpages.
A deliberate choice
Extraction is deterministic and source-grounded rather than an LLM summary. When a deadline matters, you want to know exactly where it came from. The normalised output is shaped so a model can sit on top later without touching the scraping layer.
Work with me
Hiring for a backend, mobile or cloud role, or have something you want built? Tell me what you have in mind.