Scores how well a college syllabus matches what employers are actually hiring for. Sample figures.
Alignment dashboard
Four numbers at the top, because a curriculum committee will not read a model report.
Alignment by subject (sample)
Course warnings
- Subject 04 — low coveragewarning
Teaches skills the market is no longer asking for.
- Subject 05 — no matched skillswarning
Either the description is too vague to extract from, or the subject genuinely sits outside current demand. The system cannot tell those apart, and says so.
Seven steps, one run
Kicked off from the report screen and streamed back as events, because the whole run takes minutes, not seconds.
Pipeline steps
- 1Scraping jobs
- 2Extracting job skills
- 3Extracting course skills
- 4Retraining ML models
- 5Generating alignment scores
- 6Final validation
- 7Creating PDF report
| Step | Function | Can be skipped |
|---|---|---|
| 1 | scrape_jobs_from_google_jobs | yes — run on stored postings |
| 2 | extract_skills_from_jobs | yes |
| 3 | extract_subject_skills_from_supabase | no |
| 4 | retrain_ml_models | yes — scoring can use the last model |
| 5 | compute_subject_scores_and_save | no |
| 6 | final_checking | no |
| 7 | generate_pdf_report | no |
Every step is a flag, so a run can scrape nothing and rescore everything — which is what you want when the question is about the syllabus, not the market. Steps yield between batches so a long scrape does not block the event stream.
What the market is asking for
Most in-demand skills (sample frequency)
Raw extraction produces the same skill five ways. Counts are normalized, then fuzzy-deduped above a similarity threshold, then folded through an alias map — without that, a syllabus looks misaligned simply because the postings spelled a skill differently.
Why the cleaning layer exists
- Normalizecasing, punctuation, plurals
- Fuzzy deduperatio threshold
Merges near-identical skill strings before they are counted.
- Alias foldcurated
Maps known equivalents onto one canonical term.
- Fuzzy membershipat match time
A course skill counts as covered if it is close enough to a market skill, not only if it is identical.
Where the syllabus lags the market
| Skill in demand | Covered? | Subjects | Demand |
|---|---|---|---|
| Sample skill A | covered | 2 | high |
| Sample skill B | partial | 1 | high |
| Sample skill C | missing | 0 | high |
| Sample skill D | missing | 0 | medium |
A missing high-demand skill is the one finding a department can act on in a single curriculum review, which is why the table sorts this way rather than by score.
Subject 03
Scores
- Alignment score
- 52%
- Coverage
- sample
- Average similarity
- sample
- Calculated at
- sample date
Skills
- Taught
- sample list
- In market
- sample list
- Matched
- sample list
- Missing
- sample skill C
Three numbers rather than one: coverage is how much of the market's demand the subject touches, average similarity is how closely it touches it, and the score combines them. A subject can score badly for either reason, and the fix differs.
Real model metrics belong in the case study, not in a walkthrough. These figures are placeholders.
What is actually trained
- Query quality modelfilters noise
Decides whether a scraped posting is worth scoring at all. Trained on logged queries and their yield, so the scraper gets better at asking.
- Subject success modelpredicts
Scores a subject against current demand from its extracted skill set.
- Sentence embeddingsmatching
Course descriptions and job requirements embedded into one space, which is what makes near-miss matching possible.
From posting to score
- 1scrape
- 2query-quality filter
- 3extract skills
- 4embed
- 5match to syllabus
- 6score and rank
- 7validate
Retraining is step 4 of the pipeline, not a separate ritual, so the models track the postings rather than the day they were first fitted.
The five tables
Browsable in the app, because the first question anyone asks about a score is where it came from.
data model courses · jobs
course_id / course_code / course_titletextcourse_descriptiontextwhat skills are extracted fromjob_id / title / company / locationtextdescription / requirementstextvia / source / matched_keywordtextwhich query found the postingposted_at / scraped_attimestampdata model job_skills · course_skills · course_alignment_scores_clean
job_skills / course_skillstext[]the extraction output, kept separate from the source rowdate_extracted_jobs / date_extracted_coursetimestampskills_taught / skills_in_markettext[]both sides of the comparison, stored with the resultscore / coverage / avg_similaritynumericcalculated_attimestampso a score can be traced to the data behind itThe alignment table stores both skill lists alongside the score. A committee can see what the system thought the subject taught, and disagree with that rather than with the number.
Generated report
- Syllabus / job alignment reportPDF
Generated at the end of a run and uploaded to storage, returned as a link rather than a download the browser has to hold.
- Per-subject breakdowntable
- Missing skills, ranked by demandranked
- Course warningsflags
Upload a curriculum CSV or a syllabus PDF, start a run, watch the seven steps, get the report. That is the whole product surface — the rest is the pipeline behind it.