How ATS Resume Parsing Actually Works
What happens between uploading a CV and a recruiter seeing it: text extraction, section segmentation, entity extraction and search indexing — and where each stage breaks.
Published · Updated · by Maksym Shykov
Short answer
An ATS parses a CV in four mechanical stages: it extracts the text, splits it into sections by their headings, pulls out entities such as email, dates and job titles, and indexes the result for recruiters to search. Nearly every piece of CV formatting advice exists because one of those stages breaks.
Most CV advice treats the applicant tracking system as a black box with opinions. It is simpler than that. Parsing runs in stages, each one mechanical, and almost every piece of real formatting advice is downstream of a specific stage failing.
Stage 1 — Text extraction
The file is opened and characters are pulled out of it. For a DOCX this is mostly reading XML. For a PDF it is harder: a PDF does not store lines or paragraphs, it stores fragments of text with coordinates. Reconstructing “this fragment and that fragment are on the same line” is geometry, and it is where multi-column layouts fall apart — two columns can be stitched into one nonsensical line.
Here is what that looks like. A two-column CV, as a reader sees it, and the text a line-by-line extractor — this site’s included — actually gets back:
EXPERIENCE SKILLS Engineering Manager, Acme Kubernetes, Go Jan 2020 – now Terraform, AWS
EXPERIENCE SKILLS Engineering Manager, Acme Kubernetes, Go Jan 2020 – now Terraform, AWS
Every line is now a hybrid. The Experience heading shares a line with Skills, your job title has two skills glued to it, and the date range is followed by cloud providers. Nothing was lost — it was just put back together in the wrong order, which for a parser is the same thing.
Fails when: the page is an image, the text is outlined, or the layout is multi-column.
Stage 2 — Section segmentation
The text is cut into regions by looking for headings. This is why the boring heading beats the clever one: a parser matching against a list of known words finds Experience, Work Experience and Employment History. It does not find Where I've Made a Dent.
To make this concrete: these are the exact headings this site’s checker looks for. It accepts a short line — 45 characters or fewer, in any capitalisation — that contains one of them, so Relevant Work Experience counts as Experience.
| Section | Headings it accepts | Points |
|---|---|---|
| Experience | “experience”, “employment history”, “work experience”, “work history”, “professional experience” | 8 |
| Education | “education”, “academic background” | 6 |
| Skills | “skills”, “core competencies”, “technical skills”, “expertise” | 6 |
| Summary | “summary”, “profile”, “objective”, “about me”, “about” | 5 |
| Achievements | “achievements”, “key achievements”, “accomplishments”, “highlights” | 4 (bonus) |
| Projects | “projects”, “selected projects”, “side projects” | 3 (bonus) |
| Certifications | “certifications”, “certificates”, “courses”, “licenses”, “certifications & courses” | 3 (bonus) |
Commercial parsers use longer lists, but the principle is identical — and none of them has your invented heading on it.
Fails when: headings are creative, styled only by colour, or absent.
Stage 3 — Entity extraction
Within each section the parser looks for specific things: an email is a pattern, a phone number is a pattern, a date range marks the start of a role. Job title and employer are inferred from position relative to the date.
This is why dates matter more than people expect. A dated line is the anchor that tells the parser “a new job starts here”. Undated roles collapse into whatever came before them, and your three years somewhere land inside the previous employer's entry.
Fails when: dates are missing, written only as “2 yrs”, or drawn as a graphic timeline.
Stage 4 — Indexing and search
The extracted fields go into a database. A recruiter then searches it — by title, by skill, by location. Your CV is not being judged at this point; it is being queried.
Which reframes the keyword question. The goal is not to stuff terms in to please an algorithm. It is to make sure the words a recruiter would plausibly type appear somewhere you have honestly earned them. If you led Kubernetes migrations and never wrote the word “Kubernetes”, you will not be in that result set.
Fails when: the vocabulary of your CV and the vocabulary of the job ad have no overlap.
What this implies
- Formatting advice is not superstition — each rule maps to a stage above.
- Keyword advice is about vocabulary overlap, not density. Never claim a skill you lack.
- Nothing in this pipeline evaluates you. It moves you into a searchable shape, or fails to.
Related
Frequently asked questions
- Can an ATS read a two-column resume?
- Often not in the right order. Text extraction rebuilds lines from positions on the page, so two columns can be stitched together line by line, mixing a job title with a skills list. A single-column layout avoids the problem entirely.
- Do applicant tracking systems read headers and footers?
- Some parsers skip those regions or treat them unreliably. Keep your name, email and phone number in the main body of the page so they are extracted whichever parser reads the file.
- Does hiding keywords in white text help?
- No. White text is still text: it is extracted into the parsed profile a recruiter reads, where it looks like an attempt to game the search. Use the job ad’s vocabulary only where you have genuinely earned it.