Resume example (text format)

Prompt Engineer, Evaluation Jonah Lindqvist

Prompt Engineer, Support Automation

[email protected] | (503) 555-1608 | Portland, United States

Profile

Prompt engineer with eighteen months on a customer support automation platform, owning the intent classification and reply drafting prompts behind about 40,000 conversations a month. Built the team's first held-out eval set, 480 conversations sampled across nine intents and frozen before any prompt work started, plus the baseline it measured against. Shipped the reply drafting prompt that beat that baseline by 11 points of rubric score with the regression suite green.

Work Experience

03/2025 - Present, Prompt Engineer, Support Automation, Alderwood Systems, Portland, United States

  • Own the intent classification and reply drafting prompts behind about 40,000 support conversations a month.
  • Built the team's first held-out eval set: 480 conversations sampled across nine intents, deduplicated by customer, frozen before any prompt work began.
  • Shipped the reply drafting prompt that beat the frozen baseline by 11 points of rubric score, with the 210-case regression suite green.
  • Wrote the two-page rubric the eval set is scored against and trained three reviewers on it.

06/2023 - 03/2025, Support Operations Analyst, Alderwood Systems, Portland, United States

  • Built the intent taxonomy the eval set is stratified on, from a hand-coded sample of 1,200 conversations.
  • Ran the quality review queue for a team of 14 agents, including the weekly calibration session.

Education

09/2019 - 06/2023, Bachelor of Arts, Linguistics, University of Oregon, Eugene, United States

Coursework in corpus methods, statistics and experimental design.

Skills

Eval set design and stratified sampling, 75

Rubric authoring, 80

Inter-rater agreement, 65

Python, 70

SQL, 75

Error analysis, 75

Prompt version control, 65

Annotation tooling, 80

Languages

English, native

Swedish, advanced

Certificates

08/2024, Deep Learning Specialization, DeepLearning.AI

Summary

A prompt engineer resume is a one to two page document showing that a change you made produced a measurably better output: a held-out eval set, a baseline it had to beat, a stated judge, and a regression suite that caught what the fix broke. This guide gives you three adaptable versions, the evaluation block that separates a practitioner from a chat box, and the published occupations this work is paid as.

Prompt Engineer resume examples by experience level

A prompt engineer resume is a one to two page document showing that something you changed made the output measurably better, and that you can prove which change did it. That is the whole screen. Everyone applying has typed into a chat box and got a better answer on the second try. Very few can name the held-out set that told them the second answer was better in general, the baseline it beat, what did the judging, and the regression run that stopped the fix breaking something two intents over. The prompts themselves are weak evidence: short, copyable, and liable to break when the model changes. The evaluation is the artifact that does not transfer.

Resume guide for a prompt engineer resume

This guide and the corresponding prompt engineer resume example will cover:

  • Why the evaluation, not the prompt, is the evidence on this page
  • How to write an eval set: the size, the sampling and the freeze
  • Judges, rubrics and what to say about inter-rater agreement
  • Regression suites and prompt version control as resume lines
  • What the title is paid as, given that BLS does not publish it

How to write a prompt engineer resume

Six sections: contact header, summary, experience, an evaluation block, a skills snapshot, and education. One page for the first three or four years, two once you have owned an evaluation harness or a team. The evaluation block is the structural difference between this and a software engineer resume, and it earns its own space rather than being scattered through bullets.

SectionWhat it answersWhere it goes
SummaryCan this person prove a change made things better?Top, three or four lines
ExperienceWhich systems, at what volume?Reverse chronological
EvaluationWhat was measured, against what, judged by whom?Its own block
SkillsThe models, harnesses and languages, as an indexBelow evaluation
EducationThe degree and any measurement backgroundBottom, short
Expert Tip

Write the eval set before you write anything else

Draft the evaluation block first, even though it sits fourth on the page. It is the hardest part to write honestly, and the rest is easier once it exists.

Four facts make an eval set real: how many cases, where they came from, how they were stratified, and when they were frozen. "480 support conversations sampled from three months of production traffic, stratified across nine intents, frozen before any prompt work began" cannot be written by somebody who did not build one. The freeze is the part readers check: a set you kept editing while tuning against it is a training set, and a number from it means nothing.

Run the finished file through the free Rezoom ATS resume checker before you submit.

Choosing the best resume format for a prompt engineer resume

Reverse chronological, single column. The title is new enough that most candidates arrive from somewhere else, and the timeline makes the route legible: a linguist, a support operations lead, a data analyst and a QA engineer can all end up doing this work.

A combination format earns its place in one case: one or two years of language model work on top of several years in an adjacent field. Put a short relevant-work block first, then the history underneath with every date intact. Do not use a functional layout, which in a field this new reads as hiding how new the experience is.

Include your contact information

✅ Right❌ Wrong
Jonah LindqvistJonah Lindqvist
Prompt Engineer, Evaluation and Applied Language ModelsAI enthusiast and prompt wizard
(503) 555-1608, [email protected], Portland, OR(503) 555-1608, [email protected]
github.com/example, linkedin.com/in/exampleFull mailing address, date of birth, photograph

Give the title you actually hold, since postings for this work carry at least six.

Make use of a summary

Three or four lines: the system and its volume, the eval set you run, one shipped change and the number it beat.

Early career adaptable resume summary example

Prompt engineer with eighteen months on a customer support automation platform, owning the intent classification and reply drafting prompts behind about 40,000 conversations a month. Built the team's first held-out eval set, 480 conversations sampled across nine intents and frozen before any prompt work started, plus the baseline it measured against. Shipped the reply drafting prompt that beat that baseline by 11 points of rubric score with the regression suite green.

Mid career adaptable resume summary example

Prompt engineer with four years on production language model systems, owning the evaluation harness for a legal document review product used by 63 firms. Run a 2,100-case held-out eval set stratified by document type, scored against a written rubric with human adjudication on a 15% subsample at a Cohen's kappa of 0.74. Nine prompt versions shipped in the last year, each tied to a baseline, a regression run and a version identifier stamped on every production response.

Lead adaptable resume summary example

Applied AI lead with nine years across search relevance, data quality and language model evaluation, running the evaluation function behind three products with a team of five. Built the platform every prompt change is graded through: frozen eval sets, a 1,400-case regression suite in continuous integration, and a rubric review group that adjudicates disagreement. Cut median time from a proposed prompt change to a defensible ship or reject decision from eleven days to two.

Right vs wrong: the same prompt engineer summary, twice

✅ Right❌ Wrong
Run a 2,100-case held-out eval set stratified by document type, frozen before tuning.Extensive experience crafting and refining prompts for large language models.
Scored automatically against a written rubric, human adjudication on a 15% subsample at kappa 0.74.Skilled at evaluating model outputs for quality and accuracy.
Nine prompt versions shipped in the last year, each tied to a baseline and a regression run.Improved model performance through iterative prompt optimization.

The right column describes a process a reader could audit. The left describes an activity.

Outline your experience

Title, employer, city, dates, then three to five bullets. Name the system and its volume, because that sets the scale of everything else you claim.

Instead ofUse
Designed and optimized prompts for a customer service chatbotOwn the prompts behind about 40,000 support conversations a month
Improved the accuracy of model outputsTook rubric score on the held-out set from 3.1 to 3.8 of 5 against a frozen baseline
Tested prompts extensively before releaseBuilt a 2,100-case held-out eval set stratified by document type, frozen before any tuning began
Worked with engineers to deploy prompt changesPut prompts under version control with an identifier stamped on every production response
Used various prompting techniquesCut hallucinated citations in summaries from 6.2% to 0.9% of outputs by requiring quoted source spans
Adaptable resume employment history example

Prompt Engineer, Document Review, Rosehill Legal Technologies, Portland, OR, February 2024 to Present

Own the extraction and summarization prompts behind a document review product used by 63 firms, about 210,000 documents a month.

Built and maintain the 2,100-case held-out eval set, stratified by document type and jurisdiction, frozen before tuning and refreshed yearly under a documented sampling procedure.

Cut hallucinated citations in summaries from 6.2% to 0.9% of outputs by requiring quoted source spans and rejecting unsupported claims at validation.

Shipped nine prompt versions in the last year, each with a named baseline and a 640-case regression run; one was rolled back within four hours.

Wrote the rubric and trained four reviewers on it; adjudicated disagreement sits at a Cohen's kappa of 0.74.

Show the evaluation, not the prompt

This is the block that decides the screen, and most candidates do not have it because they have been iterating rather than measuring. An evaluation entry has six parts. Write one per system, and no more than three.

PartWhat it settlesExample
The eval setWhether the number generalizes2,100 cases from 14 months of production documents, stratified by type
The freezeWhether you tuned against your own testFrozen January 2025, before any prompt work
The baselineWhat the change had to beatPrompt version 7, rubric score 3.1 of 5
The metric and the judgeWho decided it was betterRubric 1 to 5, automatic scorer, human adjudication on 15%
The resultWhat actually movedRubric 3.1 to 3.8; hallucinated citations 6.2% to 0.9%
The ship recordWhether it survived contactVersion 9 live since March 2025; one rollback in June
Adaptable resume evaluation record example

Document summarization. Eval set: 2,100 cases from 14 months of production documents, stratified by type and jurisdiction, frozen January 2025. Baseline: prompt version 7 at rubric 3.1 of 5. Judge: automatic rubric scorer, human adjudication on a 15% subsample, kappa 0.74. Result: rubric 3.1 to 3.8, hallucinated citations 6.2% to 0.9%. Shipped: version 9, live since March 2025, one rollback in June.

Clause extraction. Eval set: 860 contracts across six clause types, frozen March 2024. Baseline: the regular expression extractor it replaced, recall 0.58. Judge: exact span match against two reviewers' annotations. Result: recall 0.58 to 0.86 at equal precision. Shipped: live since July 2024.

Intent classification. Supported, not owned.

Two details make that block credible rather than decorative. Give the baseline in the same sentence as the result, because a rubric score with nothing to compare it to is a number without a scale. And say plainly what you did not own: "supported, not owned" costs nothing and buys trust for the entries above it. In a field where teams are small and credit is blurry, overclaiming is the commonest way a strong file falls apart in conversation.

Expert Tip

Name the risks your eval set covers, and the ones it does not

The National Institute of Standards and Technology's Generative AI Profile, NIST AI 600-1, published in July 2024 as a companion to the AI Risk Management Framework, identifies twelve risk categories that are unique to or made worse by generative AI.

You do not need to recite them. What is worth a line is which your evaluation tests for and which it does not. "The eval set covers confabulation and information integrity; harmful content is handled by a separate classifier the safety team owns" is a sentence very few applicants can write, and it reads as someone thinking about coverage rather than scores.

The same honesty applies to the set itself. An eval set of 480 cases cannot separate a three-point difference from noise, and saying so first is worth more than the three points.

Judges, rubrics and inter-rater agreement

Naming the judge is the difference between a measurement and an impression.

The judgeWhen it is the right oneWhat to put on the resume
Exact or span match against annotationsExtraction, classification, structured outputThe annotation source and number of annotators
Automatic scorer against a written rubricOpen-ended generation with agreed criteriaThe rubric's dimensions and scale, and who wrote it
Human raters against the same rubricAnything subjective, and as adjudicationRaters, subsample size, agreement statistic

If more than one person judged, give the agreement statistic. Cohen's kappa is the usual one for two raters and it has a published interpretation: Landis and Koch, writing in Biometrics in 1977, set out the benchmark scale still in common use, where 0.41 to 0.60 is moderate agreement, 0.61 to 0.80 is substantial and 0.81 to 1.00 is almost perfect (Landis and Koch, Biometrics, volume 33, 1977).

Quoting a kappa of 0.74 and calling it substantial against that named scale tells a reader three things: you had more than one rater, you checked whether they agreed, and you know what the number means.

The regression suite and prompt version control

A prompt change that fixes one case and breaks four is the normal failure of this work, and it is invisible without a suite. A regression suite is a frozen set of cases with known-good outputs that runs on every change, separate from the eval set. The eval set answers whether the new version is better on average; the regression suite answers whether anything that used to work has stopped.

Write three facts about yours: how many cases, what triggers a run, what a failure blocks. "640 cases, running in continuous integration on every prompt change, blocking merge on any regression in the extraction cases" is a complete claim.

Version control is the other half, and it is one line that changes how a reader reads the rest: prompts in the repository, reviewed like code, with a version identifier stamped on every production response. That is what makes a complaint answerable six weeks later, which is the difference between an incident you can investigate and one you can only apologize for.

Do

Give the eval set a size, a sampling procedure and a freeze date. Name the baseline in the same sentence as the result. Say who or what judged, with the agreement statistic when more than one person did. Report the regression suite and what a failure blocks. Say what you supported rather than owned.

Iconly/Bold/Close Square Don’t

Do not paste prompts into the resume; they are short, copyable and dated. Do not quote a score with no baseline. Do not describe iterating until it looked better as evaluation. Do not claim an eval set you tuned against. Do not name every model you have called through an API; the reader will ask about the one you know least well.

Build a snapshot of your key skills

Three groups, every entry traceable to something above it.

Adaptable resume skills example

Evaluation: Eval set design and stratified sampling, Rubric authoring, Inter-rater agreement, Regression suite design, Error analysis

Systems: Python, SQL, Git, Continuous integration, Retrieval augmented generation, Structured output validation, Prompt version control

Models and tooling: Commercial and open-weight model families, Evaluation harnesses, Annotation tooling, Tracing

Name model families rather than every model you have called, and put the ones you can discuss in depth first.

List your education and certifications

Degree, field, institution, city, year. This role has no licensing body and no required credential, so the block stays short and the measurement content in it is the part worth naming.

Adaptable resume education example

Bachelor of Arts, Linguistics, University of Oregon, Eugene, OR, 2017. Coursework in corpus methods, statistics and experimental design.

Certifications: Deep Learning Specialization, DeepLearning.AI (2023).

Writing: internal evaluation handbook adopted across three product teams, 2025.

No certification qualifies anyone for this work, and a reader who sees four stacked above a thin evaluation block draws the obvious conclusion.

Choose the right layout and design

Single column, 10 to 12 point body type, clear headings, no photograph, no skill rating bars. The evaluation block wants a consistent shape rather than decoration: the same six parts in the same order for each system, so two entries can be compared without re-reading. Keep the numbers identical to anything you say in an interview, because these figures are specific enough that a discrepancy gets noticed.

Prompt engineer job market and outlook

The U.S. Bureau of Labor Statistics publishes no Occupational Outlook Handbook profile titled prompt engineer. There is no official employment count, no projected growth rate and no federal median wage for the title, and the retired wage pages still circulating online belong to an older cycle. So this page quotes no prompt engineer wage, and names the published occupations that bracket the work.

Bracketing occupation, BLS publishedMedian annual wage, May 2025Why it brackets this work
Computer and information research scientists$140,300When the work is method: designing how things get evaluated at all
Software developers$135,980When prompts are part of a shipped product with a test suite
Data scientists$120,230When the job is measurement: sampling, metrics, agreement, holdouts

Source: U.S. Bureau of Labor Statistics, Occupational Outlook Handbook, May 2025 wage data.

The closest in daily shape is data scientists, whose detail is the most useful reference available:

Measure, data scientistsValue
Jobs, 2025275,600
Projected change, 2025-3535% (much faster than average), +95,400
Projected annual openingsAbout 24,800
Median annual wage, May 2025$120,230
Typical entry-level educationBachelor's degree

Source: U.S. Bureau of Labor Statistics, Occupational Outlook Handbook, 2025-35 projections and May 2025 wage data.

Statistical insight

The title is unpublished; the occupations around it are growing fast

All three bracketing occupations are projected to grow between 2025 and 2035. Data scientists are projected to grow 35%, adding 95,400 jobs from a 2025 base of 275,600, with about 24,800 openings a year. Computer and information research scientists are projected to grow 22%, adding 8,400 from a base of 38,600, with about 2,900 openings a year. Software developers, quality assurance analysts and testers are projected to grow 10%, adding 185,400 from a base of 1,905,400, with about 106,100 openings a year (BLS, 2025-35 projections).

Read the sizes as well as the rates. The research occupation pays the most and hires the least, and its typical entry-level education is a master's degree; the developer occupation hires 36 times as many a year at a similar median. The target matters more than the title.

What salary you can expect as a prompt engineer

With no federal wage figure for the title, treat the three bracketing occupations as the band an offer sits inside, and take an occupation figure into the conversation.

At the measurement end, data scientists had a median of $120,230 as of May 2025. Where prompts are part of a shipped product with a test suite, software developers had a median of $135,980. Where the job is designing evaluation method rather than running it, computer and information research scientists had a median of $140,300, and a master's degree is the typical entry-level education for that occupation (BLS, May 2025).

What moves you inside that band is the evidence this page is built on. A candidate who can name a frozen eval set, a baseline, a judge, an agreement statistic and a rollback is priced as someone who can be trusted with a release decision.

Key takeaways for a prompt engineer resume

  1. Lead with the evaluation, not the prompt; the prompt is the least transferable thing you have.
  2. Give the eval set a size, a sampling procedure and a freeze date.
  3. Name the baseline in the same sentence as every result.
  4. Say who or what judged, with the agreement statistic when more than one did.
  5. Report the regression suite, what triggers it and what a failure blocks.
  6. Put prompts under version control and stamp the version on production responses.
  7. Say what your evaluation does not cover, and what you supported rather than owned.

Build your prompt engineer resume in 15 minutes with our AI resume builder.

Related information technology resume examples

Pair it with a matching prompt engineer cover letter.

Prompt engineer resume questions, answered

Is prompt engineer a real job title?

It is a real posting title with no federal occupational profile behind it. The Bureau of Labor Statistics publishes no Occupational Outlook Handbook profile for prompt engineers, so there is no official employment count, growth rate or median wage. The work is hired and paid under the occupations that bracket it.

What does a prompt engineer earn?

There is no published federal figure for the title, and this page will not quote a jobs-board average as if there were. The published brackets as of May 2025 were $120,230 for data scientists, $135,980 for software developers and $140,300 for computer and information research scientists (BLS, May 2025). Which one an offer resembles depends on whether the role is measurement, product engineering or method design.

What do I put on the resume if I have no formal prompt engineering job?

The same evidence, from wherever you produced it. An internal assistant built for your own team still had users, a volume, a baseline of whatever they did before, and some way of telling whether it helped. A small honest evaluation beats a large claimed one.

How big should an eval set be?

Big enough that the difference you care about is not noise, which usually means larger than people expect. A 40-case set cannot separate a three-point rubric difference from sampling variation. State the size and the sampling procedure so a reader can judge it, and say when a result is suggestive rather than conclusive.

Do I need a machine learning degree for this work?

No, and the routes in are varied: linguistics, support operations, data analysis, quality assurance and software engineering all lead here. What substitutes for the degree is measurement literacy: sampling, holdouts, agreement statistics, and knowing when a difference is not real.

Should I include the prompts themselves in a portfolio?

One or two at most, and only where they illustrate a design decision you can explain. Prompts go stale when the model underneath changes. What travels is the evaluation: the sampling procedure, the rubric, the regression suite, and the write-up of a change that did not work.