Data Engineer Hana Sugimoto
Data Engineer
[email protected] | (206) 555-1834 | Seattle, United States
Profile
Data engineer with 18 months on a three-person platform team at a freight company, running 24 AWS Glue jobs that land about 40 GB a day into Amazon S3 and Amazon Redshift. Converted the largest raw feed from gzipped JSON to partitioned Apache Parquet, which took the Amazon Athena bytes scanned on the daily operations query from 1.4 TB to 90 GB and removed the weekly query timeout the operations team had learned to work around. Hold the AWS Certified Data Engineer - Associate credential.
Work Experience
03/2025 - Present, Data Engineer, Elliott Bay Logistics, Seattle, United States
- Run 24 AWS Glue jobs landing about 40 GB a day from 9 partner file feeds into Amazon S3 and Amazon Redshift.
- Converted the largest raw feed from gzipped JSON to Apache Parquet partitioned by date and carrier; Athena bytes scanned on the daily operations query fell from 1.4 TB to 90 GB.
- Added row count and null rate checks between Glue stages so a bad upstream day fails the run instead of publishing a partial load.
- Maintain 84 tables in the AWS Glue Data Catalog, including the crawler schedule and the schema change alerts.
01/2024 - 03/2025, Data Analyst, Elliott Bay Logistics, Seattle, United States
- Wrote the SQL behind the operations reporting that later became the first Glue pipeline.
- Documented the field definitions for the carrier feeds before they were cataloged.
Education
09/2020 - 06/2024, Bachelor of Science, Computer Science, University of Washington, Seattle, United States
Coursework in database systems, distributed systems and statistics.
Skills
AWS Glue and PySpark, 70
Amazon S3 and partitioning, 70
Apache Parquet, 70
Amazon Athena, 70
Amazon Redshift, 60
SQL and window functions, 80
Python, 75
AWS Glue Data Catalog, 65
Languages
English, native
Japanese, native
Certificates
06/2025, AWS Certified Data Engineer - Associate, Amazon Web Services
Exam DEA-C01. Valid for three years from the date earned.
Summary
An AWS data engineer resume is a one to two page document naming the AWS services in your pipelines, the volume moving through them each day, and what the platform costs after you worked on it. This guide gives you three adaptable versions, the pipeline block almost nobody writes, what AWS publishes about its data engineering credential, and current Bureau of Labor Statistics pay for the published occupation this work sits closest to.
AWS Data Engineer resume examples by experience level
In short, an AWS data engineer resume is a one to two page document naming the services in your pipelines, the volume moving through them, and what the platform costs to run after you touched it.
That is three facts, yet most resumes for this title give none of them. They give a service list instead: the twenty-item row of logos on every file in the stack, which proves only that the writer has read a job posting. After all, a service named without a volume or a cost beside it is a word, not a claim.
Resume guide for an AWS data engineer resume
Specifically, this guide and the corresponding AWS data engineer resume example will cover:
- First, how to write an AWS data engineer resume, section by section
- Then why every service you name needs a volume and a cost beside it
- Also, three adaptable summaries: first pipeline role, working engineer and lead
- Finally, what AWS publishes about its data engineering credential
- What the market pays when the Handbook publishes no profile under this title
How to write an AWS data engineer resume
To start, plan six sections: contact header with a title line, summary, a pipeline block, experience, skills grouped by layer, and education with certifications. Keep one page for your first three or four years, then two once you own a platform other teams build on.
The order follows what the reader is checking, which is operational rather than theoretical: what breaks at three in the morning, who fixes it, and what the monthly bill looks like.
| Section | What it is answering | Where it goes |
|---|---|---|
| Title line | Pipeline author, or owner of a platform? | Directly under your name |
| Summary | Which services, moving how much, costing what? | Three or four lines |
| Pipeline block | Could this person be handed our data platform? | Its own block |
| Experience | What did you build, and what did it replace or remove? | Reverse chronological, with numbers |
| Skills | Ingestion, storage, transformation, orchestration, governance | Grouped by layer |
| Certifications | Which AWS credential, and is it still inside its validity period? | With education |
Open the billing console before you open the resume
Collect five numbers: how much data lands per day, how many jobs or state machines run it, how many tables are in the catalog, what your part of the account costs a month, and what it cost twelve months ago.
The last pair nobody writes, and it is the pair that gets interviews. Data engineering is the part of the stack where inefficiency arrives as an invoice, so "took $41,000 a month out of the platform bill and kept every day of history" is a result finance can verify without reading a line of the pipeline.
Run the finished file through the ATS resume checker before you submit. Cloud data postings draw very heavy volume and are parsed into a structured record first.
Choosing the best resume format for an AWS data engineer resume
For format, generally use reverse chronological, single column, with the pipeline block high on the page. However, move to two pages only once you own a platform other teams depend on.
Skip the functional format even if you came from analytics, database administration or backend development. The reader is checking how long you carried something in production, because the second year on a pipeline is when you learn what it does at month end and when an upstream system sends a day of duplicates.
Include your contact information
| ✅ Right | ❌ Wrong |
|---|---|
| Hana Sugimoto | Hana Sugimoto |
| AWS Data Engineer, 320 TB lakehouse, 140 Glue jobs, $41k a month removed | Big Data Engineer, Cloud Enthusiast |
| AWS Certified Data Engineer – Associate, 2025 | AWS certified |
| (206) 555-1834, [email protected], Seattle, WA | (206) 555-1834, [email protected] |
For example, the right column states a scale, a workload and an outcome in one line, and dates the credential. AWS publishes a validity period, so an undated one invites the reader to assume the worst.
Make use of a summary
Next, write three or four lines: the services you ran, the volume they moved, one cost number, and the credential.
Data engineer with 18 months on a three-person platform team at a freight company, running 24 AWS Glue jobs that land about 40 GB a day into Amazon S3 and Amazon Redshift. Converted the largest raw feed from gzipped JSON to partitioned Apache Parquet, which took the Amazon Athena bytes scanned on the daily operations query from 1.4 TB to 90 GB and removed the weekly query timeout the operations team had learned to work around. Hold the AWS Certified Data Engineer - Associate credential.
AWS data engineer with six years building ingestion and transformation pipelines in healthcare, running 61 AWS Glue jobs and 9 AWS Step Functions state machines that move about 1.2 TB a day out of Amazon Kinesis Data Streams and Amazon RDS into Amazon S3 and Amazon Redshift. Rewrote the claims pipeline onto partitioned Parquet with an S3 lifecycle policy and took $14,200 a month off the storage and query bill without losing a day of history. Hold the AWS Certified Data Engineer - Associate credential.
Lead data engineer with 11 years in data platforms and 4 leading a four-person team on an AWS lakehouse holding 320 TB in Amazon S3, fed by 140 AWS Glue jobs, Amazon MSK and AWS Database Migration Service, and served through Amazon Athena and Amazon Redshift. Took $41,000 a month out of the platform bill by partitioning the three heaviest tables, moving cold partitions to S3 Intelligent-Tiering and retiring two Amazon EMR clusters that each ran nightly for one job. Hold the AWS Certified Data Engineer - Associate and AWS Certified Solutions Architect - Associate credentials.
Right vs wrong: the same AWS data engineer summary, twice
| ✅ Right | ❌ Wrong |
|---|---|
| 140 AWS Glue jobs feeding a 320 TB lakehouse in Amazon S3, served through Athena and Redshift. | Experienced with AWS Glue, S3, Redshift, Athena, EMR, Lambda, Kinesis and Step Functions. |
| Took $41,000 a month off the platform bill and kept every day of history. | Optimized data pipelines for performance and cost efficiency. |
| Athena bytes scanned on the daily operations query fell from 1.4 TB to 90 GB after partitioning. | Improved query performance through best practices. |
In other words, the right column names a service, attaches a number and says what changed. By contrast, the wrong column is a keyword row the reader has seen forty times this week.
Outline your AWS data engineer experience
Employer, sector, dates, then the platform before any bullet: what lands, how much, into what, served to whom. Additionally, in experience, write three to five bullets each, every one naming a service and a number.
| Instead of | Use |
|---|---|
| Built and maintained ETL pipelines using AWS services | Run 61 AWS Glue jobs and 9 Step Functions state machines moving 1.2 TB a day into S3 and Redshift |
| Optimized data storage and reduced costs | Partitioned the three heaviest tables and moved cold partitions to S3 Intelligent-Tiering, removing $41,000 a month |
Lead Data Engineer, Kestrel Health Systems, Seattle, WA, April 2022 to Present
Lead a four-person team on an AWS lakehouse of 320 TB in Amazon S3, fed by 140 AWS Glue jobs, Amazon MSK and AWS Database Migration Service, and served through Amazon Athena and Amazon Redshift to 11 analytics teams.
Took $41,000 a month off the platform bill: partitioned the three heaviest tables by date and payer, moved cold partitions to S3 Intelligent-Tiering, and retired two Amazon EMR clusters that each ran nightly for one job now handled by Glue.
Rewrote the claims pipeline from gzipped JSON into partitioned Apache Parquet, cutting the bytes an ordinary Athena query scans by more than 90 percent and removing a timeout the reporting team had worked around for two years.
Put row count and null rate checks between every Glue job in the 9 Step Functions state machines, so a bad upstream day fails the run rather than publishing half a day of claims.
Own Lake Formation permissions across 640 catalog tables and 9 teams, including the column rules keeping member identifiers out of the analytics workspace.
Name the services, the volumes and the cost you took out
In practice, this is the section almost nobody writes, and it separates an engineer who has run a platform from one who has finished a tutorial on the same services.
Name the services precisely. AWS publishes the service list its own data engineering exam covers, a fair map of what this job is made of: AWS Glue, Amazon Athena, Amazon EMR and AWS Lake Formation for analytics; Amazon Kinesis Data Streams, Kinesis Data Firehose and Amazon MSK for streaming; Amazon S3, S3 Tables and S3 Glacier for storage; Amazon Redshift, RDS, Aurora and DynamoDB for databases; and AWS Step Functions, Amazon MWAA and Amazon EventBridge for orchestration (AWS Certified Data Engineer – Associate exam guide, retrieved 17 September 2026).
Instead, name the four or five you actually operated and say what each does. "Amazon MSK for the clinical event stream, Glue for transformation, Step Functions for orchestration, Athena and Redshift for serving" is a sentence a reader can draw. On the other hand, a row of twenty logos is not.
Volumes and query costs for AWS pipelines
Attach a volume to each one. Gigabytes a day landing, events a second, tables in the catalog, jobs on the schedule, teams querying it. Volume is what makes a service name mean something: Glue moving 40 GB a day and Glue moving 1.2 TB a day are different jobs with different failure modes.
Then give the cost. This is what separates the file, and it is specific to this occupation: in most engineering roles waste is invisible, while here it arrives monthly as a bill somebody in finance is already annoyed about.
Amazon Athena SQL queries are priced on data scanned, and AWS's own worked example shows the lever: querying a 3 TB uncompressed text file costs twelve times what the same data costs after it is compressed to 1 TB and converted to Apache Parquet, because Athena then reads only the column the query needs (AWS, Amazon Athena pricing, retrieved 17 September 2026). Also, partitioning and file format are not housekeeping. They are the bill.
Storage costs and the facts to include
Storage is the same. S3 Intelligent-Tiering moves an object untouched for 30 days to the Infrequent Access tier and one untouched for 90 days to Archive Instant Access, with optional Archive Access and Deep Archive Access tiers configurable out to 730 days, and objects smaller than 128 KB are not eligible for automatic tiering (AWS, Amazon S3 documentation, retrieved 17 September 2026). An engineer who knows that last clause has met the problem it causes.
| The fact | Why the reader wants it | How to write it |
|---|---|---|
| Services operated | Separates running a pipeline from listing a stack | Glue, Step Functions, MSK, Athena, Redshift, Lake Formation |
| Volume and scale | Makes the service name mean something | 1.2 TB a day landing, 320 TB in S3, 640 catalog tables |
| File format and partitioning | The single biggest lever on query cost | Partitioned Parquet by date and payer |
| Storage lifecycle | Shows you manage data after it stops being new | Cold partitions on S3 Intelligent-Tiering |
| Governance | Whether you can be trusted with regulated data | Lake Formation column rules across 9 teams |
| Cost removed | The result finance can verify | $41,000 a month, history kept |
Platform, current role
Ingestion: Amazon MSK for the clinical event stream at about 4,100 events a second, AWS Database Migration Service for six operational databases, Glue jobs for 22 partner file feeds.
Storage and catalog: 320 TB in Amazon S3, partitioned Apache Parquet by date and payer, 640 tables in the AWS Glue Data Catalog, cold partitions on S3 Intelligent-Tiering.
Transformation and orchestration: 140 AWS Glue jobs inside 9 AWS Step Functions state machines, with row count and null rate checks between stages and a dead letter path.
Serving and governance: Amazon Athena and Amazon Redshift for 11 analytics teams, AWS Lake Formation permissions including column rules on member identifiers.
Cost: $41,000 a month removed over four quarters, with no reduction in retained history.
If you inherited the platform, say inherited and say what it looked like. If a separate team owned the account, the networking or the security boundary, say so. After all, a reader who runs an AWS account tests that boundary in ten minutes, and being straight about it costs nothing.
Build a snapshot of your key AWS data engineer skills
Twelve to eighteen entries, grouped by layer rather than by vendor, because that is how the reader thinks about their own platform.
Ingestion: Amazon MSK, Kinesis Data Streams and Firehose, AWS Database Migration Service, change data capture
Storage and format: Amazon S3, partitioning strategy, Apache Parquet, S3 Intelligent-Tiering and lifecycle policies, AWS Glue Data Catalog
Transformation: AWS Glue and PySpark, SQL and window functions, Amazon EMR, incremental and backfill patterns
Orchestration: AWS Step Functions, Amazon MWAA, EventBridge schedules, retries and dead letter queues, data quality checks between stages
Governance and cost: AWS Lake Formation, IAM least privilege, KMS encryption, cost allocation tags, query cost monitoring
List your education and certifications
Degree first if you have one, then certifications with the exact name and the year.
Meanwhile, a bachelor's degree is the typical entry-level education for database administrators and architects, the published occupation this work sits closest to (BLS Occupational Outlook Handbook, retrieved 17 September 2026). Many teams take proven pipeline work instead, but the line clears a filter at large employers.
AWS states the validity period, so date the credential
The credential written for this job is the AWS Certified Data Engineer – Associate, exam DEA-C01. AWS describes the target candidate as having the equivalent of 2 to 3 years in data engineering and 1 to 2 years hands-on with AWS services, and sets the exam at 65 questions over 130 minutes (AWS, retrieved 17 September 2026).
Its four domains tell you where to weight the page: data ingestion and transformation at 34 percent, data store management at 26 percent, data operations and support at 22 percent, data security and governance at 18 percent (AWS, DEA-C01 exam guide, retrieved 17 September 2026). Roughly the proportion they deserve on a resume.
AWS states that the certification is valid for three years and that you recertify by passing the latest version of the exam (AWS, retrieved 17 September 2026). So write the year. An undated AWS credential is read as one that has run out.
Bachelor of Science, Computer Science, University of Washington, Seattle, WA, 2015
Certifications:
AWS Certified Data Engineer - Associate, Amazon Web Services, 2025
AWS Certified Solutions Architect - Associate, Amazon Web Services, 2024
Choose the right layout and design
Lastly, use a single column, 11 point body type, plain headings, no photograph, no architecture diagram. That is because a diagram is unreadable at the size it ends up on a resume, and the pipeline block does the same job in words a parser can index.
Name four or five services and say what each does in your pipeline. Attach a daily volume to the ingestion layer. Give the file format and the partitioning key. Say what the platform costs now and what it cost before. Say which team owned the account if it was not yours. Date every AWS credential.
Do not print a twenty-logo service row; it is the commonest filler on these files. Do not claim a service you have only read about, because the screen is a whiteboard and a failure mode. Do not describe an inherited pipeline as one you designed. Do not put an architecture diagram on a resume.
AWS data engineer job market and outlook
The Bureau of Labor Statistics publishes no Occupational Outlook Handbook profile under the title AWS data engineer, and none named for a cloud provider, so no employment count or ten-year projection carries this job title.
Instead of titles, the Handbook classifies by the work, and the closest published occupation is database administrators and architects. That profile is unusually useful here, because BLS splits it into two halves that behave differently, and the split maps onto the two versions of this job.
| Measure | Database administrators | Database architects |
|---|---|---|
| Median annual wage, May 2025 | $104,620 | $139,500 |
| Lowest 10 percent | Under $60,230 | Under $86,240 |
| Highest 10 percent | Over $163,320 | Over $204,000 |
| Employment, 2025 | 75,000 | 69,500 |
| Projected change, 2025-35 | 0 percent, -100 | 9 percent, +6,500 |
Source: U.S. Bureau of Labor Statistics, Occupational Outlook Handbook, May 2025 wages and 2025 to 2035 projections, published 27 August 2026. Overall, combined median $126,760, employment 144,500, about 7,300 openings a year.
One occupation, two halves, and all of the growth is in one of them
Database administrators and architects held 144,500 jobs in 2025 at a combined median of $126,760, and the occupation is projected to grow 4 percent by 2035 with about 7,300 openings a year. Underneath that average, administrators sit at 0 percent, losing 100 jobs, while architects sit at 9 percent, adding 6,500 (BLS Occupational Outlook Handbook, 2025 to 2035 projections). All the growth is in the design half.
The pay gap between the halves is $34,880 at the median, and among architects, computer systems design employed 23 percent at a median of $148,680 (BLS, May 2025).
That split is the argument for the pipeline block. An engineer who keeps existing pipelines running is described by the flat half. One who decides the partitioning, the file format, the lifecycle policy and the orchestration is described by the half with all the growth.
What salary you can expect as an AWS data engineer
No federal median is attached to this job title, since no published occupation is named after a cloud vendor, and negotiating from a figure invented for a title BLS has never defined is a bad idea.
Therefore, use the published occupation and its split. For instance, database administrators earned a median of $104,620 in May 2025, the lowest 10 percent under $60,230 and the highest over $163,320. Similarly, database architects earned $139,500, the lowest 10 percent under $86,240 and the highest over $204,000 (BLS, May 2025).
Three things in the pipeline block move an offer inside that range: the volume you operated at, whether you designed the storage layout or inherited it, and whether you can name a cost figure you removed. The last is the rarest, and it is the only number on the page a non-technical decision maker can evaluate alone.
Key takeaways for an AWS data engineer resume
- Above all, name four or five services and say what each does, not twenty in a row.
- Then attach a daily volume to the ingestion layer and a size to the storage layer.
- In addition, give the file format and the partitioning key; they are the query bill.
- Finally, state the cost you removed and confirm you kept the history.
- Say what happens when a run fails, and who it wakes.
- Date the AWS credential, since AWS publishes a three year validity period.
- Mark what you inherited, because the whiteboard round finds that line anyway.
Build your AWS data engineer resume in 15 minutes with our AI resume builder.
Related information technology resume examples
- Data engineer resume
- Data scientist resume
- Data analyst resume
- SQL developer resume
- DevOps engineer resume
- Machine learning engineer resume
- Python developer resume
- Solution architect resume
- Software engineer resume
- Tableau developer resume
- Power BI developer resume
- Business intelligence and IT resume
Pair it with a matching AWS data engineer cover letter.
AWS data engineer resume questions, answered
How long should an AWS data engineer resume be?
One page for your first three or four years. Two once you own a platform other teams build on, because volumes, orchestration and governance do not compress into one line. Three pages is a design document nobody asked for.
Should I list every AWS service I have touched?
No. The twenty-logo row is the most common filler on these resumes and reads as a copy of the posting. Name the four or five you operated, with what each does and a number attached. If a posting names one you have used lightly, add it once and be ready to say how lightly.
Is the AWS Certified Data Engineer - Associate worth taking?
It is the credential this job's postings name most often, and AWS publishes enough about it that a reader knows what passing it means. AWS describes the target candidate as having 2 to 3 years in data engineering and 1 to 2 years hands-on with AWS, and states the certification is valid for three years (AWS, retrieved 17 September 2026). Put the year on the page.
What do I write with no professional AWS pipeline yet?
Build one small pipeline end to end on a public dataset and describe it as this guide describes work: what lands, in what format, partitioned how, orchestrated by what, and what it costs a month. A pipeline with a cost figure on it, however small, shows the habit employers are hiring for.
How do I show cost savings if my employer will not share billing figures?
Give the proportional change instead. "Bytes scanned on the daily operations query fell from 1.4 TB to 90 GB" is a cost result in a unit you are allowed to say, and any reader who runs Athena converts it instantly. Never invent a dollar figure you cannot support.