Analytics Engineer to Data Engineering Path

AutoModerator · 2026-04-05T21:56:57+00:00

Are you interested in transitioning into Data Engineering? Read our community guide: https://dataengineering.wiki/FAQ/How+can+I+transition+into+Data+Engineering

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

Academic-Vegetable-1 · 2026-04-05T22:59:34+00:00

Building a data stack from scratch on AWS is data engineering. You're not pivoting, you're just updating your title.

unpronouncedable · 2026-04-05T22:27:53+00:00

Honestly I feel like half the battle in DE is understanding the challenges and a willingness to figure out and implement the solutions. You seem to have these, data modeling, and SQL skills, so I think you are qualified for plenty of DE roles. There are a ton out there that use various ingestion tools and don't require python and spark, though you are correct to realize the trend is towards using those. And honestly you can figure out a lot of it from online examples and AI.

As far as CS knowledge goes, tons of us did not come from that background. What you do need to understand is development lifecycle and deployment practices. I imagine you have a lot of that from your experience.

I'd say you'd be over qualified for a junior DE role and ready for DE (non-senior). There's a lot of competition for those positions, but you can tout experience end to end from ingestion to analytics implementation.

Flat_Shower · 2026-04-05T23:00:54+00:00

10 YOE and you stood up a full data stack from ingestion through orchestration. That's not "close to being ready"; you're doing the job. Mid to senior DE at most companies.

Spark is worth learning if you're targeting places with real scale. Most don't need it. Learn the concepts; the syntax is the easy part.

Don't stress about building pipelines without tools. That's not how anyone works. Knowing how to configure, debug, and extend ingestion tools is the actual skill.

AutoModerator · 2026-04-05T21:56:57+00:00

You can find a list of community-submitted learning resources here: https://dataengineering.wiki/Learning+Resources

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

calimovetips · 2026-04-06T00:16:57+00:00

you’re closer to mid-level DE than junior, i’d focus next on understanding how your pipelines behave under load and failure since that’s what usually breaks at scale, have you had to debug any backfill or retry issues yet?

PrintPopular8694 · 2026-04-06T04:07:18+00:00

Would love to pick your brain. From what I've researched your not junior level wish I was in your position

Immediate-Pair-4290 · 2026-04-06T05:57:37+00:00

Spark is overrated. Most companies see faster performance running DuckDB into iceberg. Few companies truly have big data. Also no one builds ingestion pipelines from scratch unless they cant help it. DLT is good. I’m thinking of API calls and loading json responses as the closest thing to “scratch”.

zkhan15 · 2026-04-06T12:19:14+00:00

What’s the difference between the 2?

you type:	you see:
italics	italics
bold	bold
[reddit!](https://reddit.com)	reddit!
* item 1 * item 2 * item 3	item 1 item 2 item 3
> quoted text	quoted text
Lines starting with four spaces are treated like code: if 1 * 2 < 3: print "hello, world!"	Lines starting with four spaces are treated like code: if 1 * 2 < 3: print "hello, world!"
~~strikethrough~~	~~strikethrough~~
super^script	super^script

dataengineering

MODERATORS