
Senior Data Engineer w/ NLP & ML
- Remote
- Prague, Praha, Hlavní město, Czechia
- €35 - €41 per hour
- Jimmy Technologies
If you are passionate about Data, AI, Python, and building next-generation ML workflows, this role with our client offers an exciting opportunity to work on cutting-edge R&D projects!
Job description
Our partner, a Fortune 500 company, is building an AI incubator to scout, incubate, and validate client and internal ideas on a 3–5 year horizon. The team develops technology roadmaps and prototypes that turn into advanced client solutions — and continually pushes cutting-edge AI into new products, services, and capabilities.
We're looking for a Senior Data Engineer to design, build, and operate reliable, production-ready data pipelines for large-scale data and ML use cases. You'll bridge raw data ingestion, transformation, and quality assurance with downstream machine learning workflows — with particular emphasis on NLP/NLU.
This is a long-term, remote-first contract position with required US working hours overlap (2-6 PM CET).
Responsibilities
Build & harden pipelines: design and maintain reliable ingestion pipelines that move raw data into production-ready platforms.
ETL/ELT: own data cleaning, transformation, schema design, quality validation, and lineage tracking.
Large-scale processing: develop data workflows using Databricks and PySpark.
ML enablement: prepare high-quality datasets, features, and pipelines for NLP, NLU, and broader ML use cases.
Azure infrastructure: operate and integrate Blob Storage, databases, compute resources, MLflow, and Azure AI/ML services.
CI/CD & MLOps: implement automated deployment, testing, monitoring, and promotion gates.
Own production readiness: ensure pipeline reliability, observability, performance, and data quality end-to-end.
Work Conditions
Type: Full-time & Long-term contract work
Start Date: ASAP
Location: Remote (99%) in Europe; must be able to travel freely within Europe for workshops.
US Time Zone Overlap: Required (2 PM - 6 PM CET)
Contract with European LLC
Job requirements
Strong Data Engineering background — ETL/ELT, pipeline development, schema design, orchestration.
Hands-on experience building reliable, production-ready ingestion pipelines with proper hardening, validation, and monitoring.
Strong Python and PySpark skills for large-scale data processing and transformation.
Practical Databricks experience — scalable data engineering and distributed processing.
Hands-on Azure experience: Blob Storage, databases, compute resources, MLflow, Azure AI/ML services.
Solid grip on data quality, lineage tracking, observability, and production readiness.
ML lifecycle awareness — able to support model development, experimentation, deployment, and monitoring.
Experience supporting NLP/NLU use cases, including dataset and pipeline prep for language applications.
Excellent problem-solving skills and keen attention to detail.
Ability to participate in the discussions and lead the technical discussions
Have a consultancy mindset → always try to find a solution for the client
or
All done!
Your application has been successfully submitted!
You've already applied for this job
We appreciate your interest in this position. Unfortunately, you have already applied for this job.
