Skip to content

Senior Data Engineer w/ NLP & ML

  • Remote
    • Prague, Praha, Hlavní město, Czechia
  • €35 - €41 per hour
  • Jimmy Technologies

If you are passionate about Data, AI, Python, and building next-generation ML workflows, this role with our client offers an exciting opportunity to work on cutting-edge R&D projects!

Job description

Our partner, a Fortune 500 company, is building an AI incubator to scout, incubate, and validate client and internal ideas on a 3–5 year horizon. The team develops technology roadmaps and prototypes that turn into advanced client solutions — and continually pushes cutting-edge AI into new products, services, and capabilities.

We're looking for a Senior Data Engineer to design, build, and operate reliable, production-ready data pipelines for large-scale data and ML use cases. You'll bridge raw data ingestion, transformation, and quality assurance with downstream machine learning workflows — with particular emphasis on NLP/NLU.

This is a long-term, remote-first contract position with required US working hours overlap (2-6 PM CET).

Responsibilities

  • Build & harden pipelines: design and maintain reliable ingestion pipelines that move raw data into production-ready platforms.

  • ETL/ELT: own data cleaning, transformation, schema design, quality validation, and lineage tracking.

  • Large-scale processing: develop data workflows using Databricks and PySpark.

  • ML enablement: prepare high-quality datasets, features, and pipelines for NLP, NLU, and broader ML use cases.

  • Azure infrastructure: operate and integrate Blob Storage, databases, compute resources, MLflow, and Azure AI/ML services.

  • CI/CD & MLOps: implement automated deployment, testing, monitoring, and promotion gates.

  • Own production readiness: ensure pipeline reliability, observability, performance, and data quality end-to-end.

Work Conditions

  • Type: Full-time & Long-term contract work

  • Start Date: ASAP

  • Location: Remote (99%) in Europe; must be able to travel freely within Europe for workshops.

  • US Time Zone Overlap: Required (2 PM - 6 PM CET)

  • Contract with European LLC

Job requirements

  • Strong Data Engineering background — ETL/ELT, pipeline development, schema design, orchestration.

  • Hands-on experience building reliable, production-ready ingestion pipelines with proper hardening, validation, and monitoring.

  • Strong Python and PySpark skills for large-scale data processing and transformation.

  • Practical Databricks experience — scalable data engineering and distributed processing.

  • Hands-on Azure experience: Blob Storage, databases, compute resources, MLflow, Azure AI/ML services.

  • Solid grip on data quality, lineage tracking, observability, and production readiness.

  • ML lifecycle awareness — able to support model development, experimentation, deployment, and monitoring.

  • Experience supporting NLP/NLU use cases, including dataset and pipeline prep for language applications.

  • Excellent problem-solving skills and keen attention to detail.

  • Ability to participate in the discussions and lead the technical discussions

  • Have a consultancy mindset → always try to find a solution for the client

or