Sigma Software Group Logo

Sigma Software Group

Senior Data Engineer (Databricks Migration)

Posted 24 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Warsaw, Warszawa, Masovian
Senior level
In-Office or Remote
Hiring Remotely in Warsaw, Warszawa, Masovian
Senior level
Lead migration of a large-scale analytics platform from BigQuery to Databricks. Design and implement Lakehouse architectures, build and optimize PySpark data pipelines, implement medallion layers and Unity Catalog governance, troubleshoot complex SQL/Spark workloads, contribute to architecture and CI/CD, and support production deployments for retail analytics workloads.
The summary above was generated by AI
Company Description

Are you a Senior Data Engineer passionate about building scalable, high-performance data platforms and working with modern Lakehouse technologies? Join Sigma Software’s Data Engineering Center of Excellence and contribute to the modernization of an enterprise-scale analytics ecosystem for the retail domain.
We are looking for a Senior specialist with strong Databricks, PySpark, and cloud data engineering expertise to participate in the migration of a large-scale analytical platform from BigQuery to Databricks. You will collaborate with international teams, contribute to architectural decisions, and help shape reliable and scalable data solutions.
We at Sigma Software create opportunities for continuous learning, technology growth, and meaningful engineering impact while working on complex international projects.

CUSTOMER

Our Customer is a leading retail technology company specializing in AI-driven pricing optimization solutions for enterprise retailers. The company helps businesses improve profitability and competitiveness through advanced analytics, automation, and intelligent pricing strategies. Their platform combines business intelligence with sophisticated algorithms to support data-informed pricing decisions at scale for global retail organizations.

PROJECT

The project focuses on the strategic migration of a large-scale analytical platform from a legacy BigQuery ecosystem to a modern Databricks Lakehouse architecture. The platform processes high-volume retail datasets, machine learning workloads, analytics pipelines, and customer-specific business logic.
As part of the modernization initiative, the engineering team is implementing scalable Spark-based processing, Delta Lake architecture, medallion data layers, and modern governance practices. The role offers an opportunity to work with distributed data processing systems, optimize large-scale workloads, and contribute to the evolution of an enterprise-grade data platform.

Key Technologies: Databricks, Apache Spark, PySpark, Delta Lake, Python, SQL, Airflow, GCP, CI/CD, Unity Catalog

Job Description

  • Participate in the migration of a large-scale analytical platform from BigQuery to Databricks
  • Design and implement scalable Lakehouse architectures using Databricks and Delta Lake
  • Analyze existing ETL / ELT workloads and define migration approaches
  • Develop and optimize data pipelines processing large volumes of retail and analytical data
  • Implement incremental processing strategies and scalable transformation frameworks
  • Build and maintain Spark-based data processing solutions using PySpark
  • Design and maintain medallion architecture layers including Bronze, Silver, and Gold
  • Implement data governance and security best practices using Unity Catalog
  • Collaborate with Data Science, Analytics, Product, and Customer Engineering teams
  • Participate in architecture discussions and technical solution design
  • Develop reusable data platform components and engineering standards
  • Conduct code reviews and contribute to platform reliability and maintainability
  • Troubleshoot and optimize complex SQL and Spark workloads
  • Support production deployments and platform modernization activities

Qualifications

  • 5+ years of professional experience as a Data Engineer
  • Strong programming skills in Python and advanced SQL
  • Hands-on commercial experience with Databricks
  • Strong knowledge of Apache Spark, primarily PySpark
  • Experience designing and building modern cloud-based data platforms
  • Experience developing ETL / ELT pipelines and large-scale data processing solutions
  • Hands-on experience with Delta Lake
  • Experience with Spark Declarative Pipelines
  • Experience with cluster monitoring, metrics analysis, and performance optimization
  • Strong understanding of distributed data processing architectures
  • Solid understanding of data warehousing concepts and dimensional modeling
  • Experience with Airflow or similar orchestration tools
  • Experience optimizing complex analytical SQL workloads
  • Experience implementing CI / CD practices for data engineering platforms
  • Strong troubleshooting and performance optimization skills
  • Ability to work collaboratively in cross-functional international teams
  • Upper-Intermediate or higher English level

WILL BE A PLUS

  • Experience working with GCP cloud services
  • Experience with AWS or Azure cloud platforms
  • Experience in retail analytics or pricing optimization domains
  • Experience supporting machine learning or AI-related data workloads
  • Experience with platform modernization and cloud migration initiatives

Additional Information

PERSONAL PROFILE

  • Strong analytical and problem-solving mindset
  • Proactive and ownership-driven approach
  • Ability to work independently and collaboratively
  • Good communication and stakeholder collaboration skills
  • Passion for scalable data engineering and modern data platforms
  • Interest in continuous learning and technology innovation

Similar Jobs

3 Hours Ago
Remote
350K-457K Annually
Expert/Leader
350K-457K Annually
Expert/Leader
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Leads customer-focused design, development, deployment, and optimization of AI/ML solutions using Rovo. Architects integrations with APIs, microservices, and enterprise systems; manages production monitoring, evaluation, risk, privacy, and compliance. Influences product direction, aligns technical solutions with business goals, mentors engineers, and communicates with technical and non-technical stakeholders. The role requires extensive backend engineering and applied AI experience, including generative AI, LLMs, agent frameworks, and forward-deployed product development.
Top Skills: Agent-Based FrameworksAi/MlAPIsAtlassian IntelligenceGenerative AiJavaScriptLarge Language Models (Llms)MicroservicesPythonRovo
Yesterday
Easy Apply
Remote
Easy Apply
296K-444K Annually
Entry level
296K-444K Annually
Entry level
Cloud • Security • Software • Cybersecurity • Automation
Lead and grow GitLab’s Trusted Agentic Development engineering team. Define approaches for validating AI-generated software, improve agent feedback loops, support language-aware validation across Rust, Ruby, and Go, and drive adoption through useful platform capabilities. Partner with engineering teams, measure trust and usage, guide architecture and design, hire and develop engineers, manage performance, and lead complex work from planning through completion in a remote, asynchronous environment.
Top Skills: Ai AgentsGoRubyRust
Yesterday
Remote
101K-101K Annually
Junior
101K-101K Annually
Junior
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Manages operational execution of Pfizer’s clinical trial data-sharing program, from request submission through data package upload. Reviews and triages requests, coordinates stakeholders, tracks timelines, performs quality control, maintains program systems and databases, prepares metrics for leadership, supports compliance, and identifies process improvements. The role also supports the broader Expanded Access and Clinical Disclosure Team and requires occasional travel to Pfizer’s Warsaw site.
Top Skills: AIExcelMicrosoft PowerpointMicrosoft WordSharepoint

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account