Build and maintain NitroAI's data engineering infrastructure: own and scale the Airflow codebase, design and productionize pipelines, onboard new data sources, operate large-scale Spark pipelines on AWS Glue and support migration to Databricks, improve commercial pharma data models and lineage, and maintain internal Python packages while mentoring analytics teams.
Veeva Systems is a mission-driven organization and pioneer in industry cloud, helping life sciences companies bring therapies to patients faster. As one of the fastest-growing SaaS companies in history, we surpassed $3B in revenue in our last fiscal year with extensive growth potential ahead.
At the heart of Veeva are our values: Do the Right Thing, Customer Success, Employee Success, and Speed. We're not just any public company – we made history in 2021 by becoming a public benefit corporation (PBC), legally bound to balancing the interests of customers, employees, society, and investors.
As a Work Anywhere company, we support your flexibility to work from home or in the office, so you can thrive in your ideal environment.
Join us in transforming the life sciences industry, committed to making a positive impact on its customers, employees, and communities.
The NitroAI team is seeking a Senior Data Engineer to build and maintain the data engineering infrastructure that powers our analytics delivery. You'll work alongside data scientists and analytics teams to productionize manual solutions, build governed datapipelines, and deliver data to customers at scale. This is a high-impact, high-ownership role in a startup-like environment within Veeva — with immediate influence over how NitroAI delivers data across a growing multi-tenant platform.
What You'll Do
- Own the Airflow codebase end-to-end (MWAA, ~100 DAGs currently): build reusable templates, scale patterns, enforce standards, improve testing infrastructure
- Serve as delivery teams’ go-to on pipeline architecture and troubleshooting
- Own data onboarding when we connect to a new source including connection setup, schema discovery, initial pipeline design
- Operate and extend large-scale Spark pipelines on AWS Glue, including multi-TB joins and compaction jobs, and support migration of those workloads to Databricks as we move platforms
- Drive data model improvements around our commercial pharma data (patient claims, KOL and HCP data, CRM activity) with a focus on structure, lineage, and how it flows through the platform
- Contribute to and maintain our internal Python package used across the data team
Requirements
- 5+ years building data models and pipelines
- 2+ years production Airflow experience
- 2+ years Spark at scale: hands-on experience with multi-TB datasets; AWS Glue experience preferred
- Strong python and SQL
- AWS fluency in S3, ECS/Fargate, Glue, IAM, Secrets Manager
- Clear communicator who can teach: onboarding analytics team to the codebase and developing team Airflow capability is a core part of this job, not a side responsibility
Nice to Have
- Experience working with privacy-sensitive or governed data
- Databricks experience – we’re actively migrating Glue workloads there
- Proficiency with Claude Code
- Exposure to data science workflows and ML pipeline tooling
- Background in life sciences or healthcare
- Experience with Open Meta data or similar data catalog and data dictionary tooling
Interviewing with Veeva
- Follow the application process and submit your resume.
- Within 3 days, you will receive a link to a personality assessment administered by a third party.
- Once you complete the assessment, our team will review your full application package and follow up via email with our decision.
- If moving to the interview stage, the process is as follows:
- A conversation with the hiring manager
- A practical case exercise
- A final conversation with our group's Senior Leader.
- Once all interviews are complete, the manager will be in touch with a final decision.
We value your time and believe in a transparent hiring process. Here is the process you can expect.
Perks & Benefits
- Medical, dental, vision, and basic life insurance
- Flexible PTO and company paid holidays
- Retirement programs
- 1% charitable giving program
Compensation
- Base pay: $115,000 - $175,000
- The salary range listed here has been provided to comply with local regulations and represents a potential base salary range for this role. Please note that actual salaries may vary within the range above or below, depending on experience and location. We look at compensation for each individual and base our offer on your unique qualifications, experience, and expected contributions. This position may also be eligible for other types of compensation in addition to base salary, such as variable bonus and/or stock bonus.
#LI-RemoteUS#LI-MidSenior
Veeva’s headquarters is located in the San Francisco Bay Area with offices in more than 15 countries around the world.
Veeva is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity or expression, religion, national origin or ancestry, age, disability, marital status, pregnancy, protected veteran status, protected genetic information, political affiliation, or any other characteristics protected by local laws, regulations, or ordinances. If you need assistance or accommodation due to a disability or special need when applying for a role or in our recruitment process, please contact us at [email protected].
Similar Jobs
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Lead design and operation of large-scale data pipelines and data APIs. Build Spark/PySpark workflows on Databricks, optimize job performance, manage data quality and observability, and develop MCP servers and AI-agent integrations. Mentor engineers, define standards, support production incidents and on-call rotations, and collaborate with stakeholders to deliver scalable data platform products.
Top Skills:
Apache IcebergApi GatewayAWSAws LambdaAws Rds/AuroraAzureDatabricksDatadogDbtFastapiFivetranGCPGoogle BigqueryLlms/Ai AgentsMcp ServersMs Sql ServerMySQLOraclePostgresPysparkPythonS3SecretsmanagerSnowflakeSnsSparkSplunkSQLSqs
Fintech • Information Technology • Software
Own and evolve the data infrastructure powering identity and fraud products. Build and operate scalable ETL/ELT pipelines, ensure data quality and observability, optimize storage and access, support product squads, respond to production incidents, mentor engineers, and drive data platform innovation.
Top Skills:
AWSAws LambdaColumnar Data StoresDockerEksEmrGCPGoHadoopInfrastructure-As-CodeKafkaKubernetesAzureOpensearchPostgresPythonRdsRedshiftS3SnsSparkSqsVector Db
Kids + Family • Mobile
Design, build, and maintain scalable distributed data pipelines and a secure data lakehouse for streaming and batch processing to support real-time analytics, ML, and experimentation. Automate, test, and harden workflows, architect logical and physical data models, build ML model features, and collaborate with product, analytics, and data science teams to turn data into value.
Top Skills:
AirflowAWSAzureBigQueryDabsData LakehouseDatabricksDatabricks WorkflowsDbtGCPGithub ActionsLlmsPrestoPythonSnowflakeSparkSQLTerraformTrino
What you need to know about the Charlotte Tech Scene
Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.
Key Facts About Charlotte Tech
- Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
- Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
- Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
- Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
- Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus



