We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

Staff Software Engineer, ML Data Infrastructure, Autonomy

Rivian
$206500.00-$258100.00
sick time, 401(k)
United States, California, Palo Alto
Sep 17, 2026
About Rivian

Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract.

As a company, we constantly challenge what's possible, never simply accepting what has always been done. We reframe old problems, seek new solutions and operate comfortably in areas that are unknown. Our backgrounds are diverse, but our team shares a love of the outdoors and a desire to protect it for future generations.


Role Summary

Rivian's Autonomy org needs a Staff Software Engineer, ML Data Infrastructure to own how autonomy data is described, indexed and accessed. This sits in the Platform Services team in the AI Platform organization in the Autonomy team. Our fleet of 100,000+ vehicles produces a continuous stream of multi-modal drive data, plus the output of every model run against it and every simulation executed on it.

The role requires deep expertise in columnar and analytical data stores, schema and format design, and the query and access layers that ML and analytics teams depend on.

You'll work with the AI Platform, Perception, Planning, Simulation, and Vehicle Integration, Product Management, and other technology partners. Autonomy data currently lives across multiple systems because no single store serves all our access patterns: fleet-scale byte storage, petabyte-scale analytics, sub-second fleet-wide search, and small transactional state each have different cost and latency profiles. You'll own the unified metadata and query layer, decide what consolidates and what stays specialized, and land the migrations onto a unified data lake.


Responsibilities

  • Own the unified data layer for autonomy: a single interface over document metadata, high-cardinality columnar analytics, real-time search, and object-stored sensor payloads, so engineers query concepts rather than databases.

  • Design the architecture for the datalake, metadata access and the APIs that expose it, giving engineers a single access layer

  • Own the data access and format strategy, including the canonical log format used in production today and the choice of ML-native columnar and vector access going forward. Define the canonical schemas, own their evolution, and drive the migration path.

  • Design and build batch and streaming pipelines for eval data, and the storage layer behind them, to handle both the volume and diversity of metrics from on-road and simulation runs at their actual cardinality.

  • Build the indexing and discovery layer for fast semantic and metadata search across the fleet's data: scenario tagging, event indexing, embedding-based similarity search, and the query surface engineers use.

  • Improve the data mining tools that apply ML techniques to data discovery, so Perception, Behavior and Planning engineers can find rare and long-tail scenarios at fleet scale rather than searching by hand.

  • Build well-documented tools and APIs so engineers outside data infrastructure can find, slice and materialize the data they need without writing a pipeline.

  • Treat data correctness as a discipline: schema validation, contracts between producers and consumers, freshness and completeness monitoring, and alerting that catches bad data before a model trains on it.

  • Build continuous testing and monitoring for the platform, covering ingest health, data freshness, schema conformance and query performance.

  • Design to the workflows of Perception, Behavior, Planning and Simulation engineers.

  • Work with the security & privacy team on retention, access control, consent handling and regional data requirements for vehicle-collected data.

  • Set data engineering standards across Autonomy and mentor engineers on schema design, query performance and pipeline reliability.


Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering or a related field, or equivalent experience.

  • 8+ years in software engineering, with a strong focus on data infrastructure or data platform work.

  • 5+ years owning a large-scale data platform for ML or analytics workloads, including its storage and access model, not only the pipelines on top of it.

  • 5+ years with columnar and analytical stores (ClickHouse, Pinot, Druid, BigQuery, Databricks or Snowflake), including data modeling for high-cardinality data and query performance tuning.

  • 3+ years designing storage and query layers for large volumes of heterogeneous metrics, such as KPIs or performance telemetry, where diversity is as hard a problem as volume.

  • 3+ years with document or NoSQL stores: schema design, indexing strategy, sharding and the operational realities at scale.

  • 3+ years with modern data and table formats (Parquet, Arrow, Iceberg, Delta), with clear reasoning about when each applies.

  • 3+ years building batch and streaming pipelines at scale with production orchestration and real data quality gates.

  • 3+ years hands-on with Python, plus production experience in Go, C++ or Rust for performance-sensitive parts of an access layer.

  • 3+ years with queueing and event-driven systems (SQS, Kafka, Kinesis or equivalent) in a production ingest path.

  • 2+ years with production monitoring and alerting for data pipelines (Prometheus, Grafana, Datadog or CloudWatch), including automated integrity and freshness checks.

  • 2+ years migrating a production data platform between storage or format architectures incrementally, while it stayed in service.

  • Engineering leadership: setting technical vision, timelines and priorities for a project or team, acting as technical lead, and mentoring engineers.

  • Ability to turn ambiguous, high-level requirements into a detailed system design and drive it to completion unprompted.

  • Technical excellence: willingness to work through implementation detail, and a record of raising technical standards across a broader engineering organization.

  • A record of designing schemas and interfaces that other teams built on and that survived changing requirements.

  • Nice to have: applied ML for data problems, such as data mining, active learning or embedding-based retrieval for rare and long-tail scenarios.

  • Nice to have: robotics or AV data experience, including ROS/ROS 2, MCAP, rosbag and time synchronization across sensor modalities.

  • Nice to have: vector databases and embedding-based retrieval for data mining and scenario discovery.

  • Nice to have: ML training data pipeline performance, including dataloader throughput, sharding and shuffling.

  • Nice to have: feature stores, data catalogs or lineage systems.


Pay Disclosure

The salary range for this role is $206,500-$258,100 for San Francisco Bay Area based applicants. This is the lowest to highest salary we in good faith believe we would pay for this role at the time of this posting. An employee's position within the salary range will be based on several factors including, but not limited to, specific competencies, relevant education, qualifications, certifications, experience, skills, geographic location, shift, and organizational needs.

We offer a comprehensive package of benefits for full-time and part-time employees, their spouse or domestic partner, and children up to age 26, including but not limited to paid vacation, paid sick leave, and a competitive portfolio of insurance benefits including life, medical, dental, vision, short-term disability insurance, and long-term disability insurance to eligible employees. You may also have the opportunity to participate in Rivian's 401(k) Plan and Employee Stock Purchase Program if you meet certain eligibility requirements. Full-time employee coverage is effective on their first day of employment. Part-time employee coverage is effective the first of the month following 90 days of employment. More information about benefits is available at rivianbenefits.com.



Equal Opportunity

Rivian is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, sex, sexual orientation, gender, gender expression, gender identity, genetic information or characteristics, physical or mental disability, marital/domestic partner status, age, military/veteran status, medical condition, or any other characteristic protected by law.

Rivian is committed to ensuring that our hiring process is accessible for persons with disabilities. If you have a disability or limitation, such as those covered by the Americans with Disabilities Act, that requires accommodations to assist you in the search and application process, please email us at candidateaccommodations@rivian.com.

Candidate Data Privacy and Technology

Rivian may collect, use and disclose your personal information or personal data (within the meaning of the applicable data protection laws) when you apply for employment and/or participate in our recruitment processes ("Candidate Personal Data"). This data includes contact, demographic, communications, educational, professional, employment, social media/website, network/device, recruiting system usage/interaction, security and preference information. Rivian may use your Candidate Personal Data for the purposes of (i) tracking interactions with our recruiting system; (ii) carrying out, analyzing and improving our application and recruitment process, including assessing you and your application and conducting employment, background and reference checks; (iii) establishing an employment relationship or entering into an employment contract with you; (iv) complying with our legal, regulatory and corporate governance obligations; (v) recordkeeping; (vi) ensuring network and information security and preventing fraud; and (vii) as otherwise required or permitted by applicable law.

Rivian may share your Candidate Personal Data with (i) internal personnel who have a need to know such information in order to perform their duties, including individuals on our People Team, Finance, Legal, and the team(s) with the position(s) for which you are applying; (ii) Rivian affiliates; and (iii) Rivian's service providers, including providers of background checks, staffing services, and cloud services.

Rivian may transfer or store internationally your Candidate Personal Data, including to or in the United States, Canada, the United Kingdom, and the European Union and in the cloud, and this data may be subject to the laws and accessible to the courts, law enforcement and national security authorities of such jurisdictions.

How We Use AI in Our Hiring Process: To ensure transparency, we want candidates to know that Rivian uses iCIMS Talent Cloud Artificial Intelligence (TCAI) and AI-enabled tools to assist with screening, reviewing, organizing and highlighting profiles and applications that match the key requirements for each role.

AI does not make hiring decisions: Qualified candidate applications are reviewed by a member of our team, and all decisions throughout the process are made by humans. We use AI to support efficiency and consistency, not to replace human judgment. We are committed to a fair, thoughtful, and equitable experience for every candidate.

Participation in AI profile matching is entirely voluntary. If you prefer that your profile not be used in this process, you can opt out at any time. Opting out means your profile will be excluded from automated matching and will not be surfaced for additional roles through this system. Your current application remains active and will not be affected in any way.

Please note that we are currently not accepting applications from third party application services.

Applied = 0

(web-665cd84569-2d8ll)