Torc Robotics Software Engineer II - Data Engineering at Torc Robotics building AWS-native data ingestion, ETL, and storage solutions for autonomous vehicle data. Supports scalable pipelines and data lake expansion for ML training and analytics.
Responsibilities
We are seeking a Software Engineer with experience in implementing and maintaining Linux and cloud-based data management solutions who is ready to extend that experience. In this role you will be working closely with other engineers on the implementation of AWS-native data ingestion, ETL, and storage solutions that power data driven analytics, simulation, and ML training across the enterprise.
Create robust and resilient pipelines to process massive daily volumes of data created by vehicle fleets and simulation environments
Build and support scalable pipelines as part of Torc’s Data Factory to delivery data for ML training at scale
Scale Torc’s data lake through a distributed storage system, data crawling and discovery
Promote and protect the integrity of data through validation, versioning, data provenance, and governance
Support the expansion of Torc’s data lake through acquisition of additional data sets from internal and external sources
Contribute to the design, architecture and delivery of cloud-based solutions
Collaborate with teams specializing in perception, planning, control, mapping and vehicle testing to develop solutions that support product delivery
Support the implementation of emerging cloud-based capabilities that can extend our technology stack and improve our ability to build, deploy and test safety-critical software for self-driving vehicles
Participate in the team’s on-call rotation to support our deployed systems during business hours
Here’s a list of some of the technologies we use to make all the above happen:
Bachelor’s Degree in Computer ScienceMaster’s Degree in Computer ScienceExperience deployingStrong organizationalExperience with the Databricks platformExperience with pandas
Required
Bachelor’s Degree in Computer Science, Robotics, Electrical Engineering or related technical field plus demonstrates competences and technical proficiencies typically acquired through 4+ years of experience or;
Master’s Degree in Computer Science, Robotics, Electrical Engineering or related technical field plus demonstrates competences and technical proficiencies typically acquired through 0-3+ years of experience.
Experience deploying, troubleshooting, monitoring and maintaining Linux systems
Experience building and maintaining workloads in public cloud environments
Knowledge of different database architectures, including but not limited to relational and NoSQL databases, vector stores, data warehousing and clustered, distributed data stores
Practical experience with Linux and general bash scripting
Practical experience with Docker and containerization
Implementation of workflow patterns using directed acyclic graphs (Apache Airflow, AWS Step Functions)
A strong commitment to test-driven development patterns, continuous integration and delivery, and infrastructure as code
Preferred
Strong organizational, time management, and communication skills working with a team orientation and collaborative style
Experience with the Databricks platform, particularly for serving data, visualizations and jobs
Experience with scaling data for ML and AI workloads using Ray
Experience with pandas, numpy and other Python-based data analysis libraries and tooling
Deep knowledge of AWS serverless architectures (Lambda, Batch, ECS Fargate, Glue, Athena)
Experience with data storage and acquisition patterns for robotics and advanced driver assistance systems
Own end-to-end architecture and delivery of cloud-based solutions.