What are the responsibilities and job description for the Senior Software Engineer-Data Platform position at AccrueTalent?
Hiring the first dedicated engineer for its data platform. Today the platform is nascent, solid backend infrastructure exists, but nobody owns it, and the perception/ML team is being handed data that isn't in the shape they need. This is a founding-level, zero-to-one seat: you'll define how raw maritime signals become clean, consistent, labeled datasets for perception and foundation models.
You'll architect high-throughput pipelines that ingest real-time acoustic and telemetry data, then align, resample, calibrate, and restructure it, not necessarily in the moment (it can be matched up every 30 minutes, hour, or day) but into exactly what the ML team needs. Expect to build backfill/reprocessing frameworks, lineage and versioning for reproducibility, dataset discovery APIs, data-quality instrumentation, and, likely, a labeling tool from scratch. There's real backend infra to build on, but the shape of the platform is yours to define.
What You'll Own
- Post-processing pipelines that align, resample, and calibrate multi-sensor data
- Backfill/reprocessing frameworks for new filters, syncs, label corrections, and metadata enrichment across historical data
- Lineage and versioning to guarantee experiment reproducibility
- Dataset discovery access APIs/SDKs (query by time, region, modality, labels, quality flags)
- Data-quality metrics, dashboards, and alerts; canary dataset builds
- Storage-layout optimization (columnar formats, compression, chunking, sharding, prefetching)
- Likely build a labeling tool and pre-labeling workflows for the perception team
- Ramp: month 2, a first working pipeline in place; months 3–6 — iterating and honing it to exactly what the ML/perception team needs, plus the tooling around it
Requirements
- Strong data-pipeline architecture, high-throughput, with the ability to architect the system from a blank page
- Solid knowledge of at least one cloud provider, preferably AWS
- Comfortable deploying pipelines to the cloud
- Python data tooling (PyArrow/Polars/Pandas, NumPy/SciPy), plus one of Go/Rust/TypeScript for services
- Bonus: full-stack/generalist range — able to build the tools the ML team needs end-to-end
Execution and Ownership
- Comes in and builds day one with minimal hand-holding
- Low ego; takes criticism without taking it personally
- High autonomy; startup-native
Background
- ~5–8 years; startup time counts double
- Architected data pipelines / greenfield data-platform work; dataset-as-a-product ownership is a strong signal
- Exposure to edge/sensor data (audio/sonar, video, telemetry), time sync, and geospatial context is a plus
Location and Visa
- Remote OK; LA strongly preferred, with occasional on-site visits to Torrance
- US citizenship required
Nice-to-Have
- Labeling workflows (interfaces, ontologies, consensus, QA) and label-store integrations
- Splitting/sampling strategy design (by time, platform, geography, class, SNR) to avoid leakage
- Orchestration (Airflow, Prefect) and metadata/lineage (MLflow, W&B)
- Dataset-as-a-product track record with strong lineage and documentation
- Athletic or competitive background (e.g., competitive chess) — reads as dynamic and startup-fit
Who Will Thrive Here
- The data engineer who wants to own an entire platform employee-early
- Architect-operators who can go from block diagram to shipped pipeline without hand-holding
- Zero-to-one builders who've stood up data platforms from scratch and like greenfield
- Low-ego, autonomous, comfortable in crunch