Database Systems Engineer
3 days ago
Madrid
ppWe’re looking for a Database Systems Engineer who thrives at the frontier of distributed databases, storage engines, transactional systems, and analytical query execution. /p h3Role Summary /h3 pFastS3 is building AI-native data infrastructure from the ground up. We are looking for an engineer to help design the database layer that sits above PixelDB, our distributed key/value object storage engine, and works with TEMPANO, our custom Iceberg catalog and compaction API. This role is not about building a conventional application database, a Postgres extension, or a simple lakehouse connector. It is about helping design the foundation for a system that can support transactional, analytical, and vector workloads over the same live data layer. You’ll work closely with engineering leadership to prototype and build the core components of a distributed table engine: commit logs, MVCC, stable Row IDs, secondary indexes, snapshot publication, compaction, query execution, and integration with Iceberg-style table metadata. /p h3What You’ll Be Doing /h3 ul liDesign and prototype a transactional table engine on top of PixelDB. /li liWork on MVCC, commit logs, transaction visibility, snapshot isolation, and serializable consistency. /li liDesign stable Row ID abstractions so indexes point to logical rows rather than physical storage locations. /li liBuild metadata and manifest flows that connect live transactional writes to TEMPANO and Iceberg-compatible analytical snapshots. /li liHelp design compaction, version cleanup, and garbage collection that do not block long-running analytical queries. /li liExplore query execution paths for point reads, range scans, analytical scans, joins, and vector retrieval. /li liWork on distributed storage layout, partitioning, object placement, erasure-coded data, and recovery semantics. /li liCollaborate with systems and AI engineers to support agentic workloads that require transactional, analytical, and semantic access to live data. /li liBenchmark latency, throughput, consistency behavior, write amplification, and query performance under concurrent workloads. /li /ul h3What We Need to See /h3 ul liStrong systems programming experience in Rust, C++, Go, or C. /li liPractical knowledge of database internals, storage engines, distributed databases, or query engines. /li liExperience with MVCC, WAL/commit logs, transaction processing, indexing, compaction, or recovery. /li liUnderstanding of distributed systems concepts such as consensus, replication, failure recovery, consistency, and concurrency control. /li liFamiliarity with analytical formats or engines such as Apache Iceberg, Delta Lake, Parquet, Arrow, Trino, Spark, DuckDB, or ClickHouse. /li liAbility to reason deeply about tradeoffs between OLTP, OLAP, and vector workloads. /li liComfort moving from architectural design to prototype code, benchmarks, and production-quality implementation. /li liStrong debugging, profiling, and performance‑analysis skills. /li /ul h3Ways to Stand Out /h3 ul liExperience building or contributing to database kernels, storage engines, query planners, transaction managers, or distributed SQL systems. /li liHands‑on work with PostgreSQL internals, CockroachDB, YugabyteDB, TiDB, FoundationDB, ScyllaDB, Cassandra, ClickHouse, DuckDB, Neon, or similar systems. /li liExperience with Apache Iceberg, Delta Lake, table catalogs, manifest generation, compaction, or lakehouse metadata. /li liKnowledge of vector indexes such as HNSW, IVF, PQ, ANN search, or hybrid vector/SQL execution. /li liExperience designing systems around stable logical IDs, append‑only storage, object storage, LSM trees, B‑trees, or log‑structured architectures. /li liBackground in high‑performance networking, distributed storage, RDMA, erasure coding, or object storage. /li liOpen-source contributions, papers, or deep technical writing related to databases, storage engines, or distributed systems. /li /ul h3Why Join Us? /h3 pFastS3 is building infrastructure for the next generation of AI‑native data systems. Our vision is to move beyond fragmented pipelines, replicated warehouses, and bolt‑on vector extensions by designing a data layer where transactional, analytical, and semantic workloads can operate over the same live data. You’ll join a team working at the intersection of distributed storage, database internals, lakehouse architecture, vector retrieval, and AI infrastructure. If you’re excited by the idea of building a new database architecture instead of extending yesterday’s systems, this is the role. We’re growing our team in Madrid and looking for engineers who want to build core infrastructure from first principles. /p pFastS3 is proud to be an inclusive, equal opportunity employer committed to diversity, equity, and accessibility for all. /p /p #J-18808-Ljbffr