Location: Mountain View, CA — On-site
Granica builds AI infrastructure for enterprises operating massive data environments.
Our platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI.
Granica’s products include:
Crunch — continuous optimization for enterprise lakehouse data
Myelin — stateful infrastructure for long-running AI agents
Large Tabular Models — foundation models designed for enterprise tables
Together, we are building the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.
Granica has demonstrated approximately $200K in annualized value per petabyte and verified customer value within weeks.
Granica is hiring a Senior Software Engineer to build foundational data systems for AI.
You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data. This includes systems for metadata management, table maintenance, file layout optimization, distributed compute, and workload-aware data reorganization across petabyte- and exabyte-scale environments.
This is a hands-on engineering role for someone who has deep systems experience and wants to build at the intersection of data lakes, distributed systems, storage engines, query performance, and AI infrastructure.
You will work closely with Granica Research, led by Prof. Andrea Montanari at Stanford, to translate ideas from information theory, probabilistic modeling, compression, and learning efficiency into production systems.
Build metadata and transaction systems for large-scale tabular datasets
Design systems that support time travel, schema evolution, partition evolution, and atomic consistency
Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi
Optimize file layout, clustering, compaction, data skipping, indexing, and metadata pruning
Improve performance and cost efficiency across Spark, Flink, Trino, Presto, Databricks, Snowflake-adjacent, and cloud object storage environments
Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, and read-path optimization
Build adaptive engines that learn from access patterns and workloads to reorganize data automatically
Develop distributed compute pipelines that scale predictively and remain reliable under failure
Debug performance bottlenecks across storage, metadata, query execution, network, and compute layers
Implement research-driven algorithms in compression, representation, layout optimization, and data efficiency
Contribute to open-source or publish research when appropriate
Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure
Production experience with modern data lake or lakehouse technologies such as Spark, Iceberg, Delta Lake, Hudi, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems
Hands-on experience with columnar formats such as Parquet or ORC
Understanding of metadata-driven architectures, table formats, query planning, and physical data layout
Strong programming skills in Rust, Go, C++, or similar systems-oriented languages
Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency
A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end
Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems
Experience with compaction, clustering, manifests, snapshots, metadata catalogs, schema evolution, partition evolution, or table garbage collection
Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization
Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing
Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency
Build foundational infrastructure for enterprise data and AI
Work on deep systems problems across lakehouse data, metadata, storage layout, distributed compute, and AI efficiency
Partner directly with Research, Product, Engineering, and company leadership
Help shape Crunch, Granica’s production data optimization platform for enterprise-scale lakehouse environments
Work with a small, high-caliber team solving high-value infrastructure problems at massive scale
Have direct influence on architecture, product direction, customer outcomes, and company growth
Competitive salary, meaningful equity, and performance bonus for top performers
401(k) with company match, comprehensive health coverage, and unlimited PTO
Daily catered meals in our Mountain View office
Support for research, publication, and conference participation
At Granica, you'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.