About

Software engineer building petabyte-scale lakehouse and AI data infrastructure at Huawei Cloud, with ownership from technical design through implementation, benchmarking, and production hardening. Author of 50 merged Apache Hudi PRs and three merged Lance PRs; the Flink work reduced end-to-end TPC-H write time by up to 26% and inter-operator data transfer by 29–36%. Current work includes a two-tier distributed cache for Lance/Ray and a query-adaptive, incrementally updatable ANN index targeting tens of billions of vectors. Creator and operator of open-source AI developer tooling, with 8,400+ agent runs recorded across 15 repositories.

Technical skills

Languages: Java, Python, SQL; Scala (working knowledge)

Distributed data and retrieval: Apache Hudi, Apache Spark, Apache Flink, Lance, Ray, PostgreSQL, Parquet; lakehouse internals, distributed caching, vector search and ANN

Performance and reliability: Serialization, checkpointing, partitioning/load balancing, profiling, memory/GC optimization, TPC-H benchmarking, monitoring

AI developer tooling: Codex, Claude Code, multi-agent orchestration, task decomposition, human-in-the-loop (HITL) workflows, durable state/recovery, run budgets, trajectory tracing, cost analytics

Engineering delivery: Technical designs/RFCs, automated testing, CI/release engineering, Maven, Docker, Linux

Experience

May 2023 – Present

Software Engineer / Big Data Engineer

Huawei

Develop core lakehouse, distributed data, and AI retrieval infrastructure for Huawei Cloud; support teammates through code and RFC reviews and technical tutorials.

• Designed and implemented a two-tier distributed cache for Lance vector indices with an in-memory L1, persistent per-worker NVMe L2, and IVF-partition prewarming and routing; integrated it with Ray and benchmarked it on a real cluster.

• Developing a distributed, incrementally updatable, query-adaptive ANN index targeting tens of billions of vectors; authored three merged upstream Lance PRs that bounded IVF-training memory and fixed temporary-resource lifecycle issues, and helped resolve an out-of-memory problem in a three-billion-vector workload.

• Across 50 merged Apache Hudi PRs, authored RFC-84 and implemented its Flink serialization redesign, cutting TPC-H write time by up to 26% and inter-operator data transfer by 29–36%; released in Hudi 1.0.2.

• Four additional Flink optimizations ([1], [2], [3], and [4]) improved throughput by 10% and reduced GC overhead by 30%; released in Hudi 1.0.1.

• Diagnosed serial buffer flushing during Flink checkpoints, contributing to a fix that improved checkpoint performance 7×; corrected writer load balancing and timestamp, CDC-merge, and cross-engine record-key defects.

• Simplified setup across 900+ Hudi configuration parameters; designed and implemented partition-level TTL for automated retention and storage-cost control, supported by automated tests and a 100,000-partition load test.

• Drove Hudi’s Spark DataSource V2 migration; authored RFC-98 and an initial Copy-on-Write read path to enable additional query pushdowns.

Feb 2022 – May 2023

Software Engineer / ML Engineer

Digital Research (computer vision startup)

• Designed the event-driven Python backend for a production truck-monitoring system processing ~20,000 images per day; the deployed system reduced fleet idle time by 12%. Also delivered customer reporting/data-visualization and internal monitoring interfaces.

Open-source engineering

chipping-orchestrator – Creator & Maintainer | Apr 2026 – Present

• Designed and shipped a Python issue-to-PR platform using Codex and Claude Code, with task decomposition, dependency-aware scheduling, isolated git worktrees, automated verification, independent review, and human-controlled merge.

• Built GitHub-persisted workflow state, pause/resume, credential isolation, and crash recovery for interrupted rebases and squashes. Added cumulative PR-size limits and per-issue lifetime run budgets, with explicit human authorization for oversized changes.

• Added trajectory tracing and PostgreSQL/Streamlit dashboards for token usage, cost, and workflow outcomes. Shipped 14 releases through v0.12.0 and operated the platform across 15 repositories, recording 8,400+ agent runs, including 2,700+ independent review-agent runs.

Education

PhD, Geophysics

Trofimuk Institute of Petroleum Geology and Geophysics SB RAS

MSc, Computational and Applied Mathematics

Novosibirsk State University

Certificates

• Deep Learning Specialization, Coursera (2021)

Software copyrights

• Software for solving nonstationary thermohydraulic problem applied to reactors and experimental stands with sodium, lead and lead-bismuth coolants. Version 1.1. HYDRA-IBRAE/LM/V1.1 (2018)

Hobbies

• Going to the gym on a regular basis (since 2019).

• Reading developmental psychology and business management books.

- Meg Jay "The Defining Decade"

- Tom DeMarco "The Deadline: A Novel about Project Management"

- Alexey Markov "Hoolinomics" and "Greedology" [in Russian]

- Eliyahu Goldratt, and Jeff Cox "The Goal: A Process of Ongoing Improvement"