Responsibilities:
Design, develop, deploy, and operate stateful real-time data pipelines using Apache Flink, including keyed state, windowing, watermarks, timers, side outputs, and custom sources/sinks.
Build and maintain Change Data Capture (CDC) pipelines using technologies such as Flink CDC or Debezium, including snapshot and incremental processing, schema evolution, and downstream idempotency.
Design and operate Apache Kafka and Kafka Connect pipelines, including topic and partition strategies, source/sink connectors, schema management, converters, and dead-letter routing.
Own and optimize end-to-end data flows between Kafka, Flink, and data stores such as PostgreSQL, TimescaleDB, ClickHouse, and BigQuery.
Design data models, optimize storage and processing strategies, and ensure reliability, scalability, and performance of high-volume data systems.
Troubleshoot and resolve complex production issues, including checkpoint failures, backpressure, state growth, sink performance issues, autoscaling challenges, and recovery failures.
Improve platform reliability through delivery guarantees, schema migrations, observability, alerting, and operational best practices.
Contribute to infrastructure and deployment processes, including Helm charts, ArgoCD applications, Kubernetes configurations, monitoring dashboards, and alerting rules.
Participate in code reviews, mentor engineers, and influence technical decisions around architecture, scalability, and system evolution.
Requirements:
Bachelors degree in Computer Science, Software Engineering, Information Systems, or a related technical field, or equivalent practical experience.
5+ years of production experience with Java, including strong knowledge of modern Java features (records, sealed types, switch expressions, streams).
Strong understanding of JVM internals, including heap behavior, garbage collection, class loading, and runtime optimization.
Hands-on experience building and operating production-grade streaming data platforms using Apache Flink, Kafka Streams, Spark Structured Streaming, or similar technologies.
Proven experience with Apache Flink, including keyed state, RocksDB vs. heap state backends, savepoints, Flink Kubernetes Operator, and autoscaler tuning.
Deep understanding of Apache Kafka, including delivery guarantees, transactional producers, consumer groups, offset semantics, and schema registries.
Experience designing, deploying, and operating Kafka Connect pipelines in distributed environments.
Production experience with CDC technologies such as Flink CDC or Debezium, including database log-based replication, snapshots, incremental processing, schema evolution, backfills, and failure handling.
Strong experience with databases such as PostgreSQL, ClickHouse, and BigQuery (or equivalent), including schema design, query optimization, partitioning strategies, and performance troubleshooting.
Experience with TimescaleDB, hypertables, PostGIS, and spatial data processing patterns.
Experience with cloud-native deployment environments, Kubernetes, monitoring, and production operations.
Strong problem-solving skills and ability to work independently in a fast-paced environment.
cv+6566@hrhome.co.il