Repair Hadoop Pipeline & Hive Output

via Freelancer ·

Budget / Salary$30–250
TypeFreelance project
LocationRemote
Posted1 hour ago
Our production datalake pipeline on Hadoop is misbehaving: Spark and Flink jobs crash intermittently, completed stages fail to commit their results to the Hive tables, and overall throughput has slowed to a crawl. I need someone to dive in, trace the failures, and leave me with a clean, fully functioning data flow.

You will have direct access to the existing Spark and Flink code, YARN cluster dashboards, Hive metastore, and any relevant logs. The immediate goal is to identify and correct the root causes of the job failures and the missing Hive writes, then tune the pipeline so it can keep up with daily load without time-outs or excessive retries.

Acceptance criteria – once you are done:
• All scheduled Spark and Flink jobs finish successfully when triggered manually and by the scheduler.
• Output partitions appear in the target Hive table with correct row counts on at least two consecutive test runs.
• End-to-end execution time returns to normal (or faster) baseline and remains stable for 48 hours of monitoring.
• A short hand-off document summarises the changes, configs touched, and any recommended follow-up work.

If this sounds straightforward to you, let’s get started—I’m ready to grant cluster access as soon as we agree on an approach.
mysql big data sales hadoop elasticsearch hive spark etl big data
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.