AI Video Engineer for ROLLCALL Backend

via Freelancer ·

Budget / Salary$30–250
TypeFreelance project
LocationRemote
Posted1 hour ago
Senior AI Video / GPU Engineer Needed – Wan2.2 + RunPod A100 + PyTorch/CUDA

I need an experienced AI/GPU engineer to finish and productionize an existing AI video-generation backend for a platform called ROLLCALL.

This is NOT a website design job. The website is already built. I need someone who specializes in GPU inference, Python, PyTorch/CUDA environments, AI video models, and production API deployment.

CURRENT SYSTEM

We already have:

RunPod
NVIDIA A100-SXM4 80GB GPU
Wan2.2 I2V A14B
Approximately 118GB of Wan2.2 model files already downloaded
Persistent /workspace storage
Python
PyTorch/CUDA
FastAPI/Uvicorn worker
FFmpeg
Existing website integration
Existing REST API running on port 3010

Current API routes include:

GET /v1/health
POST /v1/generate
GET /v1/jobs/{job_id}
GET /v1/files/{filename}

The website can communicate with RunPod successfully.

GPU health verification works.

The generation API successfully accepts a request, creates a real job_id, returns that job ID to the website, and allows the website to poll the job status.

CURRENT TECHNICAL PROBLEM

Wan2.2 is currently failing during Python startup because of a FlashAttention/PyTorch ABI compatibility problem.

Current error includes:

flash_attn_2_cuda...so: undefined symbol...

I do NOT want someone randomly reinstalling packages until something works.

I want an engineer who understands PyTorch, CUDA, NVIDIA GPUs, FlashAttention, compiled Python/CUDA extensions, ABI compatibility, and production AI inference environments.

YOUR JOB

First audit the complete existing environment before changing anything.

Determine the correct compatible combination of:

NVIDIA driver
CUDA
Python
PyTorch
torchvision/torchaudio where applicable
FlashAttention
Wan2.2 dependencies

Then repair or rebuild the Python runtime cleanly if necessary.

You must also:

Preserve the existing ~118GB Wan2.2 model files.
Preserve existing ROLLCALL project files and generated videos.
Verify Wan2.2 imports successfully.
Verify generate.py starts successfully.
Verify T5, VAE and Wan2.2 model checkpoints.
Verify A100 CUDA inference.
Verify FFmpeg.
Review and stabilize the FastAPI worker.
Verify generation requests immediately return a valid job ID.
Verify job status can be polled reliably.
Implement reliable queued, generating, processing/stitching, completed and failed states.
Return useful error information when a generation fails.
Ensure completed MP4 files are accessible to the website.
Verify image-to-video generation using Wan2.2 I2V A14B.
Test the entire system from the ROLLCALL website through RunPod and back.
Make the worker start/recover correctly after a pod/container restart.
Create a reproducible startup/deployment configuration so we do not have to manually repair the environment again.
Document the exact final Python/PyTorch/CUDA/FlashAttention/package versions.
IMPORTANT

Do NOT delete, replace or redownload the approximately 118GB Wan2.2 model unless there is a genuine technical reason and I approve it first.

Do NOT consider the project finished simply because the health endpoint returns HTTP 200.

DEFINITION OF DONE

I will consider this project complete only when this complete workflow works:

ROLLCALL Website → RunPod API → Job Created → Wan2.2 I2V A14B → NVIDIA A100 GPU Generation → MP4 Created → Job Completed → Video Returned → Video Plays Correctly on ROLLCALL Website

A real end-to-end video generation must successfully complete before final acceptance.

I also want the completed environment documented and reproducible so a future RunPod restart does not require rebuilding everything manually.

REQUIRED EXPERIENCE

Please apply only if you have strong experience with several of the following:

Python, PyTorch, CUDA, NVIDIA A100/H100, RunPod, Wan2.1/Wan2.2, FlashAttention, Hugging Face, diffusion/video models, FastAPI, REST APIs, Linux, FFmpeg, Docker and production GPU inference.

Experience deploying large AI video models such as Wan2.x, HunyuanVideo, CogVideoX, Stable Video Diffusion or similar models is strongly preferred.

IMPORTANT SCREENING QUESTION

Start your proposal with:

A100-WAN22

Then answer this question:

If flash_attn_2_cuda.so produces an undefined symbol error when importing Wan2.2, what would you inspect before reinstalling FlashAttention?

Explain specifically how you would verify PyTorch version, CUDA compatibility, NVIDIA driver, Python version, FlashAttention build/wheel compatibility and ABI compatibility.

Do not simply answer “I will reinstall FlashAttention.”

Generic proposals will not be considered.

I am looking for someone who can inspect the entire stack, identify the root cause, make the correct repair once, and deliver a production-ready system.
python cuda flash animation pytorch sd-wan fastapi hugging face nvidia jetson
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.