Bilingual MCQ Shuffle PDFs
Budget / Salary₹600–1,500
TypeFreelance project
LocationRemote
Posted49 minutes ago
I’m working with a technical reference book that contains roughly 9,245 multiple-choice questions presented side-by-side in Hindi and English. Every question shows four options and, in many cases, an embedded diagram or mathematical formula. The answer-and-explanation box sits separately at the bottom of each original page.
Your assignment is to script the entire extraction in Python (feel free to lean on PyPDF2, pdfplumber, Tesseract OCR, or any other robust libraries) so that:
• Every question, its four options, and any diagram are cleanly captured. Diagrams must remain as images; please don’t attempt to convert them to text.
• Answers and explanations are pulled out and stored apart from the questions.
• The full pool of ±9,200 questions is randomly shuffled, disregarding the book’s existing chapter order.
• The shuffled bank is broken into self-contained mock tests of exactly 100 questions each.
• Inside each test PDF, questions with their options appear first. Once the student reaches the end of those 100, a new page follows that lists the full answer key accompanied by its explanations. Hindi and English text must stay intact and legible.
Deliverables
1. A five-page pilot sample that proves diagrams render correctly, Hindi glyphs are preserved, and the layout meets the brief.
2. A folder of finished mock-test PDFs (around 93 files given the total count).
3. The reusable Python script plus any helper assets so I can re-run the workflow if the source book is updated.
Acceptance criteria will hinge on text fidelity (no broken matras or missing English characters), image clarity, correct shuffling, and exact 100-question segmentation with matching answer keys.
If you already have experience automating bilingual PDF extraction or handling diagram-heavy content, this should be a straightforward yet sizeable batching job. I’ll review the sample for accuracy before green-lighting the full run.
Your assignment is to script the entire extraction in Python (feel free to lean on PyPDF2, pdfplumber, Tesseract OCR, or any other robust libraries) so that:
• Every question, its four options, and any diagram are cleanly captured. Diagrams must remain as images; please don’t attempt to convert them to text.
• Answers and explanations are pulled out and stored apart from the questions.
• The full pool of ±9,200 questions is randomly shuffled, disregarding the book’s existing chapter order.
• The shuffled bank is broken into self-contained mock tests of exactly 100 questions each.
• Inside each test PDF, questions with their options appear first. Once the student reaches the end of those 100, a new page follows that lists the full answer key accompanied by its explanations. Hindi and English text must stay intact and legible.
Deliverables
1. A five-page pilot sample that proves diagrams render correctly, Hindi glyphs are preserved, and the layout meets the brief.
2. A folder of finished mock-test PDFs (around 93 files given the total count).
3. The reusable Python script plus any helper assets so I can re-run the workflow if the source book is updated.
Acceptance criteria will hinge on text fidelity (no broken matras or missing English characters), image clarity, correct shuffling, and exact 100-question segmentation with matching answer keys.
If you already have experience automating bilingual PDF extraction or handling diagram-heavy content, this should be a straightforward yet sizeable batching job. I’ll review the sample for accuracy before green-lighting the full run.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.