Active
Developer Tools
2024 - 2025Lead DeveloperCN Question Extractor
Automated academic question parsing and question bank formatting engine.
Executive Overview
Academic institutions, test preparatory centers, and study app developers frequently manage thousands of questions locked within messy PDFs, scanned papers, and Word files. CN Question Extractor automates the parsing and structuring of this content into machine-readable datasets.
The Problem & Motivation
Manual data entry of exam papers into databases is expensive, slow, and prone to typographic errors. Existing OCR and text extraction software outputs raw unsegmented text blocks that require hours of human cleanup to separate questions from multiple-choice options.
Architecture & System Design
A multi-stage parsing pipeline built with Python, combining structural heuristics, regex state machines, and document layout analysis.
Architectural PipelineRaw Document -> Layout Extraction -> State Machine Tokenizer -> Regex Rule Filter -> Schema Validation -> Structured JSON
State Machine Parsing: Tracks parser state across question titles, stems, mathematical formulas, numbered options (A/B/C/D), and marks allocation.
JSON-Schema Validation: Guarantees that extracted items strictly match the standardized schema required by modern database ingestion endpoints.
Technical Challenges & Solutions
1Challenge: Disentangling multi-line questions with nested sub-questions (e.g. 1(a), 1(b)(i)) and varied numbering styles.
Engineered Solution: Developed an adaptive hierarchy tokenizer that dynamically identifies indentation depth and alphanumeric numbering conventions.
2Challenge: Handling corrupted OCR text with irregular whitespace and non-standard bullet characters.
Engineered Solution: Engineered a normalization pre-processor that standardizes unicode characters, cleans typographical ligatures, and fixes broken line wraps before parsing.
Key Outcomes & Impact
- Accelerated exam paper digitization speed by more than 10x compared to manual transcriptions.
- Successfully extracted over 15,000 standardized questions across multiple academic curricula into clean JSON databases.
- Achieved a 98.4% accuracy rate on standardized test paper layouts.