Project overview
We revolutionized the data extraction process for a Legal Tech firm by replacing fragile, single-purpose scrapers with a robust Java and Spring Boot platform. Beyond OCR/OpenCV for documents, most of the codebase focused on web scraping and regex-based data extraction, integrating anti-captcha APIs and multiproxy rotation to work around anti-bot blocks and keep collection stable across difficult public sources.
Challenge
High maintenance cost of hundreds of fragile web scrapers, manual data entry from PDFs, and frequent blocks from sites with anti-bot protection.
Solution
Unified Java/Spring Boot platform with parameterizable scraping, regex parsers, OCR/OpenCV, anti-captcha integration, and multiproxy rotation.
Tech Stack
- Java
- Spring Boot
- Automation
- Web Scraping
Technical scope
- Java and Spring Boot
- Parameterizable scraping with regex
- OCR/OpenCV and document analysis
- Anti-captcha and multiproxy rotation
