Inetum logo

Senior Python Data Engineer (OCR & Document Processing)- remote

Inetum
On-site
Posted 15 days ago

AI summary

Designs and builds scalable data ingestion and document processing solutions for transforming unstructured information into AI-ready data, focusing on OCR and document intelligence.

Eligible from: Unclear

Job Description

Mission

Design, build, and optimize scalable data ingestion and document processing solutions that transform large volumes of unstructured insurance data into structured, AI-ready information. Enable downstream AI and retrieval systems by leveraging OCR, document intelligence, vector databases, and cloud-native data pipelines.

Responsibilities:

  • Design and implement scalable data ingestion pipelines for processing high volumes of unstructured documents, including PDFs, scans, emails, and Office files.
  • Integrate, configure, and optimize OCR and document extraction technologies to maximize text extraction accuracy and document understanding.
  • Build automated workflows for document parsing, text cleaning, normalization, semantic chunking, and metadata enrichment.
  • Develop connectors and integrations for document sources such as SharePoint, email systems, and enterprise repositories.
  • Design and maintain vector database schemas and retrieval mechanisms to support Retrieval-Augmented Generation (RAG) solutions and AI applications.
  • Ensure document processing pipelines meet enterprise security, compliance, performance, and availability requirements.
  • Implement monitoring, validation, and quality-control mechanisms to identify and manage low-confidence OCR and extraction results.
  • Optimize data processing workflows for scalability, reliability, and low-latency operations.
  • Collaborate with AI Engineers, Backend Engineers, and Platform teams to deliver end-to-end AI-powered document processing solutions.
  • Develop and maintain cloud-native data ingestion solutions on public cloud platforms.

Requirements

Profile

Professional Experience

  • 5-10 years of experience in Data Engineering, Data Processing, Document Intelligence, or related fields.
  • Proven experience building scalable data ingestion and processing pipelines.
  • Experience working with large volumes of unstructured and semi-structured data.
  • Experience designing cloud-based data solutions.

Technical Skills

  • Strong programming skills in Python.
  • Strong SQL knowledge.
  • Hands-on experience with AWS services, including:
    • S3
    • Step Functions
    • CloudWatch
  • Experience processing unstructured documents such as:
    • PDF
    • Word
    • Excel
    • PowerPoint
    • Email content
  • Experience building connectors and integrations with enterprise content repositories (e.g., SharePoint).
  • Experience with OCR and document extraction tools (AWS Textract or equivalent).
  • Experience designing and implementing data ingestion and transformation pipelines.
  • Familiarity with vector databases and Retrieval-Augmented Generation (RAG) concepts.
  • Experience with software development best practices:
    • Git
    • CI/CD
    • Automated testing

Nice to Have

  • Experience with Vector Databases.
  • Experience with RAG architectures and AI/LLM-based applications.
  • Experience with Azure cloud services.
  • Experience with Databricks.
  • Experience in Insurance, Banking, or other regulated industries.

Ready to Apply?

Take the next step in your career journey

Apply Now

About the job

Posted on
Sep 16, 2026
Job type
Full-time
Location
Bucharest, Bucharest, roOn-site

Explore more

Browse more jobs like this

Work arrangement

Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.