You are viewing a preview of this job. Log in or register to view more details about this job.

Computer Vision & Layout Intern

Computer Vision & Layout Intern (Multimodal Translation)

The AI Publisher Platform implements a highly scalable, decoupled Layered Microservices Architecture designed to unify autonomous asset generation, complex cross-lingual processing, and state-heavy real-time human-in-the-loop workflows. This structural decoupling ensures that data-heavy generative pipelines operate independently of the client-facing interfaces, guaranteeing platform high-availability and extreme sub-second UI responsiveness.

Role Objective

Develop the automated image manipulation and structural canvas translation engines capable of programmatically detecting, isolating, and inpainting text assets embedded inside book covers and internal illustrations.

Key Responsibilities

  • Design robust image preprocessing and layout processing pipelines utilizing OpenCV and advanced cloud-based OCR frameworks.
  • Integrate Generative AI image-to-image mechanisms (Stable Diffusion Inpainting, ControlNet) to programmatically erase text assets.
  • Calculate relative bounding-box coordinate mappings to feed exact text overlay positions directly to client-facing UI layers.
  • Optimize asset asset manipulation tasks to run efficiently within asynchronous worker environments.

Required Technical Qualifications

  • Current enrollment in an MS or PhD program focusing on Computer Vision, Machine Learning, or Graphic Systems.
  • Proficiency in Python and deep familiarity with structural image manipulation tools (OpenCV, PIL, Scikit-Image).
  • Academic or professional experience modifying deep generative weights or interacting with diffusion models.