Faster chat, better deals — Get the App

AI Data Engineer

Indeed

Company

Job typeFull-time
Workplace typeOnsite
Experience levelNo experience limit
Education levelNo degree limit

Description

Job Summary: The Data Engineer at Credicorp Capital will ensure data availability, quality, and scalability for Generative AI solutions by building lightweight and reliable pipelines. Key Responsibilities: 1. Design, build, and maintain data ingestion and transformation pipelines. 2. Model AI-oriented data and prepare datasets for RAG/embeddings. 3. Orchestrate workflows and expose feeds/endpoints for AI consumption. **Credicorp Capital** invites you to Turn Challenges into Opportunities and join us as our next **AI** **Data Engineer** for the **Gen AI Innovation** team in **Lima, Peru.** ***Mission:*** The Data Engineer ensures data availability, quality, and scalability for Generative AI solutions through the construction of lightweight and reliable pipelines. Their purpose is to prepare, transform, and expose data enabling optimal operation of agents, copilots, and automations, in coordination with AI and business teams. ***Responsibilities:*** * Design, build, and maintain ingestion, transformation, and orchestration pipelines. * Model AI-oriented data and prepare datasets for RAG/embeddings. * Apply quality controls and traceability. * Maintain CI/CD for data pipelines and versioned documentation. * Orchestrate workflows and expose feeds/endpoints for consumption by agents/copilots and analytical products. ***Requirements:*** * University degree in Systems Engineering, Computer Science, Data Science, or related fields. * 3–5 years of experience in data engineering or cloud-based data integration. * Experience building pipelines and data models for analytics or AI (including dataset preparation for RAG/LLMs). * Participation in on-premises-to-cloud migrations and data workflow orchestration. * Technical English proficiency (Intermediate level preferred). ***Software and Technologies:*** * Azure Data Factory, Databricks, PySpark/Spark SQL, Delta Lake, Azure Synapse/Storage. * Advanced SQL (SQL Server, PostgreSQL) and Python (intermediate–advanced). * Version control and CI/CD: Git, GitHub/Bitbucket, GitHub Actions / Azure DevOps / Jenkins. * Integration: REST/JSON APIs; consumption from internal and external sources. * Airflow for additional orchestration (preferred); Power BI for supporting visualizations (preferred). ***Other Knowledge Areas:*** * Concepts of data lineage, quality, and governance (for alignment only). * Design of AI-oriented data models (tabular, semi-structured) and technical implementation using optimized formats (Parquet/Delta), partitioning, schemas, and metadata. * Batch and streaming over Delta Lake; audit/versioning patterns. * Security and privacy (PII: basic anonymization/pseudonymization; access controls). * Data integration from APIs, databases, and external sources. * Performance and cost optimization on cloud platforms. ***Preferred Knowledge:*** * General exposure to AWS or GCP (e.g., reading equivalent services, simple deployments, or foundational knowledge). * Basic understanding of LLMOps for data preparation and inference enablement. * Familiarity with lightweight orchestration (GitHub Actions, Data Factory, Airflow). * Certifications: AZ-900; DP-203, Databricks Data Engineer Associate, DP-900, PL-300. ***This posting is open to persons with disabilities.***

Some content was automatically translated

Posted by

María García

Indeed · HR

Location

María García

Indeed · HR

Similar jobs

AI Data Engineer job by Indeed in 2026 | ok.com