AI Data Enablement Engineer
hace 22 horas
Terrassa
Our Client's Digital Finance IT is building an AI-enablement layer on top of our enterprise data platform to enable business users across Finance to interact with governed data in natural language. We're hiring a Data Enablement Engineer to design, build, and operate the trusted datasets, semantic models, and embedded AI experiences that make this possible. This is a data platform engineering role, not a data science or model-building role. You will spend your time engineering the data foundation that makes AI reliable — semantic layers, governed data products, and embedded natural-language analytics — not training models. • Design and build AI-ready data products on Databricks — trusted datasets with well-defined business semantics, KPIs, hierarchies, and business glossary alignment, • Implement semantic layers and governed datasets that support both traditional BI consumption and natural-language querying by business users, • Deploy and operate Databricks Genie spaces with Unity Catalog, tuning them for accuracy, adoption, and business relevance, • Build RAG pipelines and conversational analytics applications grounded in governed enterprise data — including Streamlit or Databricks Apps that let business users query data without writing SQL, • Engineer robust ETL/ELT pipelines (dbt, Airflow, Snowpark, PySpark) that produce and maintain the trusted data these AI experiences depend on, • Implement data governance — RBAC, row/column-level security, masking, lineage, auditability, catalog and metadata management — in a regulated pharma environment, • Optimize cost and performance on both the data platform side (warehouse sizing, cluster tuning, query optimization) and the AI side (token usage, caching, model routing), • Partner with Finance business stakeholders to translate domain requirements into semantic models and governed data products they can trust, • 5+ years hands-on data engineering on cloud data platforms — Databricks demonstrated in real project delivery, not skill-list-only, • Direct hands-on experience with Databricks Genie — you have built, configured, and tuned these in production or advanced pilots, with specific reference to the flavors used (Cortex Analyst / Search / Agents / LLM Functions, or Genie spaces with semantic models), • Semantic layer / trusted data product delivery — you have built governed datasets that business users can rely on, with KPI definitions, hierarchies, and business glossary alignment, • dbt, PySpark, Snowpark, SQL, Python — strong across the modern data stack, • Orchestration with Airflow, Databricks Workflows, or equivalent, • Data governance in regulated environments — RBAC, RLS, masking, lineage, auditability, • Experience integrating structured and unstructured data (PDFs, SharePoint/Teams content, enterprise knowledge sources) into AI-enablement workflows, • Pharma, life sciences, or regulated financial services domain experience, • Veeva CRM, IQVIA, SAP, or clinical data source integration, • Streamlit or Databricks Apps for business-facing analytics, • Databricks Data Engineer Professional certification, • LangChain, LlamaIndex, or equivalent RAG frameworks, • Cost optimization on both compute (warehouse/cluster) and LLM (tokens/caching/routing) dimensions, • Data Scientists — this role is not model training, fine-tuning, LoRA/RLHF, or ML research, • Pure Data Engineers who list Cortex or Genie as a skill but haven't shipped it in production, • AI/GenAI engineers whose center of gravity is LangChain agents or RAG-over-documents, without a strong governed data platform foundation, • Computer vision, NLP model builders, or multi-agent orchestration specialists — wrong shape for this role #J-18808-Ljbffr