Data Engineer
Role Summary
The owns data clarity and measurement effectiveness across AI and automation initiatives. AI and automation can only deliver reliable, measurable value when the underlying data is understandable, accurate, accessible, and consistently measured.
The role is responsible for source data semantics, extraction, transformation, reusable pipelines, reliable data models, and the technical definition of how approved business outcomes are calculated, sourced, baselined, and reproduced.
This is a hands-on data engineering role rather than a data strategy or architecture position. The role focuses on building reliable data solutions and does not determine which business outcomes or priorities the organization should pursue.
The position is an individual contributor role with no direct reports.
Required Competencies
- Hands-on data engineering experience rather than primarily data architecture or strategy experience.
- Ability to take a business metric from its initial definition through to a reproducible and defensible result.
- Comfortable working with messy, incomplete, inconsistent, or semantically ambiguous source data.
- Willingness to identify and challenge unclear measurement definitions rather than making assumptions without clarification.
- Ability to understand unfamiliar APIs or data sources and correctly implement authentication, incremental loading, and failure handling.
- Strong written English skills, with the ability to explain data limitations, findings, and remediation recommendations to both technical and non-technical audiences.
Environment and Tools
The makes information from systems of record usable for automation, AI agents, reporting, and business decision-making.
The role works primarily with APIs, data pipelines, storage layers, and data surfaces rather than being limited to a single platform. Client projects may introduce unfamiliar systems, so adaptability and strong data engineering fundamentals are highly valued.
Required and Transferable Skills
- Strong SQL skills plus at least one scripting or programming language such as Python or PowerShell.
- Experience building data pipelines against APIs and systems of record, including authentication, incremental loading, and failure handling.
- Strong understanding of data modeling, data semantics, and data definitions.
- Ability to create reproducible and defensible data calculations and metrics.
- Familiarity with source control, technical documentation, and production support practices.
Systems and Technologies
Experience with the following is beneficial but not mandatory, as training may be provided:
- Microsoft Fabric
- Azure Data Factory or comparable data pipeline platforms
- Microsoft 365, SharePoint, and Microsoft Graph
- Microsoft Purview, sensitivity labels, retention, and data loss prevention
- ConnectWise Manage and BMS as source systems accessed through APIs
- Power BI and comparable visualization platforms
- RMM and other operational data sources
The role will support modern AI and agent-based platforms. The is not expected to build AI agents but should understand the data requirements of AI systems and design data solutions that support reliable retrieval, automation, and grounding.
Key Responsibilities
- Build and maintain ETL and ELT pipelines that reliably move data between systems of record and structured storage using Microsoft Fabric, Azure data services, APIs, and related technologies.
- Transform relational, semi-structured, and unstructured data into reliable formats that can be consumed by automation, AI agents, retrieval systems, and reporting platforms.
- Create and document reusable data definitions, models, and transformations to ensure common business concepts are calculated consistently.
- Translate approved business metrics into reproducible technical calculations with clearly documented sources, filters, assumptions, and limitations.
- Establish reliable baselines before significant workflow or process changes so performance can be measured consistently afterward.
- Assess and improve Microsoft 365 and SharePoint environments where permission issues, oversharing, duplicate or outdated content, or poor metadata quality affect data readiness and AI adoption.
- Implement appropriate data-handling and governance controls using tools such as Microsoft Purview, sensitivity labels, retention policies, and data loss prevention.
- Build and maintain reporting and visualization solutions that make operational and business measurements visible and actionable.
- Develop data dictionaries, pipeline documentation, measurement definitions, and technical support documentation.
- Monitor and support production data pipelines during business hours, troubleshooting failures and addressing data quality or reliability issues.
Qualifications
- Demonstrated experience building production data pipelines, integrations, or transformations using SQL and at least one scripting or programming language such as Python or PowerShell.
- Strong fundamentals in relational databases, APIs, semi-structured data, and unstructured content.
- Ability to identify, investigate, and address data quality and semantic issues.
- Working knowledge of cloud-based data services such as Microsoft Fabric, Azure Data Factory, or comparable platforms.
- Ability to build a metric end to end by identifying the source, defining the calculation, validating the result, and making the process reproducible.
- Working familiarity with Microsoft 365 and SharePoint data, permissions, and content structures.
- Practical reporting and visualization experience, including Power BI or comparable tools.
- Comfortable with source control, task tracking, documentation, testing, and production support practices.
Nice to Have
- Advanced experience with Microsoft 365, SharePoint, or Microsoft Graph
- Experience supporting Copilot or AI-readiness initiatives
- Microsoft Purview and data governance experience
- Experience working with retrieval systems or grounding AI solutions using business data
- Experience with data quality and content remediation projects
- Managed service provider or professional services experience
- Experience working with AI and automation projects
What Success Looks Like in the First 90 Days
- Core operational data baselines are documented and reproducible, with definitions, sources, filters, and assumptions clearly identified.
- At least one automation or business workflow has a defensible before-and-after measurement supported by the data engineering process.
- High-priority data or content readiness issues have been assessed, documented, and accompanied by a measurable remediation plan.
- At least one reusable data, measurement, or data-readiness pattern has been developed and documented for future projects.
- Production data pipelines are reliable, documented, and supportable by other members of the team.