SITS HESA Data Infrastructure Rebuilt: From Six-Week Submission Crisis to Two-Day Scheduled Task
A mid-size UK university was spending six weeks every year firefighting HESA submission errors - errors that originated in SITS months earlier and went completely undetected until the January deadline. Registry staff were manually exporting CSVs, admissions teams were re-keying UCAS data by hand, and senior leaders had no live view of enrolment or student retention. Numlytics audited the full SITS HESA data infrastructure, fixed the data before building a single dashboard, and delivered working automated pipelines within six weeks.
The Challenge: A SITS HESA Data Infrastructure Built on Manual Workarounds
The university's data problems were not new - they had accumulated over years of under-investment in higher education data infrastructure. By the time Numlytics engaged, the institution was managing four compounding pain points simultaneously, each making the others worse.
- HESA errors discovered at submission time. The registry team had no automated HESA Data Futures validation running during the year. JACS/HECoS code errors, missing withdrawal dates, and referential integrity failures only surfaced in January when there was no time to investigate root causes, only to patch manually under pressure.
- SITS data never reaching Power BI. Analysts were manually exporting SITS data to CSV and uploading it to Excel models. Reports were typically five to ten days stale by the time stakeholders saw them. Finance, registry, and academic teams were working from different versions of the same data a classic broken SITS Power BI integration pattern.
- UCAS applications re-keyed into SITS by hand. Admissions staff handling over 14,000 applications per cycle were manually transferring UCAS application records into SITS. The process took weeks, introduced transcription errors, and created mismatched records that caused downstream enrolment problems a UCAS to SITS integration gap that had never been closed.
- No student retention visibility. The Head of Student Experience had no way to identify at-risk students until pastoral tutors flagged them typically three to four weeks after disengagement had begun. By that point, interventions had limited impact on withdrawal rates.
- No university data governance framework. Data ownership was undefined. There were no validation rules enforced at point of entry in SITS, no data dictionary, and no monitoring in place. Each team had developed its own workaround, making the data landscape increasingly fragmented over time.
The Numlytics Solution: Rebuilding the SITS HESA Data Infrastructure in Four Phases
Numlytics approached the engagement in four sequential phases, each building on a validated foundation before introducing the next layer. The principle throughout was straightforward: fix the data before building the dashboards, not the other way around.
-
01Full Data Audit and SITS Source Mapping
Numlytics began with a structured two-week audit of the university's data landscape mapping every source system (SITS:Vision, UCAS, Agresso finance, Moodle VLE, and HR), assessing SITS field configuration against current HESA Data Futures coding manual requirements, and categorising all identified data quality issues by severity and HESA risk. The audit produced a prioritised delivery roadmap agreed with IT, registry, and the Deputy Vice-Chancellor's office. No pipeline was written until the source landscape was fully understood a principle central to any successful SITS HESA data infrastructure rebuild.
-
02SITS to Power BI ETL Pipeline via Azure Data Factory
Numlytics built a secure, scheduled Azure Data Factory pipeline connecting SITS:Vision directly to a centralised SQL data warehouse, with HESA field mapping validated at source. Student records, module enrolments, assessment marks, and progression data now refresh automatically overnight. The manual CSV export process was retired entirely. Registry, finance, and academic teams access the same validated data through Power BI dashboards - with a timestamp showing exactly when each dataset was last refreshed. This was the SITS Power BI integration the institution had needed for years.
-
03UCAS to SITS Integration Pipeline
Numlytics designed and built a full UCAS to SITS integration pipeline using the UCAS API and standard application file formats. Application records now map directly to SITS programme and module fields, with automated matching logic handling the applicant-to-student transition. Where UCAS and SITS details conflict a common occurrence with name variants and address changes - the pipeline surfaces exceptions for human review rather than silently creating duplicates. Manual re-keying of all 14,000+ annual applications was eliminated entirely.
-
04HESA Data Futures Automated Validation Dashboard
Numlytics built a nightly HESA Data Futures validation pipeline running automated quality checks against the current HESA coding manual flagging errors at the point they occur in SITS rather than accumulating them for eleven months. A live Power BI quality dashboard gives the registry team a daily view of HESA risk: errors by type, count, severity, and the SITS record responsible. The January HESA submission became a scheduled task, not a crisis response.
-
05Student Retention Analytics and Governance Framework
Numlytics built a student data lifecycle dashboard pulling attendance (from the timetabling system), VLE engagement (from Moodle), and assessment data into a single Power BI model giving the student experience team a live early-warning view updated daily. In parallel, Numlytics designed and implemented a university data governance framework: named data owners for each SITS entity, validation rules enforced at point of entry, a HESA-aligned data dictionary, and a monitoring dashboard tracking data health metrics over time.
The Results: What the SITS HESA Data Infrastructure Rebuild Delivered
- Manual CSV exports - data 5–10 days stale
- HESA errors found in January each year
- 14,000+ UCAS applications re-keyed by hand
- 6 weeks of HESA submission prep every cycle
- No at-risk student visibility
- No data ownership or governance framework
- 40 staff-hours wasted per week
- Automated nightly refresh - data always current
- HESA errors flagged nightly, year-round
- UCAS-to-SITS fully automated, zero re-keying
- 2-day submission prep a scheduled task
- Daily at-risk student dashboard live
- Named data owners and governance framework
- 40 staff-hours per week reclaimed
Before Numlytics, our HESA submission was a six-week ordeal every year. We had errors we couldn't trace, data we couldn't trust, and a registry team running on coffee and adrenaline. The pipeline Numlytics built changed everything our HESA data quality dashboard now catches issues the moment they appear in SITS, not in January.
We'd been told fixing our SITS data management would take 12 months and a full system replacement. Numlytics audited our actual data flows and had the first pipelines live in under six weeks. They understood SITS, they understood HESA Data Futures, and they understood what we actually needed, which is rare in this sector.
Why the University Chose Numlytics for Its SITS HESA Data Infrastructure
Before engaging Numlytics, the university evaluated three options: a general IT consultancy, expanding its in-house data team, and a specialist higher education data analytics partner. The decision came down to sector depth, delivery speed, and the ability to show working results within weeks - not months.
| Criteria | Numlytics | General IT Consultancy | In-House Only |
|---|---|---|---|
| SITS:Vision hands-on experience | ✓ Active UK university client — no ramp-up | ✗ Learned SITS on this engagement | Partial — strong registry knowledge, weaker engineering |
| HESA Data Futures & coding manual knowledge | ✓ Current, submission-cycle tested | ✗ Required dedicated HESA onboarding time | Registry strong; data engineering application weaker |
| UCAS to SITS integration pattern | ✓ Pre-built, tested pipeline architecture | ✗ Custom build from scratch — 4–5 months estimated | ✗ Not resourced — deprioritised annually |
| Time to first working pipeline | ✓ 6 weeks | 6–12 months quoted | Indefinite — no dedicated capacity |
| University data governance framework | ✓ Delivered both policy and tooling | Policy documentation only — no delivery layer | Policy drafted; never implemented at system level |
| OfS, TEF, NSS, POLAR4 familiarity | ✓ Fluent from day one | Terminology learned during engagement | ✓ Strong internally |
| Ongoing HESA submission support | ✓ Included — annual cycle partner | Charged separately per engagement | In-house team only — high pressure period |
| Engagement cost vs outcomes | ✓ Fixed scope — delivered in defined timeline | T&M — extended significantly beyond estimate | Headcount cost without specialist delivery |
What Numlytics Delivered - Complete SITS HESA Data Infrastructure
Every component below was designed, built, tested, and handed over with full documentation and training for the university's internal team. This is what a complete SITS HESA data infrastructure engagement looks like in practice.
- Azure Data Factory pipeline: SITS → SQL warehouse → Power BI - nightly refresh, automated scheduling, monitoring alerts on pipeline failure, and full audit trail of every data transformation.
- UCAS-to-SITS automated integration pipeline - full application lifecycle from UCAS submission to SITS enrolment, with exception handling and manual review queue for partial matches.
- HESA nightly validation pipeline - automated checks run against current HESA Data Futures coding manual schema, all error types categorised by HESA entity and severity, live Power BI quality dashboard accessible to registry team 24/7.
- Student retention Power BI dashboard - attendance, VLE engagement, and assessment signals combined into a single at-risk model, refreshed daily, accessible to personal tutors and student experience leads.
- Enrolment and finance dashboards - live enrolment counts by programme, mode, level, and campus; finance position updated overnight from Agresso integration.
- University data governance framework - named data owners per SITS entity, validation rules enforced in SITS configuration, HESA-aligned data dictionary, monthly data health scorecard, and quarterly governance review process.
- Documentation and handover - full technical documentation of all pipelines, Power BI semantic model documentation, data dictionary in SharePoint, and two-day training session for IT and registry staff.