<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>CCAO Data Team</title>
<link>https://ccao-data.github.io/blog/</link>
<atom:link href="https://ccao-data.github.io/blog/index.xml" rel="self" type="application/rss+xml"/>
<description>Cook County Assessor&#39;s Office Data Team homepage and blog</description>
<generator>quarto-1.10.19</generator>
<lastBuildDate>Tue, 06 Oct 2026 00:00:00 GMT</lastBuildDate>
<item>
  <title>How the CCAO Data team manages data infrastructure</title>
  <dc:creator>Jean Cochrane</dc:creator>
  <link>https://ccao-data.github.io/blog/posts/data-architecture/</link>
  <description><![CDATA[ 





<p>Over the past 8 years, the Data team at the <a href="https://www.cookcountyassessoril.gov/" target="_blank">Cook County Assessor’s Office</a> (CCAO) has built a data stack that allows us to do modern data science, analytics, and engineering in a public-sector office whose source-of-truth data systems were not designed to support our work. It took us years of mistakes to get our stack right, so we’re taking the time to document what we’ve learned in the hopes that we might help other public servants do similar work.</p>
<section id="about-our-team-small-cross-functional-technical-and-autonomous" class="level2">
<h2 class="anchored" data-anchor-id="about-our-team-small-cross-functional-technical-and-autonomous">About our team: Small, cross-functional, technical, and autonomous</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ccao-data.github.io/blog/posts/data-architecture/assets/ccao-data-team-photo-cropped.jpg" class="img-fluid figure-img"></p>
<figcaption>The CCAO Data team on a field trip to the National Public Housing Museum in Chicago</figcaption>
</figure>
</div>
<p>Before we dive into the technical details of our data stack, it’s important to know some background about our team.</p>
<p>The CCAO Data team is a relatively new team that formed shortly after the election of Assessor Fritz Kaegi in Fall 2018. The team has been small and cross-functional from its inception: We have ranged in size from two to seven people, and our team members include data scientists, engineers, and analytics experts.</p>
<p>The CCAO Data team has a few core mandates:</p>
<ul>
<li><strong>Data science</strong>: Deliver initial value estimates for all single-family homes, small apartment buildings, and condos in Cook County by training statistical models to predict sale prices.</li>
<li><strong>Data engineering</strong>: Maintain our <a href="https://datacatalog.cookcountyil.gov/stories/s/gzdr-q7c4" target="_blank">open data assets</a> and build internal data products to help subject-matter experts on other teams do their work more efficiently and with greater impact.</li>
<li><strong>Data analytics</strong>: Report on the accuracy and fairness of our office’s assessments compared to real-world sales.</li>
</ul>
<p>Beyond these core mandates, we often act as internal consultants, helping other teams figure out ways to use data to improve their work. Some examples of this type of work include the property tax simulation software <a href="https://github.com/ccao-data/ptaxsim/" target="_blank">PTAXSIM</a>; the model explainability app <a href="https://www.cookcountyassessoril.gov/home-value-report" target="_blank">HomeVal</a>; and Tableau dashboards like the <a href="https://www.cookcountyassessoril.gov/cook-county-housing-market-tracker" target="_blank">Cook County Housing Market Tracker</a>. CCAO leadership have granted our team a high degree of autonomy, which means we are often empowered to pursue the collaborations that we think will have the greatest impact on the goals of the office.</p>
<section id="our-teams-composition-and-mandate-influence-our-stack" class="level3">
<h3 class="anchored" data-anchor-id="our-teams-composition-and-mandate-influence-our-stack">Our team’s composition and mandate influence our stack</h3>
<p>The requirements for our data stack are downstream of our team’s composition and mandate:</p>
<ul>
<li><p>Because our team is <strong>small</strong>, our infrastructure can be lean, since an outage will only affect a limited number of users.</p></li>
<li><p>Because our team is <strong>cross-functional</strong>, our infrastructure needs to be easy to maintain with minimal effort by team members who have varying degrees of familiarity with DevOps and MLOps.</p></li>
<li><p>Because our team is <strong>technical</strong>, we want to be able to keep as much of our work as possible under version control; we want to automate the boring stuff, like continuous integration and deployment (CI/CD); and we want programmatic access to our data so that we can use it in scripts and apps.</p></li>
<li><p>Because our team is <strong>autonomous</strong>, we want as much control over the design and operation of our data systems as possible.</p></li>
<li><p>Because our core mandate revolves around <strong>data science, engineering, and analytics</strong>, we sometimes need access to larger compute environments in order to run batch jobs, and we want columnar data storage to make aggregate queries more performant.</p></li>
</ul>
<p>It’s also important to note that <strong>the universe of our data is medium-sized</strong>. The basic unit of our work is the <em>parcel</em>, which is a unit of taxable real property that refers to a home, a commercial space, or vacant land. Cook County is home to roughly 1.9 million parcels, and since most parcel attributes are recorded once per year, our database tables tend to have record counts in the tens-to-hundreds of millions. That’s enough data that it’s often difficult to analyze on a laptop, but it’s much smaller than the data that most large private-sector data teams maintain, particularly teams that maintain real-time or event-driven data.</p>
<p>If your team doesn’t look like ours, or if your data look different from ours, then your stack will probably be different, too. Regardless, if you see some similarities between your team and ours, then we hope you’ll get some value out of understanding what we’ve built.</p>
<p>On that note, let’s dive into the stack.</p>
</section>
</section>
<section id="our-data-stack-at-a-glance" class="level2">
<h2 class="anchored" data-anchor-id="our-data-stack-at-a-glance">Our data stack at a glance</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ccao-data.github.io/blog/posts/data-architecture/assets/dataflow-diagram.svg" class="img-fluid figure-img"></p>
<figcaption>Data-flow diagram showing our data stack</figcaption>
</figure>
</div>
<p>The core of our data stack centers around our AWS Athena data lake. We extract raw data from the CCAO’s source-of-truth vendor database, along with a number of other public and private data sources, and load it into the data lake for cleaning and transformation. The data lake offers us programmatic access to our data, allowing us to build complex analytics, data science workflows, and apps on top of our data using state-of-the-art tools at a very low cost.</p>
<p>Let’s step through each component of our stack, starting with the most important data source: iasWorld, our system of record.</p>
<section id="source-of-truth-vendor-database-iasworld" class="level3">
<h3 class="anchored" data-anchor-id="source-of-truth-vendor-database-iasworld">Source-of-truth vendor database: iasWorld</h3>
<p>The CCAO system of record is a database called <strong>iasWorld</strong>. A private contractor (Tyler Technologies) maintains this database along with an associated web app that allows CCAO staff to edit data in the system.</p>
<p>Beyond editing data, the iasWorld web app allows CCAO staff to perform core office functions like reviewing building permits and adjusting assessments based on appeals. While we often describe iasWorld as a database, since that’s mostly how the Data team interacts with it, iasWorld is actually a <strong>cloud Enterprise Resource Planning (ERP) system</strong> that encapsulates all of the essential business logic of the CCAO, from certifying assessments to adjudicating appeals. The iasWorld database and web app are also shared among the other Cook County offices that are part of the property tax system, including the Clerk, the Treasurer, and the Board of Review.</p>
<p>While iasWorld is the system of record for the Data team, there are a few reasons why the Data team does not use it directly for our day-to-day analytics and data science workloads:</p>
<ul>
<li>The iasWorld database is a row-wise, relational Oracle database, so <strong>aggregate queries are slow</strong>.</li>
<li>There is often heavy load on the iasWorld database during the workday, so <strong>our query volume could overload the system</strong>.</li>
<li>The iasWorld database has a basic query interface, but <strong>it does not support the programmatic SQL queries that we need</strong> for analytics and data science.</li>
<li>We prefer to work with views that clean iasWorld’s raw tables and join them to other data sources, but <strong>our team does not have permission to create views in iasWorld</strong>, let alone introduce external data into the system.</li>
</ul>
<p>These limitations of iasWorld are understandable, since the database wasn’t designed to support the types of data science and analytics workloads that our team runs every day. In order to enable those workloads, we need a read-only mirror that can handle a high volume of aggregate, programmatic queries, and we need it to support custom views that we can test and update as underlying data change. We also need to combine the data in our system of record with other public and private data sources in order to properly model parcel characteristics and report on the performance of our assessments.</p>
<p>To meet all of these requirements, we built ourselves a data lake.</p>
</section>
<section id="data-lake-aws-athena-s3-glue" class="level3">
<h3 class="anchored" data-anchor-id="data-lake-aws-athena-s3-glue">Data lake: AWS Athena (S3 + Glue)</h3>
<p>Our <a href="https://docs.aws.amazon.com/athena/latest/ug/what-is.html" target="_blank">AWS Athena</a> data lake is the foundation of the day-to-day work of the Data team. The data lake stores all of the raw, intermediate, and final tables and views that allow us to integrate public data sources with data from our system of record. It also allows us to query our data programmatically and at scale, which in turn enables us to deliver a degree of analysis and automation that would be otherwise impossible.</p>
<p>Athena is an Amazon Web Services (AWS) product that provides a serverless SQL query engine using <a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html" target="_blank">AWS S3</a> for file storage and <a href="https://docs.aws.amazon.com/glue/latest/dg/what-is-glue.html" target="_blank">AWS Glue</a> for schema management. To query some data in Athena, we store structured data files in S3, usually in the compressed Parquet file format, and then we configure Glue to read those files as tables. Athena then provides an interface for querying our Glue tables using SQL, and it also allows us to store common SQL transformations as views.</p>
<p>There are a number of advantages to using a serverless query engine like Athena for our data lake:</p>
<ul>
<li><strong>Low cost</strong>: Athena charges us per gigabyte of memory that it scans during query execution, and S3 charges per gigabyte of files stored per month, so we only pay for what we use. Costs per gigabyte are generally quite low for both services, which means we almost never have to worry about exceeding our budget.</li>
<li><strong>Fast aggregation</strong>: Athena is able to read from columnar data sources like Parquet, allowing it to execute aggregate queries with greater efficiency than a row-wise database. This is helpful when running analytics and data science workflows, which often rely on aggregate queries.</li>
<li><strong>Horizontal scaling</strong>: Athena queries are parallelizable, and AWS manages the underlying infrastructure, so the query engine’s compute resources scale automatically to support our query load. In practice, this means that we never need to worry about accidentally slowing down our database with expensive queries.</li>
<li><strong>No maintenance</strong>: Since Athena is serverless, there is no underlying infrastructure that we have to manage. This frees up our staff to work on more important problems, and it means that our team doesn’t need an experienced database administrator in order to do our work.</li>
<li><strong>Programmatic access</strong>: Like all AWS services, Athena exposes a web API, and there are many well-maintained integrations that make it easy to use the API in different contexts. We rely heavily on Athena’s <a href="https://github.com/pyathena-dev/PyAthena" target="_blank">Python</a> and <a href="https://github.com/DyfanJones/noctua/" target="_blank">R</a> integrations.</li>
</ul>
<p>Again, these advantages are downstream of our team’s composition and mandate. If our data lake had more than a few users, or if our data were larger, then we might end up spending so much on storage and query execution that we would benefit from moving our infrastructure in-house and hiring database admins to maintain it. Similarly, if our team were less technical, then programmatic access to our data lake would be less important.</p>
<p>Athena is the foundation for the data science and analytics work that we do every day, but it’s not useful without data. To get data into our data lake, we need to extract it from external data sources and load it into the data lake.</p>
</section>
<section id="extractionloading-el-r-scripts-and-spark" class="level3">
<h3 class="anchored" data-anchor-id="extractionloading-el-r-scripts-and-spark">Extraction/loading (“E/L”): R scripts and Spark</h3>
<p>There are two main ways that we extract and load data into our data lake:</p>
<ol type="1">
<li><strong>Spark ingest for frequent iasWorld extraction</strong>: Since we need to extract a medium-sized amount of data from iasWorld and load it into our data lake on a daily basis, we use a <a href="https://spark.apache.org/" target="_blank">Spark</a> cluster to speed up the extract/load process for iasWorld data. We maintain this Spark code in a dedicated repo called <a href="https://github.com/ccao-data/service-spark-iasworld/" target="_blank"><code>service-spark-iasworld</code></a> and we run it on an on-premises Ubuntu server using a <a href="https://docs.docker.com/compose/intro/compose-application-model/" target="_blank">Docker Compose service</a> that we schedule using cron.</li>
<li><strong>R scripts for infrequent external data extraction</strong>: In contrast to iasWorld, which other teams update with new data every day, most of the external data sources that we use are small and infrequently-updated. For example, we use Census data to derive neighborhood characteristics that we feed into our valuation models, and the Census only publishes data once per year, at resolutions that are less granular than parcels (we typically use tract-level Census data). Since these data are small, and since we ingest them infrequently, it’s more important that their extraction/loading processes are easy to maintain than that they are maximally efficient, so we define them as R scripts and we run them manually when needed. We store these R scripts in the <a href="https://github.com/ccao-data/data-architecture/tree/master/etl" target="_blank"><code>etl/</code> subdirectory of our <code>data-architecture</code> repo</a>.</li>
</ol>
<p>Regardless of which extraction/loading process we use for a particular data source, the result is always the same: A parquet file stored in S3. Once that raw data is in S3, it’s ready to be cleaned up and combined with other data sources.</p>
</section>
<section id="transformation-t-dbt" class="level3">
<h3 class="anchored" data-anchor-id="transformation-t-dbt">Transformation (“T”): dbt</h3>
<p>Like most data teams, we maintain a large number of SQL transformations that clean our raw data tables, join them together, and make them easier to query. To manage these transformations, we use a tool called <a href="https://docs.getdbt.com/docs/introduction?version=1.11" target="_blank">dbt</a>.</p>
<p>To motivate our use of dbt, let’s look at a simple example: our modeling views. Our <a href="https://github.com/ccao-data/model-res-avm" target="_blank">residential model</a> and our <a href="https://github.com/ccao-data/model-condo-avm" target="_blank">condo model</a> both need access to roughly 100 different parcel attributes (“features”) in order to accurately predict sale prices in the Cook County real estate market. Many features (like building size and age) are stored in iasWorld, but many others (like distance to parks and transit stops) come from external data sources; even the features that come from iasWorld may be stored in disparate tables. When we train our valuation models and use them to predict sale prices, we need to pull all of these data together, which requires a hefty SQL query.</p>
<p>In theory, we could write our modeling queries directly in our modeling code, but then we wouldn’t be able to reuse those queries anywhere else without copying them over entirely. We would prefer to have dedicated views in our data lake that encapsulate these queries so that we can test them and reuse them.</p>
<p>Athena allows us to save our modeling queries as views, and <strong>dbt allows us to create and edit those views automatically</strong>. Using dbt, we can write a SQL query for our residential model input view and save it to <a href="https://github.com/ccao-data/data-architecture/blob/7ced41c14c39e4452f97aadd3f98a204ca7928a8/dbt/models/model/model.vw_card_res_input.sql" target="_blank">a file in our dbt project</a>; then, we can define some attributes of that view in a <a href="https://github.com/ccao-data/data-architecture/blob/7ced41c14c39e4452f97aadd3f98a204ca7928a8/dbt/models/model/schema.yml#L749" target="_blank">YAML config file</a> and use the <a href="https://docs.getdbt.com/reference/commands/build" target="_blank"><code>dbt build</code> command</a> to create an Athena view based on the SQL query and its YAML attributes. With dbt, we don’t have to push files to S3 or create and run Glue crawlers, since dbt automates those steps for us. Instead, the view is entirely defined by regular files, and we can track those files in our version control system (GitHub) and set up an automated job to update the view whenever its associated files change.</p>
<p>Beyond managing transformations, dbt also has some convenient features for testing and documenting our data. We use dbt to define and run <a href="https://docs.getdbt.com/docs/build/data-tests?version=1.11#overview" target="_blank">data tests</a> that confirm basic assumptions about our data, and we also use it to autogenerate a <a href="https://ccao-data.github.io/data-architecture/" target="_blank">data catalog</a> based on the YAML config files that define our tables and views.</p>
<p>Another way of thinking about it is that <strong>dbt applies the principle of <a href="https://en.wikipedia.org/wiki/Infrastructure_as_code" target="_blank">infrastructure-as-code</a> to the data in our data lake</strong>. Using dbt, we can document the relations in our data lake using version-controlled queries and metadata files; we can automatically update those relations whenever their definitions change; and we can test and document our relations. This framework allows us to make a high volume of changes to a large number of transformations while having confidence that all of those changes will be reviewed, applied, and recorded correctly in a central source of truth (namely, GitHub).</p>
<p>Though the company that maintains dbt offers a paid service for this type of automated workflow, we use the free and open source version of dbt because we have the technical capacity to maintain our own automations. While this means that we don’t need to pay for dbt, it also means that we need an automation and orchestration layer in order to put it into production.</p>
</section>
<section id="automation-and-orchestration-github-actions-plus-vms" class="level3">
<h3 class="anchored" data-anchor-id="automation-and-orchestration-github-actions-plus-vms">Automation and orchestration: GitHub Actions, plus VMs</h3>
<p>The <a href="https://docs.github.com/en/actions/get-started/understand-github-actions" target="_blank">GitHub Actions</a> platform allows us to run dbt whenever we push changes to the metadata files that define the transformed tables and views in our data lake. More broadly, it provides a general-purpose remote compute environment that we use for a variety of small automated tasks.</p>
<p>The basic unit of work for GitHub Actions is the <em>workflow</em>, which is a YAML config file defining a sequence of scripts that run on particular triggers, called <em>events</em>. The most common events that trigger our workflows include:</p>
<ul>
<li><a href="https://docs.github.com/en/actions/reference/workflows-and-actions/events-that-trigger-workflows#schedule" target="_blank"><code>schedule</code></a>: Run the workflow at specific times based on a cron expression.</li>
<li><a href="https://docs.github.com/en/actions/reference/workflows-and-actions/events-that-trigger-workflows#push" target="_blank"><code>push</code></a> / <a href="https://docs.github.com/en/actions/reference/workflows-and-actions/events-that-trigger-workflows#pull_request" target="_blank"><code>pull_request</code></a>: Run the workflow when a git event (commit to a branch or update to a pull request) occurs in the repo where the workflow is defined.</li>
<li><a href="https://docs.github.com/en/actions/reference/workflows-and-actions/events-that-trigger-workflows#workflow_dispatch" target="_blank"><code>workflow_dispatch</code></a>: Run the workflow on demand, with an optional set of input variables that the workflow caller can specify.</li>
</ul>
<p>When an event occurs that a workflow is configured to listen for, GitHub will spin up a temporary compute environment to run the scripts defined in that workflow. In this way, GitHub Actions is both an <em>automation</em> layer and an <em>orchestration</em> layer: “<strong>Automation</strong>” in that it can run tasks in a remote compute environment without requiring us to run commands directly, and “<strong>orchestration</strong>” in that it can trigger sequences of tasks that should run at specific times, or when certain prerequisite tasks execute successfully.</p>
<section id="examples" class="level4">
<h4 class="anchored" data-anchor-id="examples">Examples</h4>
<p>We maintain a large number of workflows for a variety of automated jobs. Some examples include:</p>
<ul>
<li><strong>Rebuilding relations in our data lake using dbt when their definitions change</strong>. The <a href="https://github.com/ccao-data/data-architecture/blob/85a1a92839088cd39ebf61c4408b8e772ef52d98/.github/workflows/build_and_test_dbt.yaml" target="_blank"><code>build-and-test-dbt</code></a> workflow rebuilds all new and modified relations in our data lake. We configure it to run on every push to a pull request or to the main branch in our dbt repo. If we push a commit to a pull request, then the workflow will create relations in a temporary staging database, while pushes to the main branch will trigger changes in our production database. In this way, we never have to directly modify our data lake resources, and indeed, we configure our Athena permissions so that no one can modify production relations outside the context of this workflow.</li>
<li><strong>Running weekly data integrity tests</strong>. In addition to defining data transformations in dbt, we also use it to define a suite of data integrity tests that check some basic assumptions about our data, like uniqueness constraints and expected row counts. We schedule the <a href="https://github.com/ccao-data/data-architecture/blob/85a1a92839088cd39ebf61c4408b8e772ef52d98/.github/workflows/test_dbt_models.yaml" target="_blank"><code>test-dbt-models</code></a> workflow to run these tests once per week, and to alert us if any tests fail.</li>
<li><strong>Running our valuation models on demand</strong>. The <a href="https://github.com/ccao-data/model-res-avm/blob/6bd10a58b8a89d72fd6f95e68253ab059cd700bb/.github/workflows/build-and-run-model.yaml" target="_blank"><code>build-and-run-model</code></a> workflow allows us to train a model in a remote compute environment and use it to generate a fresh set of predicted sale prices. We’ll cover this workflow in more detail in the section Data science workloads: AWS Batch + EC2 below.</li>
</ul>
</section>
<section id="advantages" class="level4">
<h4 class="anchored" data-anchor-id="advantages">Advantages</h4>
<p>There are a number of advantages to our use of GitHub Actions as an automation/orchestration layer:</p>
<ul>
<li><strong>Generous free tier</strong>. GitHub provides unlimited workflow minutes for public repos and a generous monthly budget for private repos, even without a paid subscription. For our first few years on GitHub, we were able to limit our use of workflows in private repos such that we didn’t need to pay for our workflow use at all. We eventually decided to procure a paid GitHub subscription for other reasons, but the generous free tier helped us stand up our stack quickly.</li>
<li><strong>Job definitions are stored alongside the code that they run</strong>. The workflow definition that is stored as a YAML config file in GitHub is the authoritative definition of that job, and the workflow updates automatically whenever we push changes to its config file. This makes it easy to inspect all of the code that our automated jobs are running.</li>
<li><strong>No infrastructure management required</strong>. Since GitHub manages the compute environments that run our workflows, we don’t need to worry about maintaining servers to run our code, and we can treat each run of a workflow as if it were running on a dedicated machine.</li>
</ul>
</section>
<section id="limitations" class="level4">
<h4 class="anchored" data-anchor-id="limitations">Limitations</h4>
<p>GitHub Actions has its own limitations to consider, too:</p>
<ul>
<li><strong>Lack of centralization</strong>. By default, GitHub workflows can’t reach services that are only accessible to users inside the Cook County network, including iasWorld.<sup>1</sup> As a result, a handful of our jobs must run on separate platforms in order to access County services like iasWorld. These platforms include an Ubuntu server and a remote Windows VM, which require us to schedule jobs using cron or the Windows Task Scheduler. This lack of centralization can make it hard for new team members to understand where jobs are running.</li>
<li><strong>Simplistic orchestration</strong>. The <code>schedule</code> event is simplistic compared to the equivalent features of more sophisticated orchestrators like <a href="https://airflow.apache.org/docs/apache-airflow/stable/index.html" target="_blank">Airflow</a>. A workflow defined in one repo can’t depend on the successful execution of a workflow defined in a different repo, and scheduled workflows run based on availability of GitHub compute resources, which means that GitHub often delays scheduled workflows by up to three hours.</li>
<li><strong>Limited compute capacity on the free tier</strong>. The free tier for GitHub Actions includes a generous amount of compute resources for workflows that run in public repos (at the time of writing, 4 vCPUs and 16GB memory), but data science workloads often need even more compute. <a href="https://docs.github.com/en/enterprise-cloud@latest/actions/concepts/runners/larger-runners" target="_blank">Larger runners</a> are only available for GitHub users on paid plans.</li>
</ul>
<p>Ultimately, GitHub Actions serves us well as an automation platform and orchestrator for simple tasks. For more complex tasks like data science workloads, however, we need more powerful compute environments than what GitHub’s free hosted runners can offer us. For those use cases, we decided to build our own solution.</p>
</section>
</section>
<section id="data-science-workloads-aws-batch-ec2" class="level3">
<h3 class="anchored" data-anchor-id="data-science-workloads-aws-batch-ec2">Data science workloads: AWS Batch + EC2</h3>
<p>We use <a href="https://docs.aws.amazon.com/batch/latest/userguide/what-is-batch.html" target="_blank">AWS Batch</a> as an abstraction layer on top of <a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html" target="_blank">AWS EC2</a> in order to run heavy data science workloads like our residential and condo valuation models.</p>
<p>AWS EC2 is a service that provides on-demand rental of remote compute environments, priced per second of usage. AWS Batch is a service that builds on top of EC2, providing a convenient interface for running batch compute jobs on ephemeral EC2 machines. To define a job in AWS Batch, you provide a Docker image for the job and you specify the amount of resources (CPU and memory) that the job requires; then when you’re ready to trigger the job, Batch will provision an EC2 instance that meets your resource requirements, execute the job on that instance using the Docker image you provided, and then tear down the EC2 instance when the job completes (or fails).</p>
<p>In order to trigger Batch jobs on demand, we use a GitHub workflow called <a href="https://github.com/ccao-data/model-res-avm/blob/6bd10a58b8a89d72fd6f95e68253ab059cd700bb/.github/workflows/build-and-run-model.yaml" target="_blank"><code>build-and-run-model</code></a> that listens for a <code>workflow_dispatch</code> event. The workflow creates or updates a Batch job configuration based on a set of input variables that the workflow caller specifies, then it starts a new job, printing a link to job logs that allow us to monitor execution. Using this workflow, we can trigger an unlimited number of model runs with the click of a button on a dropdown menu in the GitHub Actions interface.</p>
<p>Batch and EC2 are good fits for our data science workloads because they give us <strong>maximum flexibility with minimal configuration</strong>: We can use any of the <a href="https://aws.amazon.com/ec2/instance-explorer/" target="_blank">hundreds of EC2 instance types</a> as the basis for our compute environments; we can configure those environments with all necessary dependencies using a Dockerfile; and we can let Batch take care of setting up and tearing down the compute environment whenever we run a job.</p>
<p>Batch also allows us to easily hook our jobs into other AWS services, which comes in handy for logging and alerting.</p>
</section>
<section id="logging-and-alerting-cloudwatch" class="level3">
<h3 class="anchored" data-anchor-id="logging-and-alerting-cloudwatch">Logging and alerting: CloudWatch</h3>
<p>We use <a href="https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/WhatIsCloudWatch.html" target="_blank">AWS CloudWatch</a> as a remote log sink for all of our jobs and services. CloudWatch has a huge number of features, but we mostly use it for storing logs in a central location where they are easy to query programmatically.</p>
<p>These logs are useful for monitoring and debugging services, but they are also useful as a source of information about the status of our scheduled jobs. Since our orchestration layer is so simple (as we discussed in a previous section), we don’t have centralized job status reporting; instead, we maintain a repo called <a href="https://github.com/ccao-data/service-alerts/" target="_blank"><code>service-alerts</code></a> with a scheduled workflow that runs every three hours and queries the CloudWatch logs for our services to make sure they produced the expected logs without running into errors. This system requires that we maintain custom code in order to run healthchecks against our services, which could be a long-term maintenance risk, but the CloudWatch API mitigates some of that risk by making the code easy to write and maintain.</p>
</section>
<section id="infrastructure-management-terraform" class="level3">
<h3 class="anchored" data-anchor-id="infrastructure-management-terraform">Infrastructure management: Terraform</h3>
<p>Since our stack relies on many AWS resources, we use <a href="https://developer.hashicorp.com/terraform/intro" target="_blank">Terraform</a> to save the state of our AWS infrastructure in version-controlled config files. To do this, we maintain a private repo in which we store Terraform config files containing definitions for all of our resources, excluding IAM user accounts and their attached policies and groups.</p>
<p>While it’s common to use Terraform to automate the provision of AWS resources, we mostly avoid that approach, since it would require giving automated jobs permission to create and destroy many types of resources in our AWS account. This type of automation would increase the risk that the job might leak powerful credentials or accidentally incur downtime by deleting resources. Instead, we make most of our changes directly in the AWS console, and then we update our Terraform config after the fact to make sure it is up-to-date with the changes.<sup>2</sup></p>
</section>
<section id="analytics-workloads-tableau" class="level3">
<h3 class="anchored" data-anchor-id="analytics-workloads-tableau">Analytics workloads: Tableau</h3>
<p><a href="https://www.tableau.com/" target="_blank">Tableau</a> is our preferred data analytics platform. We connect Tableau to our Athena data lake in order to pull data into dashboards, and we use those dashboards as the basis for our internal and external reporting.</p>
<p>Some examples of public-facing Tableau dashboards that pull from the data lake include the <a href="https://www.cookcountyassessoril.gov/cook-county-housing-market-tracker" target="_blank">Cook County Housing Market Tracker</a> and our <a href="https://www.cookcountyassessoril.gov/historical-analysis-property-tax-spikes-2021-2023" target="_blank">Historical Analysis of Property Tax Spikes</a>. For a complete list of all of our public Tableau dashboards, see the <a href="https://www.cookcountyassessoril.gov/dashboard" target="_blank">Data Dashboards page on the CCAO website</a>.</p>
<p>Tableau is a good fit for interactive data products, but it’s more limited in its ability to deliver raw data. We sometimes use it to deliver raw data to internal stakeholders, where we have some ability to train other teams to use it correctly, but we don’t rely on it for delivering raw data to the public. In order to give the public direct access to as much of our raw data as possible, we need a dedicated open data portal.</p>
</section>
<section id="open-data-tyler-data-insights-fka-socrata" class="level3">
<h3 class="anchored" data-anchor-id="open-data-tyler-data-insights-fka-socrata">Open data: Tyler Data &amp; Insights (FKA Socrata)</h3>
<p>We use the Tyler Data &amp; Insights platform, formerly known as Socrata, to host our open data portal.</p>
<p>In order to push data to the portal, we schedule a GitHub workflow called <a href="https://github.com/ccao-data/data-architecture/blob/85a1a92839088cd39ebf61c4408b8e772ef52d98/.github/workflows/socrata_upload.yaml" target="_blank"><code>upload-open-data-assets</code></a> to run twice a month. The workflow runs a Python script that extracts metadata about our open data tables from our dbt config, queries the relevant tables in our data lake, and uses the <a href="https://dev.socrata.com/publishers/soda-producer/soda-producer-basics" target="_blank">SODA Producer API</a> to push new data to the portal. To make sure the upload works properly each month, we schedule a separate workflow called <a href="https://github.com/ccao-data/data-architecture/blob/master/.github/workflows/test_open_data_assets.yaml" target="_blank"><code>test-open-data-assets</code></a> to compare data in our data lake to data in the portal, ensuring that they match. This workflow sends us an alert if the assets in our data portal don’t have similar row counts to the underlying tables in our data lake, which can indicate that the upload failed to complete.</p>
<p>The open data portal is the outermost layer of our data stack, pulling data from our data lake and delivering it directly to the public. See <a href="https://datacatalog.cookcountyil.gov/stories/s/gzdr-q7c4" target="_blank">this guide</a> for an overview of all of the assets that we publish on our open data portal.</p>
</section>
</section>
<section id="benefits-of-our-stack-low-cost-high-velocity" class="level2">
<h2 class="anchored" data-anchor-id="benefits-of-our-stack-low-cost-high-velocity">Benefits of our stack: Low cost, high velocity</h2>
<p>Aside from a few proprietary services, our stack leans toward tools that are <strong>free and open source</strong>. The tools that we do pay for, like AWS and GitHub Actions, are often built on top of open-source standards and programming languages. This keeps our costs extremely low, since we pay for only a small set of proprietary tools, and it gives us the option of porting our code to run on other platforms if we wanted to switch to a cheaper competitor.</p>
<p>We also directly control the majority of our infrastructure, meaning we rarely need to ask for help from any internal teams or outside contractors in order to make changes to our stack. This helps us move fast in a bureaucratic environment where other teams are often busy with their own priorities.</p>
<p>As we mentioned above, this stack works well for us, but it only works because we have the technical capacity to maintain it. Many government teams choose to buy off-the-shelf services where we have chosen to build them internally, and that may be the right decision for their team.</p>
</section>
<section id="limitations-of-our-stack" class="level2">
<h2 class="anchored" data-anchor-id="limitations-of-our-stack">Limitations of our stack</h2>
<p>While our stack delivers a lot of benefit for us, it isn’t perfect, and we continue to make improvements to it. There are a few major structural weaknesses of the stack that we hope to address in the future.</p>
<section id="our-orchestration-layer-will-not-scale" class="level3">
<h3 class="anchored" data-anchor-id="our-orchestration-layer-will-not-scale">Our orchestration layer will not scale</h3>
<p>There’s a big problem with our orchestrator: Namely, that we don’t really have one. To see why this is a problem, let’s think about the jobs we run.</p>
<p>Most of our scheduled jobs depend on the daily Spark job that ingests data from iasWorld, our source-of-truth database. Currently, we configure all of our automated jobs to run on independent schedules, and we assume that upstream jobs like the Spark ingest will complete within a certain window of time.</p>
<p>This system is brittle, and it breaks easily. If upstream jobs fail, or if they run slower than expected, then the schedules for any downstream jobs will no longer be valid, causing our data products to display stale data. What’s worse, this type of failure can happen silently, since downstream jobs can often run successfully using old data.</p>
<p>For most of the history of the CCAO Data team, it made sense to stick with a simple-but-brittle orchestration layer, since our team maintained very few scheduled jobs that depended on the state of other jobs. Orchestration layers are notoriously complex systems that can be difficult to manage, and we preferred to save our capacity for higher-priority tasks.</p>
<p>However, as time has passed and our team has built trust with other departments, we find ourselves maintaining more and more automated data products that other teams rely on. As a result, we are starting to experience the costs of our lack of orchestration in the form of decreased reliability of those data products. Eventually, we will likely need to accept the additional maintenance burden of a sophisticated orchestrator like Airflow in order to continue expanding our suite of data products.</p>
</section>
<section id="automated-jobs-are-scattered-across-compute-environments" class="level3">
<h3 class="anchored" data-anchor-id="automated-jobs-are-scattered-across-compute-environments">Automated jobs are scattered across compute environments</h3>
<p>As we mentioned above in the section on automation and orchestration, we run automated jobs in a few different compute environments. GitHub Actions is our preferred compute environment, but we have to run some jobs on our Windows VMs and our Ubuntu server in order to access private data. We also run data science jobs on AWS Batch due to our historical reliance on the free tier of GitHub Actions, which didn’t provide us access to larger runners.</p>
<p>In the short term, these scattered compute environments don’t pose a major problem, though it’s often annoying to deploy code to our Windows VMs and our Ubuntu server given that those platforms don’t allow for automated CI/CD. However, these scattered compute environments are a long-term maintenance risk: The more compute environments we have, the harder it will be for the team to keep track of all of the jobs we’re running.</p>
<p>Eventually, we would like to migrate all of our jobs over to GitHub Actions so that they all run in the same environment. We could do this with a combination of self-hosted runners and larger runners. This migration would involve trading some staff time for the benefit of centralization, since self-hosted runners will require us to manage our own infrastructure rather than leave it up to GitHub.</p>
</section>
</section>
<section id="you-too-can-procure-the-tools-for-this-stack" class="level2">
<h2 class="anchored" data-anchor-id="you-too-can-procure-the-tools-for-this-stack">You too can procure the tools for this stack</h2>
<p>While we built most of our stack from the ground up in AWS, we had to go through procurement in order to get access to our AWS environment in the first place. For some government teams, procurement might be the hardest part of copying our stack.</p>
<p>In order to procure AWS, we needed to coordinate with other departments both inside and outside our organization. We worked with our Budget and IT departments to prep the required materials, and we were subject to the rules set by the Cook County Office of the Chief Procurement Officer.</p>
<p>In Cook County, some <a href="https://www.cookcountyil.gov/service/procurement-process" target="_blank">procurements</a> are competitive, involving time-consuming bids and requests for proposals. Others are non-competitive, which can be faster because there is insufficient vendor competition, an emergency, or a small cost. Thankfully, we were able to work with our executive leadership to make the case that our AWS account should be subject to non-competitive procurement, since there was only one vendor who was familiar enough with the CCAO and its processes to allow us to move as quickly as we needed.</p>
<p>To help with the vendor Statement of Work for our project, we listed out the responsibilities of the Data team that would require a data lake. For each responsibility, we listed the inputs, outputs, data model constraints, and data model requirements associated with the tasks. We also documented information about the operational importance of each task, task frequency, and whether the task was a current or future responsibility of the Data Department.</p>
<p>Our final statement of work leveraged the above task descriptions to build a six-week project timeline. The Statement of Work was signed by all parties – the CCAO, the Office of the Chief Procurement Officer, and our vendor – on June 23, 2021. Development work began in early August, and concluded in September of 2021.</p>
<p>There were <strong>three important conditions</strong> that helped us procure our way to this stack:</p>
<ol type="1">
<li><strong>Motivated and trusting organization leadership</strong>. Our leadership team wanted us to move fast, and they trusted our expertise. As a result, they granted us discretion to lead the procurement process, and they helped us coordinate across teams to speed it up. Many government teams do not have this kind of buy-in from their executive leadership, so they are forced to spend time building trust and persuading leadership.</li>
<li><strong>In-house technical capacity</strong>. We knew that we could build most of the infrastructure we needed once we had access to a cloud environment. Because of this capacity, we were able to keep our initial ask small and focused on standing up a simple environment that we could build upon. Teams with less technical capacity will likely need to contract out the building of the resources we describe above, which could lead to a larger, slower procurement process.</li>
<li><strong>Pre-existing AWS use in the County</strong>. Other Cook County offices had already used AWS at the time we asked to procure our own account. If we had been seeking entirely new software, we may have needed to make a more extensive case for our work.</li>
</ol>
<p>If you are a government team trying to procure a stack like ours, don’t hesitate to reach out to us to discuss. You can contact us via email at Assessor.Data at cookcountyil dot gov, or you can reach our Chief Data Officer Nicole Jardine on <a href="https://www.linkedin.com/in/nicole-jardine-data" target="_blank">LinkedIn</a>.</p>


</section>


<div id="quarto-appendix" class="default"><section id="footnotes" class="footnotes footnotes-end-of-document"><h2 class="anchored quarto-appendix-heading">Footnotes</h2>

<ol>
<li id="fn1"><p>Note that we <em>could</em> use GitHub workflows to reach services inside the County network if we switched to <a href="https://docs.github.com/en/actions/concepts/runners/self-hosted-runners" target="_blank">self-hosted runners</a>. See Limitations: Automated jobs are scattered across compute environments for more details.↩︎</p></li>
<li id="fn2"><p>There are a handful of cases in which we <em>do</em> use Terraform to automatically provision AWS resources, most notably in the workflows that trigger Batch jobs to run our valuation models (<a href="https://github.com/ccao-data/actions/blob/f4c8e4b39575a9371883014426452b3c4acfdfda/.github/workflows/build-and-run-batch-job.yaml" target="_blank">see the reusable workflow definition for those jobs here</a>). In these cases, we can tightly scope the permissions for the IAM role that the workflow assumes, since it only needs to manage a few discrete types of AWS resources.↩︎</p></li>
</ol>
</section></div> ]]></description>
  <guid>https://ccao-data.github.io/blog/posts/data-architecture/</guid>
  <pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate>
  <media:content url="https://ccao-data.github.io/blog/posts/data-architecture/assets/dataflow-diagram.svg" medium="image" type="image/svg+xml"/>
</item>
</channel>
</rss>
