> ## Documentation Index
> Fetch the complete documentation index at: https://cloudanix.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# EMR Clusters Should Be In VPC

### More Info:

Ensure that your Amazon Elastic MapReduce (EMR) clusters are provisioned using the AWS EC2-VPC platform instead of EC2-Classic platform (outdated from 2013.12.04) for better flexibility and control over security, better traffic routing and availability.

Medium

### Address

Cost Optimization, Security

### Compliance Standards

PCIDSS, HIPAA

### Remediation

How to ensure EMR clusters are in VPC

#### Using AWS Console

1. Navigate to the Amazon EMR console.
2. Click on the "Create cluster" button to create a new cluster, or select an existing cluster from the cluster list.
3. In the "Create Cluster" or "Edit Cluster" page, locate the "General configuration" section.
   (In the Cloudanix Console, navigate to "Misconfig" page and look for Affected Assets for "MR Clusters Should Be In VPC" Policy.)
4. Under the "Network settings" section, ensure that the "Launch mode" is set to "VPC".
5. Select the desired VPC from the available options. If you haven't created a VPC yet, you can create one using the Amazon VPC service before configuring your EMR cluster.
6. Select the appropriate Subnet(s) within the selected VPC where you want the EMR cluster instances to be launched.
7. Optionally, you can configure additional network settings such as the VPC security groups and EMR managed scaling rules.
8. Review and configure the remaining settings for your EMR cluster, such as instance types, EC2 key pair, software configuration, etc.
9. Once you have completed the configuration, click on the "Create cluster" or "Save" button to launch or update the cluster.

### Triage and Remediation

<Tabs>
  <Tab title="Remediation">
    ### Remediation

    <AccordionGroup>
      <Accordion title="Using Console" defaultOpen="true">
        For Redshift, the equivalent control is “Redshift clusters must be launched in a VPC (not EC2-Classic) and associated only with approved VPCs/subnets.”\
        Redshift now launches only into VPC by default in most regions, but older clusters might still be outside or in an incorrect VPC/network setup. You can’t “move” an existing Redshift cluster to a different network; you must recreate it in the desired VPC.

        Below are console-based steps, assuming two scenarios:

        ***

        ## 1. Ensure new Redshift clusters are in a VPC

        ### A. Create or verify a VPC and subnets

        1. In the AWS Management Console, go to **VPC**.
        2. If you already have a compliant VPC, note:
           * VPC ID
           * Subnet IDs (in at least two AZs if you want multi-AZ style resilience)
           * Route tables, and whether they allow required connectivity.
        3. If you need a new VPC:
           1. Click **Your VPCs > Create VPC**.
           2. Choose **VPC only** or **VPC and more**, give it a name and CIDR (e.g., `10.0.0.0/16`).
           3. Create at least 2 **private subnets** for Redshift:
              * Go to **Subnets > Create subnet**, select your VPC and define CIDR blocks (e.g., `10.0.1.0/24`, `10.0.2.0/24`).
           4. Set up **NAT gateway / routing** if the cluster needs outbound internet access.
           5. Optionally create **VPC endpoints** (for S3 etc.) if you want private access.

        ### B. Create a Redshift subnet group

        1. Go to **Amazon Redshift** console.
        2. In the left panel, choose **Configurations > Subnet groups**.
        3. Click **Create cluster subnet group**.
        4. Provide:
           * **Name** and **Description**
           * **VPC**: select the VPC you prepared.
           * **Subnets**: add the private subnets for Redshift.
        5. Click **Create cluster subnet group**.

        ### C. Create or configure a security group for Redshift

        1. Go to **VPC** console → **Security groups**.
        2. **Create security group**:
           * Select the same **VPC**.
           * Add an **inbound rule** for Redshift:
             * Type: **Redshift** (TCP 5439) or a custom TCP rule for the cluster port.
             * Source: a specific CIDR or security group (e.g., application servers’ SG), not `0.0.0.0/0` unless justified.
        3. Save the security group ID for use when creating the cluster.

        ### D. Create the Redshift cluster in the VPC

        1. Go to **Amazon Redshift** console.
        2. Click **Create cluster**.
        3. Under **Cluster configuration**, set usual options (identifier, node type, etc.).
        4. Under **Network and security**:
           * **VPC**: select your target VPC.
           * **Subnet group**: select the subnet group you created.
           * **Publicly accessible**:
             * Usually **No** (private cluster); set **Yes** only if absolutely necessary.
           * **VPC security groups**: select the security group you configured.
        5. Finish the wizard and create the cluster.

        Result: all new clusters are properly in a VPC and use only the allowed network resources (VPC, subnets, SGs).

        ***

        ## 2. Remediate an existing Redshift cluster not in the correct VPC

        You cannot directly move a Redshift cluster from EC2-Classic or another VPC. You must snapshot and recreate.

        ### A. Take a snapshot of the existing cluster

        1. Go to **Amazon Redshift** console → **Clusters**.
        2. Select the non-compliant cluster.
        3. Choose **Actions > Create snapshot**.
        4. Enter a name and confirm.
        5. Wait until the snapshot status is **available**.

        ### B. Restore the snapshot into a VPC

        1. In **Redshift** console, go to **Snapshots**.
        2. Select the snapshot you just created.
        3. Click **Actions > Restore from snapshot** (or **Restore snapshot**).
        4. In the restore wizard:
           * **Cluster identifier**: new name (e.g., `mycluster-vpc`).
           * Under **Network and security**:
             * **VPC**: select the compliant VPC.
             * **Subnet group**: choose the subnet group you created.
             * **Publicly accessible**: typically **No**.
             * **VPC security group**: choose the secure SG.
           * Adjust port if needed (default 5439).
        5. Complete restoration and wait for the new cluster to become **available**.

        ### C. Update clients / applications to use the new cluster

        1. In the new cluster details, copy the **Endpoint**.
        2. Update:
           * Application configuration / connection strings.
           * ETL jobs, BI tools, Lambda functions, etc., to point to the new endpoint.
        3. Validate connectivity and data correctness.

        ### D. Decommission the old non-compliant cluster

        1. Once fully validated, go to **Clusters**.
        2. Select the old cluster.
        3. **Actions > Delete**:
           * Optionally keep a **final snapshot** if you want.
        4. Confirm deletion.

        ***

        By enforcing that all Redshift clusters are restored/created only in approved VPCs and subnet groups, you satisfy the “clusters must be in VPC” requirement for Redshift using the AWS console.
      </Accordion>

      <Accordion title="Using CLI">
        For Amazon Redshift, the equivalent of “EMR clusters should be in a VPC” is “Redshift clusters must be launched in a VPC (not EC2-Classic).”\
        You can’t move an existing non‑VPC Redshift cluster into a VPC; instead you snapshot it and restore into a VPC.

        Below are AWS CLI–based steps.

        ***

        ## 1. Identify Redshift clusters not in a VPC

        ```bash theme={null}
        aws redshift describe-clusters \
          --query "Clusters[?VpcId==null].[ClusterIdentifier,NodeType,ClusterStatus]" \
          --output table
        ```

        If this returns your cluster(s), they are not in a VPC and must be migrated.

        ***

        ## 2. Choose or create a VPC and subnets

        List VPCs:

        ```bash theme={null}
        aws ec2 describe-vpcs \
          --query "Vpcs[*].[VpcId,CidrBlock,IsDefault]" \
          --output table
        ```

        Pick a VPC ID (e.g., `vpc-0123456789abcdef0`).

        List subnets in that VPC:

        ```bash theme={null}
        VPC_ID=vpc-0123456789abcdef0

        aws ec2 describe-subnets \
          --filters "Name=vpc-id,Values=$VPC_ID" \
          --query "Subnets[*].[SubnetId,AvailabilityZone,CidrBlock]" \
          --output table
        ```

        Pick 2+ subnets in different AZs for production use.

        ***

        ## 3. Create a Redshift subnet group in that VPC

        ```bash theme={null}
        SUBNET_GROUP_NAME=my-redshift-subnet-group
        SUBNET_IDS="subnet-aaa,subnet-bbb"   # comma-separated

        aws redshift create-cluster-subnet-group \
          --cluster-subnet-group-name "$SUBNET_GROUP_NAME" \
          --description "Redshift subnet group in VPC $VPC_ID" \
          --subnet-ids $SUBNET_IDS
        ```

        ***

        ## 4. Create / choose security groups for Redshift

        Create a security group in the same VPC:

        ```bash theme={null}
        SG_NAME=my-redshift-sg

        aws ec2 create-security-group \
          --group-name "$SG_NAME" \
          --description "Redshift access" \
          --vpc-id "$VPC_ID"
        ```

        Open Redshift port 5439 only to needed sources (example: office IP):

        ```bash theme={null}
        SG_ID=$(aws ec2 describe-security-groups \
          --filters "Name=group-name,Values=$SG_NAME" "Name=vpc-id,Values=$VPC_ID" \
          --query "SecurityGroups[0].GroupId" \
          --output text)

        MY_IP=x.x.x.x/32

        aws ec2 authorize-security-group-ingress \
          --group-id "$SG_ID" \
          --ip-permissions IpProtocol=tcp,FromPort=5439,ToPort=5439,IpRanges="[{CidrIp=$MY_IP,Description=RedshiftAccess}]"
        ```

        ***

        ## 5. Snapshot the existing non‑VPC cluster

        ```bash theme={null}
        OLD_CLUSTER_ID=old-redshift-cluster
        SNAPSHOT_ID=${OLD_CLUSTER_ID}-pre-vpc-migration-$(date +%Y%m%d%H%M%S)

        aws redshift create-cluster-snapshot \
          --cluster-identifier "$OLD_CLUSTER_ID" \
          --snapshot-identifier "$SNAPSHOT_ID"
        ```

        Wait until snapshot is ready:

        ```bash theme={null}
        aws redshift describe-cluster-snapshots \
          --snapshot-identifier "$SNAPSHOT_ID" \
          --query "Snapshots[0].Status" \
          --output text
        # wait until status == available
        ```

        ***

        ## 6. Restore the cluster into the VPC

        Restore from snapshot specifying the subnet group and security group(s):

        ```bash theme={null}
        NEW_CLUSTER_ID=new-redshift-vpc-cluster

        aws redshift restore-from-cluster-snapshot \
          --cluster-identifier "$NEW_CLUSTER_ID" \
          --snapshot-identifier "$SNAPSHOT_ID" \
          --cluster-subnet-group-name "$SUBNET_GROUP_NAME" \
          --vpc-security-group-ids "$SG_ID"
        ```

        Optionally adjust public accessibility:

        ```bash theme={null}
        aws redshift modify-cluster \
          --cluster-identifier "$NEW_CLUSTER_ID" \
          --publicly-accessible false
        ```

        Wait for the cluster to become `available`:

        ```bash theme={null}
        aws redshift describe-clusters \
          --cluster-identifier "$NEW_CLUSTER_ID" \
          --query "Clusters[0].ClusterStatus" \
          --output text
        ```

        ***

        ## 7. Update applications and clients

        Get new endpoint:

        ```bash theme={null}
        aws redshift describe-clusters \
          --cluster-identifier "$NEW_CLUSTER_ID" \
          --query "Clusters[0].[Endpoint.Address,Endpoint.Port]" \
          --output table
        ```

        Update all applications to use this new endpoint.

        ***

        ## 8. Decommission the old non‑VPC cluster

        Once you’ve confirmed everything works:

        ```bash theme={null}
        aws redshift delete-cluster \
          --cluster-identifier "$OLD_CLUSTER_ID" \
          --skip-final-cluster-snapshot
        ```

        (Or omit `--skip-final-cluster-snapshot` and provide `--final-cluster-snapshot-identifier` if you want another final snapshot.)

        ***

        After these steps, your Redshift cluster runs inside a VPC and satisfies the “must be in VPC” requirement.
      </Accordion>

      <Accordion title="Using Python">
        For Redshift, “cluster in VPC” effectively means: each Redshift cluster must be associated with a VPC subnet group (i.e., have a non‑null `VpcId` and `ClusterSubnetGroupName`). Old “EC2-Classic” style clusters are the issue.

        Below is how to (1) detect non‑VPC Redshift clusters and (2) migrate them into a VPC using Python/boto3.

        ***

        ## 1. Prereqs

        ```bash theme={null}
        pip install boto3
        aws configure   # configure credentials + region
        ```

        You’ll also need:

        * A target VPC ID (e.g. `vpc-1234567890abcdef0`)
        * At least two private subnets in that VPC (for multi‑AZ best practice), e.g.:
          * `subnet-aaa...`
          * `subnet-bbb...`
        * A security group for Redshift, e.g. `sg-1234...` (with appropriate inbound rules from your app/BI tools and outbound allowed)

        ***

        ## 2. Detect Redshift clusters not in a VPC

        ```python theme={null}
        import boto3

        redshift = boto3.client('redshift')

        def get_non_vpc_clusters():
            paginator = redshift.get_paginator('describe_clusters')
            non_vpc = []
            for page in paginator.paginate():
                for c in page['Clusters']:
                    # old style clusters have no VpcId
                    if 'VpcId' not in c or c.get('VpcId') in (None, ''):
                        non_vpc.append(c)
            return non_vpc

        clusters = get_non_vpc_clusters()
        for c in clusters:
            print(c['ClusterIdentifier'])
        ```

        ***

        ## 3. Create / verify a Redshift subnet group for your VPC

        ```python theme={null}
        TARGET_VPC_SUBNET_IDS = [
            "subnet-aaaaaaaa",
            "subnet-bbbbbbbb",
        ]  # private subnets in same VPC

        SUBNET_GROUP_NAME = "redshift-vpc-subnet-group"
        SUBNET_GROUP_DESC = "Subnet group for Redshift VPC migration"

        def ensure_subnet_group():
            try:
                redshift.describe_cluster_subnet_groups(
                    ClusterSubnetGroupName=SUBNET_GROUP_NAME
                )
                print("Subnet group exists:", SUBNET_GROUP_NAME)
            except redshift.exceptions.ClusterSubnetGroupNotFoundFault:
                print("Creating subnet group:", SUBNET_GROUP_NAME)
                redshift.create_cluster_subnet_group(
                    ClusterSubnetGroupName=SUBNET_GROUP_NAME,
                    Description=SUBNET_GROUP_DESC,
                    SubnetIds=TARGET_VPC_SUBNET_IDS
                )

        ensure_subnet_group()
        ```

        ***

        ## 4. Migrate each non‑VPC cluster into the VPC

        You can’t “flip” an existing non‑VPC Redshift cluster into a VPC in‑place. The safe pattern:

        1. Take a snapshot of the existing cluster.
        2. Restore a new cluster *from that snapshot* into the VPC subnet group, with a VPC security group.
        3. Cut over DNS / application connections to the new endpoint.
        4. Delete the old non‑VPC cluster and old snapshot (once validated).

        ### 4.1 Snapshot the original cluster

        ```python theme={null}
        import time

        def create_manual_snapshot(cluster_id, snapshot_id):
            print(f"Creating snapshot {snapshot_id} of cluster {cluster_id}")
            redshift.create_cluster_snapshot(
                SnapshotIdentifier=snapshot_id,
                ClusterIdentifier=cluster_id
            )

            # Wait until snapshot is available
            waiter = redshift.get_waiter('snapshot_available')
            waiter.wait(SnapshotIdentifier=snapshot_id)
            print("Snapshot available:", snapshot_id)

        for c in clusters:
            old_cluster_id = c['ClusterIdentifier']
            snapshot_id = f"{old_cluster_id}-pre-vpc-migration"
            create_manual_snapshot(old_cluster_id, snapshot_id)
        ```

        ***

        ### 4.2 Restore a new VPC-based cluster from the snapshot

        You’ll map the old cluster to a new ID (e.g., append `-vpc`), in a VPC subnet group and security group.

        ```python theme={null}
        VPC_SECURITY_GROUP_IDS = ["sg-1234567890abcdef0"]  # Redshift SG in your VPC
        PUBLICLY_ACCESSIBLE = False                        # usually recommended

        def restore_cluster_in_vpc(old_cluster, snapshot_id):
            old_id = old_cluster['ClusterIdentifier']
            new_id = f"{old_id}-vpc"

            print(f"Restoring new cluster {new_id} from snapshot {snapshot_id}")

            # Use same node type, count, and parameter group where possible
            redshift.restore_from_cluster_snapshot(
                ClusterIdentifier=new_id,
                SnapshotIdentifier=snapshot_id,
                ClusterSubnetGroupName=SUBNET_GROUP_NAME,
                VpcSecurityGroupIds=VPC_SECURITY_GROUP_IDS,
                PubliclyAccessible=PUBLICLY_ACCESSIBLE,
                # Optional: reuse param group, maintenance settings, etc.
                ClusterParameterGroupName=old_cluster['ClusterParameterGroups'][0]['ParameterGroupName']
                    if old_cluster.get('ClusterParameterGroups') else None,
                Port=old_cluster.get('Endpoint', {}).get('Port', 5439),
                # other optional args: EnhancedVpcRouting, AutomatedSnapshotRetentionPeriod, etc.
            )

            # Wait for cluster to be available
            waiter = redshift.get_waiter('cluster_available')
            waiter.wait(ClusterIdentifier=new_id)
            print("New VPC cluster available:", new_id)

            # Show new endpoint
            new_desc = redshift.describe_clusters(ClusterIdentifier=new_id)['Clusters'][0]
            print("New endpoint:", new_desc['Endpoint']['Address'], new_desc['Endpoint']['Port'])

        for c in clusters:
            old_id = c['ClusterIdentifier']
            snapshot_id = f"{old_id}-pre-vpc-migration"
            restore_cluster_in_vpc(c, snapshot_id)
        ```

        ***

        ## 5. Cutover and clean up

        1. Update application connection strings to use the **new cluster’s endpoint** (hostname + port + DB name).
        2. Validate queries / workloads on the new VPC cluster.
        3. When fully satisfied, delete the old cluster and snapshot.

        ```python theme={null}
        def delete_old_cluster_and_snapshot(old_cluster_id, snapshot_id):
            print(f"Deleting old cluster: {old_cluster_id}")
            redshift.delete_cluster(
                ClusterIdentifier=old_cluster_id,
                SkipFinalClusterSnapshot=True  # or False, if you want one last snapshot
            )
            waiter = redshift.get_waiter('cluster_deleted')
            waiter.wait(ClusterIdentifier=old_cluster_id)
            print("Old cluster deleted:", old_cluster_id)

            print(f"Deleting snapshot: {snapshot_id}")
            redshift.delete_cluster_snapshot(SnapshotIdentifier=snapshot_id)

        for c in clusters:
            old_id = c['ClusterIdentifier']
            snapshot_id = f"{old_id}-pre-vpc-migration"
            # Run this ONLY after you have validated new VPC cluster
            # delete_old_cluster_and_snapshot(old_id, snapshot_id)
        ```

        ***

        ## 6. Summary of remediation

        1. Detect Redshift clusters with no `VpcId` (non‑VPC).
        2. Create a Redshift subnet group for your target VPC.
        3. For each non‑VPC cluster:
           * Take a manual snapshot.
           * Restore a new cluster from that snapshot in the VPC (subnet group + VPC SG).
           * Switch clients to the new endpoint.
           * Delete old non‑VPC cluster and its snapshot once verified.

        This ensures all Redshift clusters run in a VPC, satisfying the “cluster should be in VPC” control, using Python automation.
      </Accordion>

      <Accordion title="Using Terraform">
        ```hcl theme={null}
        # Place the Redshift cluster in a VPC by using a subnet group and VPC security groups.

        resource "aws_redshift_subnet_group" "redshift_vpc_subnets" {
          name       = "redshift-subnet-group"
          description = "Redshift subnet group in VPC"

          subnet_ids = [
            aws_subnet.REDSHIFT_SUBNET_1_ID, # replace with your subnet resource or ID
            aws_subnet.REDSHIFT_SUBNET_2_ID, # must be in the same VPC, different AZs recommended
          ]
        }

        resource "aws_security_group" "redshift_sg" {
          name        = "redshift-sg"
          description = "Security group for Redshift in VPC"
          vpc_id      = aws_vpc.REDSHIFT_VPC_ID # replace with your VPC resource or ID

          # Add ingress/egress rules as needed
          ingress {
            from_port   = 5439
            to_port     = 5439
            protocol    = "tcp"
            cidr_blocks = ["ALLOWED_CIDR_BLOCK"] # replace with appropriate CIDR or use security_group_id refs
          }

          egress {
            from_port   = 0
            to_port     = 0
            protocol    = "-1"
            cidr_blocks = ["0.0.0.0/0"]
          }
        }

        resource "aws_redshift_cluster" "redshift_cluster" {
          cluster_identifier = "YOUR_REDSHIFT_CLUSTER_IDENTIFIER"  # replace
          node_type          = "dc2.large"                         # replace
          master_username    = "MASTER_USERNAME"                   # replace
          master_password    = "MASTER_PASSWORD"                   # replace with secret handling

          cluster_subnet_group_name = aws_redshift_subnet_group.redshift_vpc_subnets.name
          vpc_security_group_ids    = [aws_security_group.redshift_sg.id]

          # other required arguments for your environment...
          # number_of_nodes       = 2
          # iam_roles             = [...]
          # encrypted             = true
        }

        ```

        Moving an existing Redshift cluster from EC2-Classic to a VPC (by adding `cluster_subnet_group_name` / `vpc_security_group_ids`) forces replacement of the cluster; this can cause downtime and potential data loss if not migrated carefully.

        For verification, `terraform plan` should show the Redshift cluster either being created with `cluster_subnet_group_name` and `vpc_security_group_ids` set, or (for an existing non‑VPC cluster) a `-/+` replacement where the new resource includes those VPC fields.
      </Accordion>
    </AccordionGroup>
  </Tab>
</Tabs>

### Additional Reading:

* \[[https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-clusters-in-a-vpc.html](https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-clusters-in-a-vpc.html)]
