> For the complete documentation index, see [llms.txt](https://mamawhocode.gitbook.io/aws/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://mamawhocode.gitbook.io/aws/services/analytics/data-processing/kinesis.md).

# Kinesis

a suite of services that helps you work with streaming data

[AWS Document](https://docs.aws.amazon.com/kinesis/index.html) | [Data Analytics](https://explore.skillbuilder.aws/learn/course/internal/view/elearning/44/data-analytics-fundamentals) | [Kinesis Streams](https://explore.skillbuilder.aws/learn/course/internal/view/elearning/157/introduction-to-amazon-kinesis-streams) | [Apache Flink](https://docs.aws.amazon.com/kinesisanalytics/latest/java/what-is.html)

## Overview

* **collect**, **process**, **analyze** video & data streams in <mark style="color:red;">`real-time`</mark>.
* Real-time data: Application logs, Metrics, Website clickstreams, IoT telemetry data.
* operate in several modes
  * Data Streams
  * Firehose
  * Managed Apache Flink (old: Data Analytics)
  * Video stream&#x20;

    <figure><img src="https://2259236002-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fuh9xZDZ53qGqmMCM44PU%2Fuploads%2Fgit-blob-bbc2902f9f9a8c23d8c75552c258c4a9fa5bcee5%2Ffigure_20230503150852.png?alt=media" alt=""><figcaption></figcaption></figure>

### Use cases

* Application monitoring
* Fraud detection
* Live game leaderboards
* IoT
* Sentiment analysis

## Features

| **Service**                                                                  | **Description**                                                                                             | **Use Case**                                                                                                                                                                                                        |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Kinesis Data Streams                                                         | capture, process & store data stream                                                                        | Ingesting data from various sources, processing data in *<mark style="color:red;">real-time</mark>*, performing *<mark style="color:red;">real-time</mark>* (millisecond) analytics                                 |
| Kinesis Firehose                                                             | built-in *<mark style="color:red;">**Transform**</mark>*, deliver streaming data *directly to AWS services* | ***near real-time*** (buffer 1 min to 15min) storing and analyzing large amounts of data over time without managing your own data pipeline, <mark style="color:red;">**simple transformation**</mark>, auto scaling |
| Kinesis Data Analytics -> Amazon Managed service for Apache Flink (**MSAF**) | Process and analyze streaming data using <mark style="color:red;">`standard SQL`</mark> queries             | Real-time data analytics on data streams without the need for specialized programming skills                                                                                                                        |
| Kinesis Video Streams                                                        | Securely stream video from connected devices to AWS for analysis and processing                             | Capturing video from security cameras, drones, and IoT sensors, and analyzing the data in real-time                                                                                                                 |

### Kinesis Data Streams

![](https://2259236002-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fuh9xZDZ53qGqmMCM44PU%2Fuploads%2Fgit-blob-ff76a8431b1373f75af14c98f358a2f6065062d7%2Ffigure_20230503164714.png?alt=media)

* Producer
  * Kinesis Agent
  * AWS SDK
  * Kinesis Producer Library (KPL)
* Consumer
  * **Kinesis Data Analytics**: use an Amazon Kinesis Data Analytics application to process and analyze using SQL or Java.
  * **Kinesis Firehose**: use an Amazon Kinesis Data Firehose delivery stream to process and store records in a destination.
  * **Kinesis Client Library** (KCL): use Kinesis Client Library to develop consumers.

#### Capacity modes

* Provisioned mode
  * Choose the number of shards
  * Scale `manually` using API
* On-demand mode
  * No nead to provision or manage capicty
  * Scale automatically based on observed throughput peak during the last 30 days.

### Kinesis Firehose

* Serverless, fully managed, automatic scaling.
* Supports *custom data transformations* using Lambda&#x20;

  <figure><img src="https://2259236002-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fuh9xZDZ53qGqmMCM44PU%2Fuploads%2Fgit-blob-7612a4dabc51fffb5de0084c039eb6c38fdac22f%2Ffigure_20230503171037.png?alt=media" alt=""><figcaption></figcaption></figure>
* Producer
* Consumer
  * AWS S3
  * Redshift
  * OpenSearch
  * 3rd party: Splunk, MongoDB, DataDog, NewRelic, HoneyCom...
  * HTTP Endpoint
* Use cases:
  * Main usage scenarios for *<mark style="color:red;">CloudWatch metric streams</mark>*: Data lake— Create a metric stream and direct it to an Amazon Kinesis Data Firehose delivery stream that delivers your CloudWatch metrics to a data lake such as Amazon S3.

### Kinesis Data Analytics/ Managed service for Apache Flink (MSAF)

* Reads and processes real-time streaming data.&#x20;

  <figure><img src="https://2259236002-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fuh9xZDZ53qGqmMCM44PU%2Fuploads%2Fgit-blob-ea3de72e95709e8c61385800174603d3a6b45506%2Ffigure_20230503163220.png?alt=media" alt=""><figcaption><p>Kinesis Data Analytics/ MSAF</p></figcaption></figure>

### Kinesis Video Streams

## Best practices

* Increase number of shards in your Kinesis Data stream to handle increase throughput/traffic (resolve `ProvisionedThroughputExceeded` problem).
* Use **random partition key** to deal with hot shard problem (unevenly distributed stream)

## Trivia

* `real-time` or `near-real time` = Kinesis Data Stream.
* Kinesis Data stream uses the partition key associated with each data record to determine which shard a data record belongs to.
* **Multiple** Kinesis Data Streams applications can consume data from **a stream**.
* Firehose does not support DynamoDB. [refer](/aws/services/analytics/data-processing/kinesis.md#kinesis-firehose)
* `PutRecords` request can support up to 500 records. Each record in the request can be as large as 1 MiB, up to a limit of 5 MiB.
* `ProvisionedThroughputExceededException`: when there is throttling, it best practices to
  * Implement retries with exponential backoff.
  * Increase Shard Count
  * Optimize Data Send Rat&#x65;**:** If possible, batch records to use the `PutRecords` API
  * Reduce the frequency and/or size of the requests.
  * Uniformly Distribute Partition Keys
* Redshift applies compression to columns to reduce storage size and improve query speed. To determine the best compression encoding for a table, use the `ANALYZE COMPRESSION` command.

## Concepts

* [Producer (upstream)](/aws/services/analytics.md): a producer `put` records into Kinesis
* [Consumer (downstream)](/aws/services/analytics.md): a consumer `get` records from Kinesis
* [Sharding](/aws/services/analytics.md): DB sharding is the processing of breaking up large tables into multiple smaller tables, or chunks called shards. So sharding is horizontal partitioning. ![shard](https://miro.medium.com/v2/resize:fit:1400/format:webp/1*-3CrSE3jsfH1AQd8BB2tcA.png)
  * [Shard](/aws/services/analytics/data-processing/kinesis.md):&#x20;
    * a shard is a unit of throughput capacity.&#x20;
    * The number of instances does not exceed the number of open shards. Each shard is processed by exactly one KCL worker and has exactly one corresponding record processor, so you never need multiple instances to process one shard. However, one worker can process any number of shards, so it's fine if the number of shards exceeds the number of instances.
* [Clickstream](/aws/services/devtools/cloudformation.md#concepts): Clickstream data is a record of a user's activity on the internet, including every click they make while browsing a website or using an application.
* [Sub-Optimal Encoding](/aws/services/management/config.md#concepts): occurs when the applied compression method is not the best fit for the data, leading to inefficiencies in storage and performance.
