Thursday, April 29, 2021

Quick review: One month study | Algorithm study | Less than 20 algorithms

 April 29, 2021

Introduction

I just could not believe what I did. I spent whole month to study system design, but I only gave myself less than 24 hours to prepare algorithm interview. It is tough situation for me to handle. 

Algorithm review

I need to review at least 100 algorithms and also it is better for me to cover those medium and level algorithms. 

I need to go to sleep first. Now it is 1:27 AM. I did take a nap during the day time, and I slept two hours from 3 PM to 5 PM. 


Leetcode discuss: 1090. Largest Values From Labels

 April 29, 2021

Here is the link. 

C# | greedy algorithm | Two hashmaps to track records | Case study

April 29, 2021
Introduction
I quickly reviewed my own C# code written in 2019, and then made a few changes to speed up the code.

Case study

	 var values = new int[] { 5, 4, 3, 2, 1 };
     var labels = new int[] { 1, 3, 3, 3, 2 };

     var result = LargestValsFromLabels(values, labels, 3, 2);
     Debug.Assert(result == 12);

The question is to ask what is maximum value to find, three numbers, total number of values with same label should not exceed 2. The answer is 12, which is sum of 5 + 4 + 3. The greedy algorithm is applied to find largest value first, and then continue to try to add next smaller one if label count constraint is not broken. The detail is written in the following.

The idea is to put values into C# Dictionary object, a hashmap variable called "map", so values = {5, 4, 3, 2, 1} will be saved into hashmap, for example, first value is 5, label is 1, so map[5] = 1. All values are saved in the following:
map[5] = 1, map[4] = 3, map[3] = 3, map[2] = 3, map[1] = 2.

Next step is to get all keys from hashmap map, and put into an array. Sort the array in descending order. Apply greedy algorithm, go over largest key first, and then choose it into the set if it's label count is still available.

The extra work is to use a hashmap called "used" to track how many numbers are chosed based on label value.

The biggest key is 5, so it is chosen. used hashmap used[5] =1; next biggest key is 4, and 4 is chosen, since value 4's label is 3, so used[3] = 1. Next biggest key is 3, and 3's label is 3, and used[3] = 1 < 2, so 3 can be added to largest value set. All three values are found with sum = 5 + 4 + 3 = 12.

Test case
values = [3,2,3,2,1]
labels = [1,0,2,2,1]
The number in values array may not be unique, so in other words, one number may have different labels. The above test case, the first and third number both are 3, but labels are different. One is 1, second one is 2. In my design, C# variable name map is defined using Dictionary<int, List>, not Dictionary<int, int>.

Simplicity
Avoid using C# SortedDictionary, since keys only needs to be sorted once after all are saved. No need to maintain a SortedDictionary data structure.

The following code was modified based on my previous practice here. The following code passes online judge.

public class Solution {
 static void Main(string[] args)
        {
            RunTestcase2();
        }

        public static void RunTestcase1()
        {
            var values = new int[] { 5, 4, 3, 2, 1 };
            var labels = new int[] { 1, 1, 2, 2, 3 };

            var result = LargestValsFromLabels(values, labels, 3, 1);
            Debug.Assert(result == 9);
        }

        public static void RunTestcase2()
        {
            var values = new int[] { 5, 4, 3, 2, 1 };
            var labels = new int[] { 1, 3, 3, 3, 2 };

            var result = LargestValsFromLabels(values, labels, 3, 2);
            Debug.Assert(result == 12);
        }
		
    public int LargestValsFromLabels(int[] values, int[] labels, int num_wanted, int use_limit)
        {
            var map = new Dictionary<int, List<int>>();

            var length = values.Length;
            for (int i = 0; i < length; i++)
            {
                var key = values[i];
                var value = labels[i];

                if (!map.ContainsKey(key))
                {
                    map.Add(key, new List<int>());
                }

                map[key].Add(value);
            }

            var usedCount = new Dictionary<int, int>();

           
            var numbers = map.Keys.ToArray();
            Array.Sort(numbers);
        
		    int index = 0;
            length = numbers.Length;
            var sum = 0;            
 
            for (int i = length - 1; i >= 0; i--)
            {
                var key  = numbers[i];
                var list = map[key];

                foreach (var item in list)
                {
                    if (!usedCount.ContainsKey(item))
                    {
                        usedCount.Add(item, 0);
                    }

                    if (usedCount[item] < use_limit)
                    {
                        index++;
                        sum += key;

                        usedCount[item]++;

                        if (index == num_wanted)
                        {
                            return sum; 
                        }
                    }
                }                
            }

            return sum; 
        }
}

Wednesday, April 28, 2021

System design: Introduction to AWS Networking (Re-uploaded for those who had audio issues)

April 28, 2021

Here is the link. 

This video provides high level overview of all AWS networking services and components and how they fit into any architecture. This covers introduction to VPC, Subnets, Route Tables, IGW, NAT, Site to Site VPN, Direct Connect, VPC Peering, Transit Gateway, VPC endpoint, PrivateLink, Route53 and CloudFront. My udemy course: [https://www.udemy.com/course/networki...​]


  • VPC Basics
  • ELB/CloudFront
  • Route 53 (DNS)
  • VPC Endpoint
  • Private Link
  • VPC Peering

  • CIDR
  • Internet Gateway
  • Subnets
  • Route tables
  • NAT
  • IP (public, private, elastic), NAT
  • Security group
  • Network ACL

System design: Apache Kafka with Spark Streaming | Kafka Spark Streaming Examples | Kafka Training | Edureka

 April 28, 2021

Here is the link. 

** Kafka Online Training : https://www.edureka.co/kafka-certific...​ ** In this Kafka Spark Streaming video, we are demonstrating how Apache Kafka works with Spark Streaming. In this video, we have discussed Apache Kafka & Apache Spark briefly. Finally, we have explained the integration of Kafka & Spark Streaming. Topics covered in this Kafka Spark Streaming Tutorial video are: 1. What is Kafka? 2. Kafka Components 3. Kafka architecture 4. What is Spark? 5. Spark Component 6. Kafka Spark Integration 7. Kafka Spark Streaming Project As mentioned in the video, you can go through these Kafka & Spark videos: Kafka Tutorial: https://www.youtube.com/watch?v=hyJZP...​ Spark Tutorial: https://www.youtube.com/watch?v=9mELE...​ Spark Streaming: https://www.youtube.com/watch?v=uD_q4...​ Subscribe to our channel to get video updates. Hit the subscribe button above. Check our complete Kafka playlist here: https://goo.gl/jEZfLj​ - - - - - - - - - - - - - - How it Works? 1. This is a 5 Week Instructor led Online Course, 40 hours of assignment and 30 hours of project work 2. We have a 24x7 One-on-One LIVE Technical Support to help you with any problems you might face or any clarifications you may require during the course. 3. At the end of the training you will have to undergo a 2-hour LIVE Practical Exam based on which we will provide you a Grade and a Verifiable Certificate! - - - - - - - - - - - - - - About the Course Edureka’s Apache Kafka certification training is designed to help you become a Kafka developer. During this course, our expert Kafka instructors will help you: 1. Learn Kafka and its components 2. Set up an end to end Kafka cluster along with Hadoop and YARN cluster 3. Integrate Kafka with real time streaming systems like Spark & Storm 4. Describe the basic and advanced features involved in designing and developing a high throughput messaging system 5. Use Kafka to produce and consume messages from various sources including real time streaming sources like Twitter 6. Get insights of Kafka Producer & Consumer APIs 7. Understand Kafka Stream APIs 8. Work on a real-life project, ‘Implementing Twitter Streaming with Kafka, Flume, Hadoop & Storm - - - - - - - - - - - - - - Who should go for this course? This course is designed for professionals who want to learn Kafka techniques and wish to apply it on Big Data. It is highly recommended for: Developers, who want to gain acceleration in their career as a "Kafka Big Data Developer" Testing Professionals, who are currently involved in Queuing and Messaging Systems Big Data Architects, who like to include Kafka in their ecosystem Project Managers, who are working on projects related to Messaging Systems Admins, who want to gain acceleration in their careers as a "Apache Kafka Administrator - - - - - - - - - - - - - - Why Learn Kafka? Kafka is used heavily in the Big Data space as a reliable way to ingest and move large amounts of data very quickly. ​LinkedIn, Yahoo, Twitter, Netflix, Uber, Goldman Sachs,PayPal, Airbnb​ ​​& other fortune 500 companies use Kafka. The average salary of a Software Engineer with Apache Kafka skill is $87,500 per year. (Payscale.com salary data). - - - - - - - - - - - - - For more information, Please write back to us at sales@edureka.co or call us at IND: 9606058406 / US: 18338555775 (toll-free). Instagram: https://www.instagram.com/edureka_lea...​ Facebook: https://www.facebook.com/edurekaIN/​ Twitter: https://twitter.com/edurekain​ LinkedIn: https://www.linkedin.com/company/edureka

Kafka decouples data pipelines









System design: Kafka Tutorial for Beginners - Setup Kafka on Hortonworks in 30 min! - Frank Kane

April 28, 2021

Here is the link. 

Explore the full course on Udemy (special discount included in the link): https://www.udemy.com/the-ultimate-ha...​ Learn to stream big data with Kafka, starting from scratch. Kafka is a powerful data streaming technology and a very hot technical skill to have right now. With Kafka, you can publish streams of data from web logs, sensors, or whatever else you can imagine to systems that manipulate, analyze, and store that data all in real time. Kafka bring a reliable publish / subscribe mechanism that is resilient and can allow clients to pick up where they left off in the event of an outage. In this tutorial, you will set up a free Hortonworks sandbox environment within a virtual Linux machine running right on your own desktop PC, learn about how data streaming and Kafka work, set up Kafka, and use it to publish real web logs on a Kafka topic and receive them in real time. Kafka is sometimes billed as a Hadoop killer due to its power, but really it is an integral piece of the larger Hadoop ecosystem that has emerged. Course Description The world of Hadoop and "Big Data" can be intimidating - hundreds of different technologies with cryptic names form the Hadoop ecosystem. With this course, you'll not only understand what those systems are and how they fit together - but you'll go hands-on and learn how to use them to solve real business problems! Learn and master the most popular big data technologies in this comprehensive course, taught by a former engineer and senior manager from Amazon and IMDb. We'll go way beyond Hadoop itself, and dive into all sorts of distributed systems you may need to integrate with. Install and work with a real Hadoop installation right on your desktop with Hortonworks and the Ambari UI Manage big data on a cluster with HDFS and MapReduce Write programs to analyze data on Hadoop with Pig and Spark Store and query your data with Sqoop, Hive, MySQL, HBase, Cassandra, MongoDB, Drill, Phoenix, and Presto Design real-world systems using the Hadoop ecosystem Learn how your cluster is managed with YARN, Mesos, Zookeeper, Oozie, Zeppelin, and Hue Handle streaming data in real time with Kafka, Flume, Spark Streaming, Flink, and Storm Understanding Hadoop is a highly valuable skill for anyone working at companies with large amounts of data. Almost every large company you might want to work at uses Hadoop in some way, including Amazon, Ebay, Facebook, Google, LinkedIn, IBM, Spotify, Twitter, and Yahoo! And it's not just technology companies that need Hadoop; even the New York Times uses Hadoop for processing images. This course is comprehensive, covering over 25 different technologies in over 14 hours of video lectures. It's filled with hands-on activities and exercises, so you get some real experience in using Hadoop - it's not just theory. You'll find a range of activities in this course for people at every level. If you're a project manager who just wants to learn the buzzwords, there are web UI's for many of the activities in the course that require no programming knowledge. If you're comfortable with command lines, we'll show you how to work with them too. And if you're a programmer, I'll challenge you with writing real scripts on a Hadoop system using Scala, Pig Latin, and Python. You'll walk away from this course with a real, deep understanding of Hadoop and its associated distributed systems, and you can apply Hadoop to real-world problems. Plus a valuable completion certificate is waiting for you at the end! Please note the focus on this course is on application development, not Hadoop administration. Although you will pick up some administration skills along the way. I hope to see you in the course soon! -Frank Who is the target audience? Software engineers and programmers who want to understand the larger Hadoop ecosystem, and use it to store, analyze, and vend "big data" at scale. Project, program, or product managers who want to understand the lingo and high-level architecture of Hadoop. Data analysts and database administrators who are curious about Hadoop and how it relates to their work. System architects who need to understand the components available in the Hadoop ecosystem, and how they fit together. Your instructor is Frank Kane of Sundog Education, bringing nine years of experience as a senior engineer and senior manager at Amazon.com and IMDb.com, where his job involved extracting meaning from their massive data sets, and processing that data in a highly distributed manner.

System design: How to Choose the Right Database? - MongoDB, Cassandra, MySQL, HBase - Frank Kane

 April 28, 2021

Here is the link. 

Explore the full course on Udemy (special discount included in the link): https://www.udemy.com/the-ultimate-ha...​ Choosing the right database for your application is no easy task. You have a wide variety of options relational databases such as MySQL, or distributed NoSQL solutions such as MongoDB, Cassandra, and HBase. NoSQL has come to mean not only SQL as many distributed database systems do in fact support SQL-style queries, as long as you are not doing complex join operations and this further blurs the lines between these systems. We will talk about how to analyze the requirements of your system in terms of consistency, availability, and partition-tolerance, and how to apply the CAP theorem to guide your choice after showing you where different database technologies fall on the sides of the CAP triangle. We will also talk about more practical considerations, such as your budget, need for professional support, and the ease of integration into the other systems already in place in your organization. Maybe you dont even need a distributed storage solution at all! Choosing the right technology for your data storage will save you a lot of pain as your application grows and evolves and making the wrong choice can lead to all sorts of maintenance problems and wasted work. Your instructor is Frank Kane of Sundog Education, bringing nine years of experience as a senior engineer and senior manager at Amazon.com and IMDb.com, where his job involved extracting meaning from their massive data sets, and processing that data in a highly distributed manner.

5:08 CAP consideration 7:55 Keep it simple 9:05 example - simple phone directory app 10:57 another example 13:07 example for Cassandra 14:54 build a massive stock trading system 15:02 care about consistency more than anything 15:14 deal with big data 16:01 discussion about choices 17:46

Scaling requirements- transaction rate -
Trading system design


System design: Build a Serverless Web App for a Theme Park: Episode 1 - AWS Virtual Workshop

 April 28, 2021

Here is the link. 

Learn how to build a complete serverless web application for a popular theme park called Innovator Island. The theme park is rolling out a mobile app that provides thousands of visitors with wait times, photo opportunities, notification alerts, and language translation for visitors who need it. In this session - the first of a five-part series - we'll cover the application scenario, the architecture, and include a brief introduction to each of the major AWS services used, like AWS Amplify Console, and the AWS Serverless Application Model. Practically, we'll cover how to set up your environment, deploy the backend, and learn how to deploy the front-end code automatically so you’re ready to set up real time messaging with customers in the next session.

Learning Objectives: - Learning about important AWS services in serverless technology - Deploying front-end with Amplify Console - Deploying backend with the AWS Serverless Application Model (SAM) To learn more about the services featured in this workshop, please visit: https://aws.amazon.com/serverless

System design: What is Amazon VPC?

 

What is Amazon VPC?

Amazon Virtual Private Cloud (Amazon VPC) enables you to launch AWS resources into a virtual network that you've defined. This virtual network closely resembles a traditional network that you'd operate in your own data center, with the benefits of using the scalable infrastructure of AWS.

Amazon VPC concepts

Amazon VPC is the networking layer for Amazon EC2. If you're new to Amazon EC2, see What is Amazon EC2? in the Amazon EC2 User Guide for Linux Instances to get a brief overview.

The following are the key concepts for VPCs:

  • Virtual private cloud (VPC) — A virtual network dedicated to your AWS account.

  • Subnet — A range of IP addresses in your VPC.

  • Route table — A set of rules, called routes, that are used to determine where network traffic is directed.

  • Internet gateway — A gateway that you attach to your VPC to enable communication between resources in your VPC and the internet.

  • VPC endpoint — Enables you to privately connect your VPC to supported AWS services and VPC endpoint services powered by PrivateLink without requiring an internet gateway, NAT device, VPN connection, or AWS Direct Connect connection. Instances in your VPC do not require public IP addresses to communicate with resources in the service. Traffic between your VPC and the other service does not leave the Amazon network. For more information, see AWS PrivateLink and VPC endpoints.

  • CIDR block —Classless Inter-Domain Routing. An internet protocol address allocation and route aggregation methodology. For more information, see Classless Inter-Domain Routing in Wikipedia.

Accessing Amazon VPC

You can create, access, and manage your VPCs using any of the following interfaces:

  • AWS Management Console — Provides a web interface that you can use to access your VPCs.

  • AWS Command Line Interface (AWS CLI) — Provides commands for a broad set of AWS services, including Amazon VPC, and is supported on Windows, Mac, and Linux. For more information, see AWS Command Line Interface.

  • AWS SDKs — Provides language-specific APIs and takes care of many of the connection details, such as calculating signatures, handling request retries, and error handling. For more information, see AWS SDKs.

  • Query API — Provides low-level API actions that you call using HTTPS requests. Using the Query API is the most direct way to access Amazon VPC, but it requires that your application handle low-level details such as generating the hash to sign the request, and error handling. For more information, see the Amazon EC2 API Reference.