Sunday, August 9, 2020

TAO: The power of the graph

 Here is the article. 


I like to read the article again. It is tough for me to learn the basics from the article. I will figure out the most important concepts in the article first. 

My notes: 

Requirement: the data set must be retrieved and rendered on the fly in a few hundred milliseconds

Facebook puts an extremely demanding workload on its data backend. Every time any one of over a billion active users visits Facebook through a desktop browser or on a mobile device, they are presented with hundreds of pieces of information from the social graph. Users see News Feed stories; comments, likes, and shares for those stories; photos and check-ins from their friends -- the list goes on. The high degree of output customization, combined with a high update rate of a typical user’s News Feed, makes it impossible to generate the views presented to users ahead of time. Thus, the data set must be retrieved and rendered on the fly in a few hundred milliseconds.

two factors to consider:

  1. the high degree of output customization
  2. a high update rate of a typical user's News Feed
  3. It is impossible to generate the views presented to users ahead of time. - The argument
Data set is not easily partitionable
photos of celebrities, request rates spike significantly
multiple this by the millions of times per second - 
read-dominated workload 

This challenge is made more difficult because the data set is not easily partitionable, and by the tendency of some items, such as photos of celebrities, to have request rates that can spike significantly. Multiply this by the millions of times per second this kind of highly customized data set must be delivered to users, and you have a constantly changing, read-dominated workload that is incredibly challenging to serve efficiently.

Memcache and MySQL

Facebook has always realized that even the best relational database technology available is a poor match for this challenge unless it is supplemented by a large distributed cache that offloads the persistent store. Memcache has played that role since Mark Zuckerberg installed it on Facebook’s Apache web servers back in 2005. As efficient as MySQL is at managing data on disk, the assumptions built into the InnoDB buffer pool algorithms don’t match the request pattern of serving the social graph. The spatial locality on ordered data sets that a block cache attempts to exploit is not common in Facebook workloads. Instead, what we call creation time locality dominates the workload -- a data item is likely to be accessed if it has been recently created. Another source of mismatch between our workload and the design assumptions of a block cache is the fact that a relatively large percentage of requests are for relations that do not exist -- e.g., “Does this user like that story?” is false for most of the stories in a user’s News Feed. Given the overall lack of spatial locality, pulling several kilobytes of data into a block cache to answer such queries just pollutes the cache and contributes to the lower overall hit rate in the block cache of a persistent store.


What is the following statement: 

the lower overall hit rate in the block cache of a persistent store.

persistent store -> block cache -> lower overall hit rate 


The use of memcache vastly improved the memory efficiency of caching the social graph and allowed us to scale in a cost-effective way. However, the code that product engineers had to write for storing and retrieving their data became quite complex. Even though memcache has “cache” in its name, it’s really a general-purpose networked in-memory data store with a key-value data model. It will not automatically fill itself on a cache miss or maintain cache consistency. Product engineers had to work with two data stores and very different data models: a large cluster of MySQL servers for storing data persistently in relational tables, and an equally large collection of memcache servers for storing and serving flat key-value pairs derived (some indirectly) from the results of SQL queries. Even with most of the common chores encapsulated in a data access library, using the memcache-MySQL combination efficiently as a data store required quite a bit of knowledge of system internals on the part of product engineers. Inevitably, some made mistakes that led to bugs, user-visible inconsistencies, and site performance issues. In addition, changing table schemas as products evolved required coordination between engineers and MySQL cluster operators. This slowed down the change-debug-release cycle and didn’t fit well with Facebook's “move fast” development philosophy.

memcache 

a general-purpose networked in-memory data store with a key-value data model

Product engineers had to work with two data stores and very different data models: a large cluster of MySQL servers for storing data persistently in relational tables, and an equally large collection of memcache servers for storing and serving flat key-value pairs derived (some indirectly) from the results of SQL queries.

How I quickly go over the content in my own words? 8/9/2020 3:53 PM

Product engineers 

two data stores - data models are different 

a large cluster of MySQL servers for storing data persistently in relational tables

an equally large collection of memcache servers - storing and serving flat key-value pairs derived (some indirectly) from the results of SQL queries

Objects and associations

In 2007, a few Facebook engineers set out to define new data storage abstractions that would fit the needs of all but the most demanding features of the site while hiding most of the complexity of the underlying distributed data store from product engineers. The Objects and Associations API that they created was based on the graph data model and was initially implemented in PHP and ran on Facebook's web servers. It represented data items as nodes (objects), and relationships between them as edges (associations). The API was an immediate success, with several high-profile features, such as likes, pages, and events implemented entirely on objects and associations, with no direct memcache or MySQL calls.


As adoption of the new API grew, several limitations of the client-side implementation became apparent. First, small incremental updates to a list of edges required invalidation of the entire item that stored the list in cache, reducing hit rate. Second, requests operating on a list of edges had to always transfer the entire list from memcache servers over to the web servers, even if the final result contained only a few edges or was empty. This wasted network bandwidth and CPU cycles. Third, cache consistency was difficult to maintain. Finally, avoiding thundering herds in a purely client-side implementation required a form of distributed coordination that was not available for memcache-backed data at the time.

All those problems could be solved directly by writing a custom distributed service designed around objects and associations. In early 2009, a team of Facebook infrastructure engineers started to work on TAO (“The Associations and Objects”). TAO has now been in production for several years. It runs on a large collection of geographically distributed server clusters. TAO serves thousands of data types and handles over a billion read requests and millions of write requests every second. Before  we take a look at its design, let’s quickly go over the graph data model and the API that TAO implements.


Design -> objects and associations, API design:

product engineers - with no direct memcache or MySQL calls.

The Objects and Associations API that they created was based on the graph data model and was initially implemented in PHP and ran on Facebook's web servers. It represented data items as nodes (objects), and relationships between them as edges (associations). The API was an immediate success, with several high-profile features, such as likes, pages, and events implemented entirely on objects and associations, with no direct memcache or MySQL calls.


likes, pages, and events implemented entirely on objects and associations, with no direct memcache or MySQL calls. 


TAO data model and API

This simple example shows a subgraph of objects and associations that is created in TAO after Alice checks in at the Golden Gate Bridge and tags Bob there, while Cathy comments on the check-in and David likes it. Every data item, such as a user, check-in, or comment, is represented by a typed object containing a dictionary of named fields. Relationships between objects, such as “liked by" or “friend of," are represented by typed edges (associations) grouped in association lists by their origin. Multiple associations may connect the same pair of objects as long as the types of all those associations are distinct. Together objects and associations form a labeled directed multigraph.

Normal description of task:

Alice checks in at the Golden Gate Bridge and tags Bob there, while Cathy comments on the check-in and David likes it.

Data items:

a user, check-in, or comment - a typed object containing a dictionary of named fields

Relationships between objects, liked by or friend of, typed edges (associations) grouped in association lists by their origin. 

multigraph - objects and associations form a labeled directed multigraph

The set of operations on objects is of the fairly common create / set-fields / get / delete variety. All objects of a given type have the same set of fields. New fields can be registered for an object type at any time and existing fields can be marked deprecated by editing that type’s schema. In most cases product engineers can change the schemas of their types without any operational work.


Associations are created and deleted as individual edges. If the association type has an inverse type defined, an inverse edge is created automatically. The API helps the data store exploit the creation-time locality of workload by requiring every association to have a special time attribute that is commonly used to represent the creation time of association. TAO uses the association time value to optimize the working set in cache and to improve hit rate.


The simplicity of TAO API helps product engineers find an optimal division of labor between application servers, data store servers, and the network connecting them.


System design: How Facebook Scale its Social Graph Store? TAO

 Here is the article. 

Social graph data is stored in MySQL and cached in Memcached

3 problems:

  1. list update operation in Memcached is inefficient. cannot append but update the whole list.
  2. clients have to manage cache
  3. Hard to offer read-after-write consistency

To solve those problems, we have 3 goals:

  • online data graph service that is efficiency at scale
  • optimize for read (its read-to-write ratio is 500:1)
    • low read latency
    • high read availability (eventual consistency)
  • timeliness of writes (read-after-write)

Data Model 

  • Objects (e.g. user, location, comment) with unique IDs
  • Associations (e.g. tagged, like, author) between two IDs
  • Both have key-value data as well as a time field

Solutions: TAO 

  1. Efficiency at scale and reduce read latency

  2. Write timeliness

    • write-through cache
    • follower/leader cache to solve thundering herd problem
    • async replication
  3. Read availability

    • Read Failover to alternate data sources




inode - index node - Unix-style file system

 The inode (index node) is a data structure in a Unix-style file system that describes a file-system object such as a file or a directory. Each inode stores the attributes and disk block locations of the object's data.[1] File-system object attributes may include metadata (times of last change,[2] access, modification), as well as owner and permission data.[3]

Storage research: A Comparison between Data Storage Solutions – SATA vs SSD & HDD

 Here is the article. 


My takeaways from the article:

  1. three main kinds of hard drives, SATA, Solid state drivers, hard disk drivers (SATA vs SSD & HDD)

Data Storage Solutions for a PC

When it comes to data storage solutions for a PC, there are three main kinds of hard drives, SATA, solid state drives, and hard disk drives. 

Hard drives are basically metal plates with a magnetic surface that store the bulk of your data. Your resume, pictures of your dog, and your favorite recipes are all stored there, assuming you haven’t put them in the cloud. 

Here we will examine the differences between SATA, SSD, and HDD drives.


SATA vs HDD and SDD

Hard disk drives (HDD) are the original hard drives. HDDs use a read/write arm on top of a spinning disc. The whole device kind of resembles a tiny, shiny vinyl record player. These drives are large, bulky, and prone to problems.

One of the most hassling problems of hard disk drives is that they tend to get “fragmented” with time. This occurs as a result of the disc becoming crowded with data that was not written on the drive sequentially. When large files are being written to a new drive, they fit nicely into the storage space. But as time goes on, and lots of files of varying sizes become written to the disk, they don’t all fit together nicely.

HDDs can still access files that are not all written in one place. It just takes more time as the drive has to spin for a little while longer. That’s why hard disk drives need to be “de-fragmented” or “de-fragged” from time to time. Files have to be sorted back together in an orderly fashion, not unlike an actual, physical filing cabinet. Solid-state drives (SSD), by comparison, do not have this problem.

SSDs store memory in flash drive where it can be easily accessed. It’s almost like random access memory (RAM) that doesn’t delete itself when the device powers off. It’s revolutionary in that for the first time, hard disk storage can now be lightning fast while fitting onto an even smaller device. That’s why laptops have gotten so much smaller and lighter in recent years. There’s no longer any need for the bulky hard disks that once resembled cross-cut sections of a cinderblock.

SATA vs HDD & HDD vs SSD

The third option for hard drives is a SATA drive. SATA drives are less expensive and more common than SSDs. However, SATA drives are also slower to boot up and slower in retrieving data than SSDs. If you’re looking for a hard drive with tons of storage space, a SATA drive may be for you, as they commonly hold terabytes of data. But take note of the fact that because SATA drives have moving parts, they are more likely to malfunction.

HDDs are similar to SATA drives in terms of the functionality. Comparing SSDs to HDDs is similar to comparing SATAs to SSDs. Let’s look at the differences in terms of reliability, speed, and lifespan.

SSD vs HDD Reliability

SSDs are more reliable than HDDs. Because SSDs don’t have moving parts (hence the term, “solid state”), there’s a lot less that can go wrong in terms of malfunctioning. And while the lifespan might generally be shorter than an HDD, solid-state drives win the battle of SSD vs HDD reliability hands down.

SSD vs HDD Speed

Without question, SSD drives are faster. Files can be written and read without the need for a spinning disc. It’s like the difference between a two-wheeled scooter that you have to push and one with an electric motor. There’s not much more to be said about SSD vs HDD speed. Systems using SSD drives feel snappier due to their ability to quickly retrieve files. HDDs just don’t work the same way.

SSD vs HDD Lifespan

By now, you may be thinking that SSDs are far superior to other types of hard drives. And in the short term, this may be true. But when it comes to SSD vs HDD lifespan, another picture arises.

SSDs work by forcing electrons through a gate in order to change their state. This creates wear and tear on the cell, gradually reducing its performance until the drive gives out. So, while HDDs become bogged down with heavy storage and need to be defragmented, they tend to last longer if you plan on using the same hardware for a number of years.

Which solution is right for my needs?

All in all, HP hard drives, HP SSDs, and HP SATA drives are all quality pieces of hardware. HP is one of the most recognizable names in tech hardware, and provides solutions for a variety of industries, including government offices.

The best option for your needs depends on what you need to get out of your drive. HDDs are the best options for large amounts of storage with a greater lifespan, while SSDs offer greater flexibility and speed. Most users opt to have both options for their work station by storing files that they’ll need for years on a HDD, and using SSDs for files that have to move between devices.



XFS file system:

 Here is the link. 



Storage research: Finding a Needle in Haystack: Facebook's Photo Storage

 Here is the link. 

NFS based design

Typical webstie

  • Small working set
  • Infrequent access of ...
Metadata bottleneck
    Each image stored as a file
    Large metadata size severely limits the metadata hit ratio

Image read performance
10 lops / image read (large directories - thousands of files)
2 iops  / image read (smaller directories - hundreds of files_
2.5 ..

Haystack based design 

Haystack store

Storage - 12 * 1TB SATA, RAID6
Filesystem
- single - 10TB xfs filesystem

Haystack
  Log structured, append only object store containing needles as object abstractions
 100 haystacks per node each 100GB in size

Haystack store - haystack file layout 
Haystack store - haystack index file layout
Haystack store - photo server
Accepts HTTP requests and translates them to corresponding Haystack operations
Builds and maintains an inore index of all images in the Haystack
32 bytes per photo (8 bytes per image vs. ~600 ybtes per inode)
5GB index

Read operation
Read 
Lookup offset/ size of the image in the incore index
Read data (-1 top)

Multiwrite(Modify)
Async append images one by one to haystack file 
Flush haystack file
Asyc append index records to the index file
Flush index file if too many dirty index records
Update incore index

Delete 
Lookup offset of the image in the incore index
Synchronously mark image as  "Deleted" in the needle header
Update incore index

Compaction
...

Haystack Directory
Logical to physical volume mapping 
URL generation 
http-> CDN->Cache->Node->Logical volume id, image id

Load balancing 
writes across logical volumes

Photo upload 
Photo download - 

conclusion 

Haystack - simple and effective storage system
 - optimized for random reads (~1 Q/o per object read)
- cheap commodity storage
8500 LOC (C++)
2 engineers 4 months from inception to initial deployment 


Slides from the talk. Here is the link. 


Storage 
 – 12x 1TB SATA, RAID6 

Filesystem 
– Single ~10TB xfs filesystem 

Haystack 
– Log structured, append only object store containing needles as object abstractions 
– 100 haystacks per node each 100GB in size





【码农】为什么我会离开amazon | 接下来何去何从 疫情对找工作的影响 | 程序员职业生涯规划

 Here is the link. 

Amazon - alexa - SDE II - Challenge to become SDE III - youtube channel - how to plan change of job -> Make a plan -> how to make plans? 

It is tough for me to work on Leetcode algorithms, and also work on system design

Youtube.com -> update one week to two weeks on youtube.com

Google/ Facebook both has two offers -> salary negotiations 

Comfortable zoom - two months to prepare for interview -> ......

Pros and cons: 

tradeoff -

pros - salary - more profit -> culture -> more if you change jobs - this is fact

- learn more from new team, new company. More challenging jobs and learn things 

- company culture is different, 

- 2.5 year experience

cons - 

domain knowledge - people know you or you know others - build trust from other people 

soft skills - it takes half year to one year to build up skills

Hack/ GraphQL, hack - modified PHP

Do not change the job very often, invest time to learn new environment, and other things ...

Long term - what is five year career plan 

Comparison 

Comparison of internal change or move to new company 

System design - new project at current job - system design is so important for ...

Saturday, August 8, 2020

CNN business: What if I am the product manage of CNN business?

August 8, 2020

I like to write 10 ideas to prepare my project to earn 10% using my $70,000 dollar cash in my TFSA account. There are so many ideas to fit into this project, but I like to work on basics first. 

CNN business

I like to spend at least a few hours to get familiar with CNN business. I like to go over in detail all product features on the page. 

I like to get ideas how to place a short term or medium term investment, and get 10% gains in a week. The project is to help me get familiar all finance news products, and compare each other. 

I have not spend more than two hours on CNN business website. It is such great tool for me as a beginner. 


CNBC finance: What if I am the product manager of CNBC finance

August 8, 2020

Introduction

As an investor, I like to spend time to read finance news; I do think that it is possible to catch those US national finance so that I can bet on a sector or a few stocks to gain 10% in less than a week. How to work on this project? 

More reading

I have to spend time to start to read top 10 finance news website. And I should think like product managers. How to read and understand news quickly? I need to be able to make a good short term and also long term decision to bet on stocks. 

Here is the link. 



The 10 Best Finance Sites to Help You Stay on Top of the Market

 Here is the article. 

  1. CNN markets
  2. Kiplinger
  3. This is money
  4. TheStreet
  5. MarketWatch
  6. Seeking Alpha
  7. Bloomberg Markets
  8. Forbes Money
  9. DealBook
  10. MyMoney




Tianpan.co: Read performance

Here is the article. 

 To optimize the read performance, denormalization is introduced by adding redundant data or by grouping data. These four categories of NoSQL are here to help.

Key-value Store 

The abstraction of a KV store is a giant hashtable/hashmap/dictionary.

The main reason we want to use a key-value cache is to reduce latency for accessing active data. Achieve an O(1) read/write performance on a fast and expensive media (like memory or SSD), instead of a traditional O(logn) read/write on a slow and cheap media (typically hard drive).

There are three major factors to consider when we design the cache.

  1. Pattern: How to cache? is it read-through/write-through/write-around/write-back/cache-aside?
  2. Placement: Where to place the cache? client side/distinct layer/server side?
  3. Replacement: When to expire/replace the data? LRU/LFU/ARC?

Out-of-box choices: Redis/Memcache? Redis supports data persistence while memcache does not. Riak, Berkeley DB, HamsterDB, Amazon Dynamo, Project Voldemort, etc.

Document Store 

The abstraction of a document store is like a KV store, but documents, like XML, JSON, BSON, and so on, are stored in the value part of the pair.

The main reason we want to use a document store is for flexibility and performance. Flexibility is obtained by schemaless document, and performance is improved by breaking 3NF. Startup’s business requirements are changing from time to time. Flexible schema empowers them to move fast.

Out-of-box choices: MongoDB, CouchDB, Terrastore, OrientDB, RavenDB, etc.

Column-oriented Store 

The abstraction of a column-oriented store is like a giant nested map: ColumnFamily<RowKey, Columns<Name, Value, Timestamp>>.

The main reason we want to use a column-oriented store is that it is distributed, highly-available, and optimized for write.

Out-of-box choices: Cassandra, HBase, Hypertable, Amazon SimpleDB, etc.

Graph Database 

As the name indicates, this database’s abstraction is a graph. It allows us to store entities and the relationships between them.

If we use a relational database to store the graph, adding/removing relationships may involve schema changes and data movement, which is not the case when using a graph database. On the other hand, when we create tables in a relational database for the graph, we model based on the traversal we want; if the traversal changes, the data will have to change.

Out-of-box choices: Neo4J, Infinitegraph, OrientDB, FlockDB, etc.


Tianpan.co: Crack the System Design Interview

 Here is the page. 

My study notes:

RPC vs RESTful style 

Generally speaking, RPC is internally used by many tech companies for performance issues, but it is rather hard to debug and not flexible. So for public APIs, we tend to use HTTP APIs, and are usually following the RESTful style.

  • REST (Representational state transfer of resources)
    • Best practice of HTTP API to interact with resources.
    • URL only decides the location. Headers (Accept and Content-Type, etc.) decide the representation. HTTP methods(GET/POST/PUT/DELETE) decide the state transfer.
    • minimize the coupling between client and server (a huge number of HTTP infras on various clients, data-marshalling).
    • stateless and scaling out.
    • service partitioning feasible.
    • used for public API.
Db Proxy 

What if we want to eliminate single point of failure? What if the dataset is too large for one single machine to hold? For MySQL, the answer is to use a DB proxy to distribute data, either by clustering or by sharding.

Clustering is a decentralized solution. Everything is automatic. Data is distributed, moved, rebalanced automatically. Nodes gossip with each other, (though it may cause group isolation).

Sharding is a centralized solution. If we get rid of properties of clustering that we don’t like, sharding is what we get. Data is distributed manually and does not move. Nodes are not aware of each other.




BitTiger: Hadoop内部原理:分布式系统如何实现?存储、计算和调度

 Here is the link. 

0:06 Distributed System; 10:01 HFS Read; 16:51 HFS Write; 26:20 Namespace (INode Tree); 31:34 FSImage; 32:53 Secondary Namenode; 35:25 Federation; 37:39 High Availability; 44:38 Map Reduce 1.0; 46:19 Yarn 52:52 Submit a Job in Yarn 56:27 ResourceManager 58:35 Spark on Yarn 1:04:19 Reference and Closing

BitTiger: Zookeeper:分布式系统入门到实战

 Here is the link. 

Basic paxos - protocol - consistency related 

Paxos is a family of protocols for solving consensus in a network of unreliable processors (that is, processors that may fail). Consensus is the process of agreeing on one result among a group of participants. This problem becomes difficult when the participants or their communication medium may experience failures.[1

ZAB - algorithm 

Raft algorithm



BitTiger: NewSQL, 从0到1,如何设计一个分布式数据库

 Here is the link. 



System design: 4 Architecture Issues When Scaling Web Applications: Bottlenecks, Database, CPU, IO

 Here is the article. 

It is the first time I like to try to see if blogger is a very good product for me to read the article. What I do is to paste the content into the editor, and then break down words, and reorganize them, make them more customized for my special need. 

Now it is 1:14 PM. 

Lets start by defining few terms to create common understanding and vocabulary. Later on I will go through different issues that pop-up while scaling web application like

  • Architecture bottlenecks
  • Scaling Database
  • CPU Bound Application
  • IO Bound Application

Determining optimal thread pool size of an web application will be covered in next blog.

Performance

Term performance of web application is used to mean several things. Most developers are primarily concerned with are response time and scalability.

  •  Response Time

    Is the time taken by web application to process request and return response. Applications should respond to requests (response time) within acceptable duration. If application is taking beyond the acceptable time, it is said to be non-performing or degraded.

  •  Scalability

    The web application is said to be scalable if by adding more hardware, application can linearly take more requests than before. Two ways of adding more hardware are

    • Scaling Up (vertical scaling) :– increasing the number CPUs or adding faster CPUs on a single box.
    • Scaling Out (horizontal scaling) :– increasing the number of boxes.

Scaling Up Vs Scaling Out

Scaling out is considered more important as commodity hardware is cheaper compared to cost of special configuration hardware (super computer). But increasing the number of requests that an application can handle on a single commodity hardware box is also important. An application is said to be performing well if it can handle more requests with-out degrading response time by just adding more resources.

Response Time Vs Scalability 

Response time and Scalability don’t always go together i.e. application might have acceptable response times but can not handle more than certain number of requests or application is handle increasing number of requests but has poor or long response times. We have strike a balance between scalability and response time to get good performance of the application.

Capacity Planning 

Capacity planning is an exercise of figuring out the required hardware to handle expected load in production. Usually it involves figuring out performance of application with fewer boxes and based on performance per box projecting it. Finally verifying it with load/performance tests.

Scalable Architecture

Application architecture is scalable if each layer in multi layered architecture is scalable (scale out). For example :– As shown in diagram below we should be able linearly scale by adding additional box in Application Layer or Database Layer.

Scaling Load Balancer

Load balancers can be scaled out by point DNS to multiple IP addresses and using DNS Round Robin for IP address lookup. Other option is to front another load balancer which distributes load to next level load balancers.

Adding multiple Load balancers is rare as a single box running nginx or HAProxy can handle more than 20K concurrent connections per box compared to web application boxes which can handle few thousand concurrent requests. So a single load balancer box can handle several web application boxes.

Scaling Database

Scaling database is one of the most common issues faced. Adding business logic (stored procedure, functions) in database layer brings in additional overhead and complexity.

RDBMS

RDBMS database can be scaled by having master-slave mode with read/writes on master database and only reads on slave databases. Master-Slave provides limited scaling of reads beyond which developers has to split the database into multiple databases.

NoSQL

CAP theorem has shown that is not possible to get Consistency, Availability and Partition tolerance simultaneously. NoSql databases usually compromise on consistency to get high availability and partition.

Splitting Database

Database can be split vertically (Partitioning) or horizontally (Sharding).

  • Vertically splitting (Partitioning) :– Database can be split into multiple loosely coupled sub-databases based of domain concepts. Eg:– Customer database, Product Database etc. Another way to split database is by moving few columns of an entity to one database and few other columns to another database. Eg:– Customer database, Customer contact Info database, Customer Orders database etc.

  • Horizontally splitting (Sharding) :– Database can be horizontally split into multiple database based on some discrete attribute. Eg:– American Customers database, European Customers database.

- moving a few columns of an entity to one database and few other columns to another database
- some discrete attribute - what is meaning discrete here?

Transiting from single database to multiple database using partitioning or sharding is a challenging task.

Architecture Bottlenecks

Scaling bottlenecks are formed due to two issues

  • Centralised component A component in application architecture which can not be scaled out adds an upper limit on number of requests that entire architecture or request pipeline can handle.

  • High latency component A slow component in request pipeline puts lower limit on the response time of the application. Usual solution to fix this issue is to make high latency components into background jobs or executing them asynchronously with queuing.

High latency component - high latency components into background jobs or executing them asynchronously with queuing. 

CPU Bound Application

An application is said to be CPU bound if application throughput is limited by its CPU. By increasing CPU speed application response time can be reduced.

Few scenarios where applications could be CPU Bound

  • Applications which are computing or processing data with out performing IO operations. (Finance or Trading Applications)
  • Applications which use cache heavily and don’t perform any IO operations
  • Applications which are asynchronous (i.e. Non Blocking), don’t wait on external resources. (Reactive Pattern Applications, NodeJS application)

In the above scenarios application is already working in efficiently but in few instances applications with badly written or inefficient code which perform unnecessary heavy calculations or looping on every request tend to show high CPU usage. By profiling application it is easy to figure out the inefficiencies and fix them.

These issues can be fixed by

  • Caching precomputing values
  • Performing the computation in separate background job.

IO Bound Application

An application is said to be IO bound if application throughput is limited by its IO or network operations and increasing CPU speed does not bring down application response times. Most applications are IO bound due to the CRUD operation in most applications Performance tuning or scaling IO bound applications is a difficult job due to its dependency on other systems downstream.

Few scenarios where applications could be IO Bound

  • Applications which are depended on database and perform CRUD operations
  • Applications which consume drown stream web services for performing its operations

Next blog will cover how to determining optimal thread pool size of an web application.




Friday, August 7, 2020

RBA.TO: Earning day - 13.88% - missing investment opportunity with my $70,000 dollars cash in TFSA account

 August 7, 2020

Introduction

As an investor and also a beginner, I should look into more carefully about research. One thing I can do as a software programmer is to prepare an Excel sheet for all Canadian company above 2 billion, and then memorize them. Today I miss the chance to make 13.88% using $70,000 dollars cash. 

Earning day

This is my favorite company, and I met people working there a few times on job fairs before in the city of Vancouver. 

Here is the article. 

Tips to share

Go to fenviz.com, and go to Screener, choose Country: Canada, and then Earnings Date: Previous 5 Day, and then the list of companies are the following:




Stock research: How to do a good research as a beginner?

 August 7, 2020

Introduction

I like to push myself to work on stock research, specially I like to work on finviz.com. What I like to do is to go over top 10 gains stocks, and then spend five minutes on each stock, ask a few questions for each stock. 

August 7, 2020


CNDT - Conduent incorporated, 82.74%

GRPN - Groupoin, Inc., 56.66%

OPGN - OpGen, Inc., 52.49%

HBP Huttig buidling products, Inc 52.12%

KZIA - Kazia Therapeutics Limited 

SG

ADMS

ANPC

SSTI

TRUR




Scalable Web Architectures: Common Patterns and Approaches

 Here is the link. 


WDC stock: Bear of the Day: Western Digital (WDC)

 Here is the article. 

I’m talking about Zacks Rank #5 (Strong Sell) Western Digital (WDC). Western Digital Corporation develops, manufactures, and sells data storage devices and solutions worldwide. It offers client devices, including hard disk drives (HDDs) and solid state drives (SSDs) for computing devices, such as desktop and notebook personal computers (PCs), security surveillance systems, gaming consoles, and set top boxes; flash-based embedded storage products for mobile phones, tablets, notebook PCs, and other portable and wearable devices, as well as automotive, Internet of Things, industrial, and connected home applications; flash-based memory wafers; and embedded storage solutions and flash products, such as multi-chip package solutions. 

Not only is WDC a Zacks Rank #5 (Strong Sell), but the Computer – Storage Devices industry ranks in the Bottom 3% of our Zacks Industry Rank. Western Digital has seen earnings estimate revisions to the downside after their last earnings report beat on EPS but missed on revenue. The company guided Q1 EPS at 45 to 65 cents per share, versus previous expectations calling for $1.20. This has caused analysts to cut expectations, bringing next year’s Zacks Consensus Expectation down from $5.35 to $4.97.

Investors looking for other stocks within the same industry don’t have many strong earnings stories to chose from. Currently, there are a few stocks which are Zacks Rank #3 (Hold) stocks. Those include Netlist (NLST) and Seagate (STX).




Thursday, August 6, 2020

Air Canada stock: Air Canada stock - waiting for August 27 US border reopen for non-essential travel

 

Enbl stock: Lessons I like to learn as a beginner

 August 6, 2020

Introduction

It is important for me to learn how to understand quarterly earning report from ENBL company. I did not expect that there is earning day, and also it is in different format compared to other oil companies. 


My story

I rushed to sell it on August 5, and it was the earning day of ENBL stock. I sold all 800 shares, since I saw there was big gain. 

After two hours, I noticed that there are another 10% gain; I decided to look up more, and then I learned that it was an earning date. 

The lesson is to learn to be patient first, and learn how to measure upper trend percentage as well. 



Is Caesars Making the Same Mistake It Made in 2007?

Here is the article. 


Stock research: Three steps to make $3000 dollars profit in three business day

August 6, 2020

Introduction

It takes a lot of time and efforts for me to push myself to learn how to invest on stock market. To make it simple, I try to write down as a beginner, how to take a few months, learn and read and then find the opportunity. 

Three steps

I like to write down three steps:
1. First step, I chased high price, and then invested over 13000 dollars on June 5, 2020. Highest price, I lost $3000 dollars. I understood that I did not know the market. I just came out $20,000 dollars loss as a passive investor. I did not really learn how to purchase stocks. 

2. Second step, I started an investment wechat group; over months, I learned to get up early, read articles, and understand the business, quarterly reports, try to learn from them good habits to study the market.

3. Make mistakes, buy high and sell low. I understood that my emotion is getting on the way. I have to take some risk, bet on US stock market, not big crash immediately. Just a normal swing. 


Is Caesars Making the Same Mistake It Made in 2007?

Here is the article. 


Business study: Irreplaceable assets on the Las Vegas Strip

Here is the article. 

Portfolio

Data:

60 properties
51K rooms
4M gaming sq ft
71K slots 
4K tables
300 F&B outlets


The 3 Largest Gambling Stocks in 2020

Here is the article.