Thursday, September 9, 2021

A Simple Framework For Mobile System Design Interviews | 20 minutes study

Sept. 9, 2021

Here is the article. 

Alex Lementuev

Disclaimer

Disclaimer


Server-side components:

  • Backend
    Represents the whole server-sider infrastructure. Most likely, your interviewer won’t be interested in discussing it.
  • Push Provider
    Represents the Mobile Push Provider infrastructure. Receives push payloads from the Backend and delivers them to clients.
  • CDN (Content Delivery Network)
    Responsible for delivering static content to clients.

Client-side components:

  • API Service
    Abstracts client-server communications from the rest of the system.
  • Persistence
    A single source of truth. The data your system receives gets persisted on the disk first and then propagated to other components.
  • Repository
    A mediator component between API Service and Persistence.
  • Tweet Feed Flow
    Represents a set of components responsible for displaying an infinite scrollable list of tweets.
  • Tweet Details Flow
    Represents a set of components responsible for displaying a single tweet’s details.
  • DI Graph
    Dependency injection graph.
  • Image Loader
    Responsible for loading and caching static images. Usually represented by a 3rd-party library.
  • Coordinator
    Organizes flow logic between Tweet Feed and Tweet Details components. Helps decoupling components of the system from each other.
  • App Module
    An executable part of the system which “glues” components together.

Providing the “signal”

  • The candidate can present the “big picture” without overloading it with unnecessary implementation details.
  • The candidate can identify the major building blocks of the system and how they communicate with each other.
  • The candidate has app modularity in mind and is capable of thinking in the scope of the entire team and not limiting themselves as a single contributor (this might be more important for senior candidates).

Deep Dive: Tweet Feed Flow

  • Architecture patterns
    MVP, MVVM, MVI, etc. MVC is considered a poor choice these days. It’s better to select a well-known pattern since it makes it easier to onboard new hires (compared to some less known home-grown approaches).
  • Pagination
    Essential for infinite scroll functionality. For more details see Pagination.
  • Dependency injection
    Helps to build an isolated and testable module.
  • Image Loading
    Low-res vs full-res image loading, scrolling performance, etc.

Add diagram here

Components

  • Feed API Service — abstracts Twitter Feed API client: provides the functionality for requesting paginated data from the backend. Injected via DI-graph.
  • Feed Persistence — abstract cached paginated data storage. Injected via DI-graph.
  • Remote Mediator — triggers fetching the next/prev page of data. Redirects the newly fetched paged response into a persistence layer.
  • Feed Repository — consolidates remote and cached responses into a Pager object through Remote Mediator.
  • Pager — trigger data fetching from the Remote Mediator and exposes an observable stream of paged data to UI.
  • “Tweet Like” and “Tweet Details” use cases — provide delegated implementation for “Like” and “Show Details” operations. Injected via DI-graph.
  • Image Loader — abstracts image loading from the image loading library. Injected via DI-graph.

Providing the “signal”

  • The candidate is familiar with the most common MVx patterns.
  • The candidate achieves a clear separation between business logic and UI.
  • The candidate is familiar with dependency injection methods.
  • The candidate is capable of designing self-contained isolated modules.

API Design

Real-time notifications

Push Notifications

  • easier to implement compared to a dedicated service.
  • can wake the app in the background.
  • not 100% reliable.
  • may take up to a minute to arrive.
  • relies on a 3rd-party service.

HTTP-polling

Short HTTP-polling

  • simple and not as expensive (if the time between requests is long).
  • no need to keep a persistent connection.
  • the notification can be delayed for as long as the polling time interval.
  • additional overhead due to TLS Handshake and HTTP-headers

Long HTTP-polling

  • instant notification (no additional delay).
  • more complex and requires more server-side resources.
  • keeps a persistent connection until the server replies.

Server-Sent Events

  • real-time traffic using a single connection.
  • keeps a persistent connection.

Web-Sockets

  • can transmit both binary and text data.
  • more complex to set up compared to Polling/SSE.
  • keeps a persistent connection.

Protocols

REST

  • easy to learn, understand, and implement.
  • easy to cache using a built-in HTTP caching mechanism.
  • loose coupling between client and server.
  • less efficient on mobile platforms since every request requires a separate physical connection.
  • schemaless — it’s hard to check data validity on the client.
  • stateless — needs extra functionality to maintain a session.
  • additional overhead — every request contains contextual metadata and headers.

GraphQL

  • schema-based typed queries — clients can verify data integrity and format.
  • highly customizable — clients can request specific data and reduce the amount of HTTP traffic.
  • bi-directional communication with GraphQL Subscriptions (WebSocket based).
  • more complex backend implementation.
  • “leaky-abstraction” — clients become tightly coupled to the backend.
  • the performance of a query is bound to the performance of the slowest service on the backend (in case the response data is federated between multiple services).

WebSocket

  • real-time bi-directional communication.
  • provides both text-based and binary traffic.
  • requires maintaining an active connection — might have poor performance on unstable cellular networks.
  • schemaless — it’s hard to check data validity on the client.
  • the number of active connections on a single server is limited to 65k.

gRPC

  • lightweight binary messages (much smaller compared to text-based protocols).
  • schema-based — built-in code generation with Protobuf.
  • provides support of event-driven architecture: server-side streaming, client-side streaming, and bi-directional streaming
  • multiple parallel requests.
  • limited browser support.
  • non-human-readable format.
  • steeper learning curve.

Pagination

Offset Pagination

  • easiest to implement — the request parameters can be passed directly to a SQL query.
  • stateless on the server.
  • bad performance on large offset values (the database needs to skip offset rows before returning the paginated result).
  • inconsistent when adding new rows into the database (Page Drift).

Keyset Pagination

  • translates easily into a SQL query.
  • good performance with large datasets.
  • stateless on the server.
  • “leaky abstraction” — the pagination mechanism becomes aware of the underlying database storage.
  • only works on fields with a natural ordering (timestamps, etc).

Cursor/Seek Pagination

  • decouples pagination from SQL database.
  • consistent ordering when new items are inserted.
  • more complex backend implementation.
  • does not work well if items get deleted (ids might become invalid).
GET /v1/feed?after_id=p1234xzy&limit=20
Authorization: Bearer <token>
{
"data": {
"items": [
{
"id": "t123",
"author_id": "a123",
"title": "Title",
"description": "Description",
"likes": 12345,
"comments": 10,
"attachments": {
"media": [
{
"image_url": "https://static.cdn.com/image1234.webp",
"thumb_url": "https://static.cdn.com/thumb1234.webp"
},
...
]
},
"created_at": "2021-05-25T17:59:59.000Z"
},
...
]
},
"cursor": {
"count": 20,
"next_id": "p1235xzy",
"prev_id": null
}
}

Authentication

Providing the “signal”

  • The candidate is aware of the challenges related to poor network conditions and expensive traffic.
  • The candidate is familiar with the most common protocols for unidirectional and bi-directional communication.
  • The candidate is familiar with REST-full API design.
  • The candidate is familiar with authentication and security best practices.
  • The candidate is familiar with network error handling and rate-limiting.

Conclusion

Things you can control

  • Your attitude — always be friendly no matter how the interview goes. Don’t be pushy and don’t argue with the interviewer — this might provide a bad “signal”.
  • Your preparation — the better your preparation is, the bigger the chance of a positive outcome. Practice mock design interviews with your peers (you can find people on Teamblind).
  • Your knowledge — the more knowledge you have, the better your chances are.
  • Gain more experience.
  • Study popular open-source projects: iOS, Android
  • Read development blogs from tech companies
  • Your resume — make sure to list all your accomplishments with measurable impact.

Things you cannot control

  • Your interviewer’s attitude — they might have a bad day or simply dislike you.
  • Your competition — sometimes there’s simply a better candidate.
  • The hiring committee — would make a decision based on the interviewers’ reports and your resume.

Judging the outcome


Book reading: Cassandra definitive guide Eben Hewitt | First 60 minutes study

Sept. 9, 2021

Introduction

It is time for me to read a book called Cassandra definitive guide. Recently I spent a lot of time to read papers, articles and lecture notes about BigTable, and it should be very important for me to read a book called Cassandra definitive guide. 

First 60 minutes study

I like to read a few pages first, so that I can understand how long it will take me to finish the book. 

Cloud girl: The battle of relational and non-relational databases | SQL vs NoSQL Explained | My notes in one page

Sept. 9, 2021

Here is the link. 

Which database is right for your application? SQL or NoSQL? Are you confused between relational and non relational databases? Want to know how they SQL and NoSQL databases are different? In this video I explain the difference and help you decide which database to use in which type of application. I am Priyanka Vergadia, Developer Advocate for Google Cloud, for more content 📌 Follow me on Twitter - https://twitter.com/pvergadia 📌 Follow me on Instagram - https://www.instagram.com/pvergadia/ 📌 Follow me on LinkedIn - https://www.linkedin.com/in/pvergadia 📌 GCPSketchnote Playlist - https://bit.ly/3jA8Ylz 📌 Visit my website to download the sketchnote image - https://thecloudgirl.dev/ #databases #nosql #sql #relationaldatabases #nonrelationaldatabases #db #dbms #rdbms


Type of databases

  • SQL - relational
  • Non-relational

SQL - Relational

Using a database with three tables: User, Product, Promo, structed to explain ACID. 
OLTP
Atomic - A 
Consistent - C
Isolated - I
Durable 

Google cloud, for example, cloud SQL, Cloud spanner 
grow vertically  

NoSQL - Non-Relational


Unstructured 
Large column  - BigTable
document - 
key value - BT
Graph - Janus + BigTable 
In memory - MemoryStore

Large data sizes
BASE - explain BASE in detail
Basically Available
Soft state
Eventually Consistency 

Horizontal scaling 

Actionable Item


Review the blog by looking up blogs using BASE protocol, and here is the blog to document detail how to understand BASE, using persistent message queue to support high availability and reduce coupling. 

Google cloud: Your Google Cloud database options, explained | 20 minutes study

Sept. 9, 2021

Here is the link. 

Relational databases 

In relational databases information is stored in tables, rows and columns, which typically works best for structured data. As a result they are used for applications in which the structure of the data does not change often. SQL (Structured Query Language) is used when interacting with most relational databases. They offer ACID consistency mode for the data, which means:

Because of these properties, relational databases are used in applications that require high accuracy and for transactional queries such as financial and retail transactions. For example: In banking when a customer makes a funds transfer request, you want to make sure the transaction is possible and it actually happens on the most up-to-date account balance, in this case an error or resubmit request is likely fine.

There are three relational database options in Google Cloud: Cloud SQL, Cloud Spanner, and Bare Metal Solution.

Cloud SQL: Provides managed MySQL, PostgreSQL and SQL Server databases on Google Cloud. It reduces maintenance cost and automates database provisioning, storage capacity management, back ups, and out-of-the-box high availability and disaster recovery/failover. For these reasons it is best for general-purpose web frameworks, CRM, ERP, SaaS and e-commerce applications.

  •  Atomic: All operations in a transaction succeed or the operation is rolled back.
  •  Consistent: On the completion of a transaction, the database is structurally sound.
  •   Isolated: Transactions do not contend with one another. Contentious access to data is moderated by the database so that transactions appear to run sequentially.
  • Durable: The results of applying a transaction are permanent, even in the presence of failures.


Cloud Spanner: Cloud Spanner is an enterprise-grade, globally-distributed, and strongly-consistent database that offers up to 99.999% availability, built specifically to combine the benefits of relational database structure with non-relational horizontal scale. It is a unique database that combines ACID transactions, SQL queries, and relational structure with the scalability that you typically associate with non-relational or NoSQL databases. As a result, Spanner is best used for applications such as gaming, payment solutions, global financial ledgers, retail banking and inventory management that require ability to scale limitlessly with strong-consistency and high-availability. 

Bare Metal Solution: Provides hardware to run specialized workloads with low latency on Google Cloud. This is specifically useful if there is an Oracle database that you want to lift and shift into Google Cloud. This enables data center retirements and paves a path to modernize legacy applications. 

Non-relational databases

Non-relational databases (or NoSQL databases) store complex, unstructured data in a non-tabular form such as documents. Non-relational databases are often used when large quantities of complex and diverse data need to be organized, or where the structure of the data is regularly evolving to meet new business requirements. Unlike relational databases, they perform faster because a query doesn’t have to access several tables to deliver an answer, making them ideal for storing data that may change frequently or for applications that handle many different kinds of data. For example, an apparel store might have a database in which shirts have their own document containing all of their information, including size, brand, and color with room for adding more parameters later such as sleeve size, collars, and so on.

Qualities that make NoSQL databases fast:

  • Typically, they are optimized for a specific workload pattern (i.e., key-value, graph, wide-column)
  •  Horizontal scaling, usually using range or hashed distributions
  • Eventual consistency: many NoSQL stores usually exhibit consistency at some later point (e.g., lazily at read time). However, Firestore uniquely offers strong global consistency.
  • Transactions:  a majority of NoSQL stores don't support cross shard transactions or flexible isolation modes. However, Firestore uniquely offers ACID transactions across shards with serializable isolation.

Because of these properties, non-relational databases are used in applications that require large scale, reliability, availability, and frequent data changes.They can easily scale horizontally by adding more servers, unlike some relational databases, which scale vertically by increasing the machine size as the data grows. Although, some relational databases such as Cloud Spanner support scale-out and strict consistency.

Non-relational databases can store a variety of unstructured data such as documents, key-value, graphs, wide columns, and more. Here are your non-relational database options in Google Cloud: 

  • Document databases: Store information as documents (in formats such as JSON and XML). For example: Firestore
  • Key-value stores: Group associated data in collections with records that are identified with unique keys for easy retrieval. Key-value stores have just enough structure to mirror the value of relational databases while still preserving the benefits of NoSQL. For example: Bigtable, Memorystore
  • In-memory database: Purpose-built database that relies primarily on memory for data storage. These are designed to attain minimal response time by eliminating the need to access disks. They are ideal for applications that require microsecond response times and can have large spikes in traffic. For example: Memorystore
  • Wide-column databases: Use the tabular format but allow a wide variance in how data is named and formatted in each row, even in the same table. They have some basic structure while preserving a lot of flexibility. For example: Bigtable
  • Graph databases: Use graph structures to define the relationships between stored data points; useful for identifying patterns in unstructured and semi-structured information. For example: JanusGraph

There are three non-relational databases in Google Cloud:

  •  Firestore: Is a serverless document database which scales on demand, is strongly-consistent, offers up to 99.999% availability and acts as a backend-as-a-service. It is DBaaS that is optimized for building applications. It is perfect for all general purpose uses cases such as ecommerce, gaming, IoT and real time dashboards. With Firestore, users can interact with and collaborate on live and offline data making it great for real-time application and mobile apps.  
  • Cloud Bigtable: Cloud Bigtable is a sparsely populated table that can scale to billions of         rows and thousands of columns, enabling you to store terabytes or even petabytes of data. It is ideal for storing very large amounts of single-keyed data with very low latency. It supports high read and write throughput at sub-millisecond latency, and it is an ideal data source for MapReduce operations. It also supports the open-source HBase API standard to easily integrate with the Apache ecosystem including HBase, Beam, Hadoop and Spark along with Google Cloud ecosystem.
  • Memorystore: Memorystore is a fully managed in-memory data store service for Redis and Memcached at Google Cloud. It is best for in-memory and transient data stores and automates the complex tasks of provisioning, replication, failover, and patching so you can spend more time coding. Because it offers extremely low latency and high performance, Memorystore is great for web and mobile, gaming, leaderboard, social, chat, and news feed applications. 

Conclusion

Choosing a relational or a non-relational database largely depends on the use case. Broadly, if your data structure is not going to change much, select a relational database. In Google Cloud use Cloud SQL for any general-purpose SQL database and Cloud Spanner for large-scale globally scalable, strongly consistent use cases. In general, if your data structure may change later and if scale and availability is a bigger requirement then a non-relational database is a preferable choice.  Google Cloud offers Firestore, Memorystore, and Cloud Bigtable to support a variety of use cases across the document, key-value, and wide column database spectrum. For more comparison resources on each database check out the overview. For more hands-on experience with Bigtable, check out our on-demand training here and learn about migrating databases to managed services check out this whitepaper.  


Here is the video to show the content as well. 

I like to take some notes here as well. 








What is Cloud Pub/Sub? - ep. 2

Sept. 9, 2021

Here is the link. 







Cloud Pub/Sub Overview - ep. 1

Sept. 9, 2021

Here is the link. 

In this first episode of Pub/Sub Made Easy, we help you get started by giving an overview of Cloud Pub/Sub. Learn how to ingest large amounts of data for analysis, simplify the development of event-driven microservices, and much more! Overview → https://goo.gle/31T8AEU Cloud Pub/Sub documentation → https://goo.gle/2BVmLhP Pub/Sub Made Easy playlist → https://goo.gle/2XQ3vwC Subscribe to the GCP Channel → https://goo.gle/GCP

My notes


Add Google pub sub service in-between all services. 




Sanas aims to convert one accent to another in real time for smoother customer service calls

Breakfast and learn: What is Cloud Pub/Sub? - ep. 2

Breakfast and learn: Bigtable in action (Google Cloud Next '17)

Wednesday, September 8, 2021

Marketing and Creative Insights from Unstructured Data: Cloud ML APIs (Cloud Next '19)

Sept. 8, 2021

Here is the link. 

How can businesses use Google Cloud to unlock unstructured data for their marketing and creative analytics? In this session we present two customer examples that use Google Cloud APIs to create actionable signals from unstructured data. Leveraging the Cloud Natural Language API, we show how a big CPG brand built a scalable data pipeline to analyze written product reviews through entity and sentiment analysis. Using the Cloud Vision API, we share the story of a creative agency that built an automated creative pipeline to gather insights at scale by analyzing visual attribute data from images.

Creative analysis at scale with Google Cloud and ML → https://bit.ly/2U1NhwY Watch more: Next '19 Databases Sessions here → https://bit.ly/Next19Databases Next ‘19 All Sessions playlist → https://bit.ly/Next19AllSessions Subscribe to the GCP Channel → https://bit.ly/GCloudPlatform Speaker(s): Sona Oakley, Sjef van Stiphout, James Rappazzo, Tim Booher

Session ID: DA223 event: Google Cloud Next 2019; re_ty: Publish; product: Cloud - Data Analytics - BigQuery, Cloud - Data Analytics - Google Marketing Platform; fullname: Sona Oakley, Sjef van Stiphout, James Rappazzo, Tim Booher;

From Blobs to Tables: Where and How to Store Your Stuff (Cloud Next '19)

Sept. 8, 2021

Here is the link. 

Google Cloud Platform offers many options for storing your data. From storage to databases to data warehousing, data is the foundation of any enterprise and any application. In this session, we’ll talk about what options exist, the strengths and trade-offs of each choice for various workloads, and demo the different experiences across services when storing your first bytes.

Data Storage with Google Cloud → http://bit.ly/2UeqfaZ Watch more: Next ‘19 All Sessions playlist → https://bit.ly/Next19AllSessions Subscribe to the Google Cloud Channel → https://bit.ly/GoogleCloud1 Speaker(s): Gabriela Ferrara, Dave Nettleton, Tobias Ternstrom Session ID: SPTL201 product:Cloud Spanner,Cloud Storage,Cloud SQL; fullname:Gabriela Ferrara,Dave Nettleton,Tobias Ternstrom;