Friday, May 14, 2021

System design: How to Program with MongoDB Using the .NET Driver? | First time - 2 hours | Learn basic concepts.

May 14, 2021

Here is the link. 

It is hard for me to learn technology over last decade. I like to read the following technical article in less than two hours. 

  1. I will install MongoDB at my home office computer. 

How to Program with MongoDB Using the .NET Driver

Oct.. 14, 2018

As more shops move to NoSQL databases, developers must learn new techniques for querying and updating data. In this article, Darko Martinović demonstrates how to work with MongoDB in .NET.

Many people that have a background in relational databases are confused with the terms NoSQL and the document database. What kind of documents are in the database and how to get data without the query language, without SQL? In my opinion, the term NoSQL does not mean that there is no schema, rather than the schema is not strictly enforced. And, of course, there is a query language as well.

During the past few years, JSON has become extremely popular. Getting data from various forms (i.e., WEB and WIN forms) became extremely easy using JSON. Furthermore, saving such data as users entered them dictates the shape of the data. The shape of the data is determined by the application itself, unlike a relational database in which the data are independent of the application. In NoSQL databases, the data are saved in the JSON document. The table record in the relational world is equivalent to the JSON document in the NoSQL world.

I think that it is easier to learn something new by comparing with something that you already know. This article will try to be a gentle introduction to the NoSQL world and will explain how to transform part of a well-known database, AdventureWorks2016, to NoSQL as shown in the image below.

For this article, I had to choose between MongoDB and Cosmos DB. Both databases provide challenges and opportunities, but, in my opinion, MongoDB has an advantage. It provides more free options than Cosmos DB. Furthermore, MongoDB is a multiplatform database with an on-premise edition that is much richer than the Cosmos DB Emulator.

Let’s get started. The first step is to set up the environment. This means installing the MongoDB database, installing the NET driver, and importing some data.

Setting up the environment

At the time of this writing, October 2018, the current version of MongoDB in 4.0.2, and the current version of the .NET driver is 2.7. To install MongoDB, start with a Google search for ‘Mongo DB download’ or click on the link. From the list of available platforms, choose Windows x64. Then fill in the form with your first and last name, e-mail address, and so on. After that, save the MSI installation file locally.

The installation is straightforward. MongoDB will be installed as a Windows service by default. You have the option of choosing the service name, the service user, and startup folders for data and log files, as shown in the image below.


If you choose the Custom installation option, you may specify an installation directory. The MongoDB installation will put the executables in C:\Program Files\MongoDB\Server\4.0\bin folder by default. Be sure to add this path to your PATH environment variable to be able to run the MongoDB shell from any folder in the command prompt.

Feel free to explore executables installed in the folder. The most important is mongod the windows service itself, mongo the shell and mongoimport, a utility that helps import various data in the database. You should also install the MongoDB GUI explorer called Compass from the link.

The next step in setting the environment is to install the NET driver. To do that, start Visual Studio. I’m using the VS 2017 Community Edition in this article.

Create a new console project using C#, and name it as you wish. The purpose of creating a new project is to show you how to reference the MongoDB .NET driver. From the project context menu, choose Manage NuGet Packages. After NuGet Package Manager appears, enter MongoDB driver in the Browse tab as shown in the image below.

Choose and install MongoDB.Driver, and the three main components for programming MongoDB by using .NET are installed. All three are published together, as shown by version number, in the image below.

MongoDB.Bson handles various types and-file formats as well as the mapping between CLR types to BSon. BSon is MongoDB representation of JSON. I will write much more about BSon later.

MongoDB.Driver.Core handles connection pooling, communication between clients and database and so on. Usually, there is no need to work with this library directly.

MongoDB.Driver is where the main API is located.

Once you download the three main libraries, you can start any new project and copy references from this first project. At this point, you can save and close this first project.

In order to follow the article, please, download the article solution from the GitHub. Open the solution by starting another instance of Visual Studio. In the solution notice the folder Data as shown in the image below.

The folder contains three JSON files that must be imported by using a MongoDB command line utility called mongoimport. Determine where on your local disk those three files are located. Start the command prompt from that location. In the command prompt window, execute the following three commands.

By executing these commands, you will create a MongoDB database named simpleTalk (camel case naming) and three collections named: adventureWorks2016, specialOffer, and product as shown in the image below from Compass. When starting Compass, it tries to establish a connection on localhost port 27017. Just press Connect and, on the left side, select the database simpleTalk.

The collection in NoSQL is similar to the table in the SQL world and provides a namespace for the document. To prove that the import was successful, start the mongo shell from the command prompt. In the command prompt, type mongo to enter into the shell.

Once you enter the shell, type use simpleTalk to switch in the context of the database. And then type show collections and you will be able to see all three imported collections.

To conclude this first section, notice the Authentication folder in the article’s solution. Inside the folder, there is a JavaScript file named AddUser.js. If you are not in the mongo shell, start it again. In the context of the simpleTalk database, execute the code of AddUser.js, as shown in the snippet below, to create a user.

All examples in the article will execute in the context of this newly created user usrSimpleTalk. The user has been granted read and write permissions to the database.

Now it’s time to talk about how to connect to the MongoDB database, how CLR types are mapped to BSon types, and the root objects in the MongoDB .NET driver.

Connecting to MongoDB

The purpose of this section is to provide information on how to connect to the MongoDB database and to examine the most important objects of MongoDB API for .NET. Those objects are the MongoDB client, the Mongo database, and the Mongo collection. Open the article solution and make sure that the startup object is Auth as shown in the image below.

The class named Auth demonstrates how to connect to the MongoDB database. MongoClient is the root object that provides that connectivity.

There are several various methods to connect to the database. The first way is to pass the database name, the username, and the password, as shown in the snippet below.

As you may have noticed by browsing the code, there is no need for opening, closing and disposing of connections. The client object is cheap, and the connection is automatically disposed of, unlike ADO.NET in which the connection is a very expensive resource. This approach is shown in the solution. In the solution, there is one configuration object located in the Configuration folder and named SampleConfig.

The MongoClient has many overloads. The most common way to instantiate the client is to pass a connection string, as shown in the snippet below.

By default, MongoDB API uses port 27017 to communicate with the database, but there are many more options. That includes connecting to multiple servers, the replica sets and so on.

At this point, I have to make a digression. One common thing that I would like to know is the number of currently opened connection. By using the mongo shell, it is possible to see that number but without further information.

If you execute in the mongo shell command db.serverStatus().connections, you will see the response in the form of a document, as shown in the image below.

The only way to get more information is to use netstat –b in the cmd window with elevated permissions and getting the result as displayed on the image below.

Let’s get back to the main topic. Down in the object hierarchy is the database object. To get a reference to the database, usually, you execute the following snippet.

The database is accessed in the context of the client. If you explore the exposed method in the client, you will notice that there are no options for creating a database. That is because the database is created automatically when it is needed and if it doesn’t exist.

In the context of the database, you can get a reference to the collection by executing snippet like shown below.

or like this one

In the context of the database, you can create, drop, or rename a collection. There is a method to create a collection, but there is no need to use it because the collection will be created automatically if does not exist. The collection and the database object are both thread safe and could be instantiated once and used globally.

As you noticed in the snippets above, the only difference in getting a reference to a collection object is by passing a type. This is the type needed in .NET to work with the collection. One is a BsonDocument that represents a dynamic schema in which you can use any document shape, and one is a so-called strongly typed schema named SalesHeader. As you will discover in the article, SalesHeader is the class that mimics the table in the AdventureWorks2016 database named Sales.SalesHeader. The option that is strongly typed is the generally preferred way when working with MongoDB in .NET.

The document object is found lower in the object hierarchy. The most general way to represent the document is to use the BSonDocument class. The BsonDocument is basically a dictionary of keys (strings) and BsonValues. BsonValues will be examined in the next section. To conclude this section, F5 to start the article’s solution. The result will display on the console screen. It represents the number of documents in the adventureWorks2016 collection, and it is determined by calling the collections method CountDocumentsAsync.

Mapping CLR types to BSON types

Change the startup object in order to follow this section. This time set Attribute Decoration located in the Serialization folder as the startup object. In the section, I will refer to a couple of simple classes which definitions can be found in the Pocos folder POCO is an acronym for ‘plain old CLR object’ and represents a class that is unencumbered by inheritance.

The first example uses two objects. The first one is of type TestSerializerWithOutAttributes, the second one of type TestSerializer. Both classes define the same properties of type bool, string, int, double and datetime. The classes’ definitions can be found in TestSerializer.cs located in the Pocos folder. The only difference between these two classes is that the second class has been decorated with attributes. In the example, I instantiate two objects and perform the basic serialization by using ToJson, as shown in the image below.

By pressing F5, the result will be displayed as shown in the image below.

As you will notice, there are a couple of differences. One of them is that decimal type, by default, is displayed as a string. It must be decorated decimal type by attribute [BsonRepresentation(BsonType.Decimal128)], as is done in the definition of class TestSerializer.

Then, if you noticed, the property OnlyDate is set by using the following snippet

Therefore, there is no time part. I am located in the UTC +1h time zone, and because it’s currently daylight savings time, the default serializer reduces the date value by two hours. To avoid such behavior, I decorated, the property with the attribute [BsonDateTimeOptions(DateOnly =true)].

In the example, I am using some other attributes, which I’ll describe. If you would like to change the element name or specify a different order, try decorating the property with BsonElement as shown in the snippet below.

As you’ll see later, every BsonDocument that is part of a collection should have an element named _id. This is a kind of primary key in the NoSQL-Bson world. Also, by default, the collection is indexed by using this field. You are free to specify your own primary key by decorating a property or field in the class definition with the attribute BsonId, as shown in the snippet below.

Finally, in order to specify that only a significant number of digits will be used when working with the double type, the following attribute is used:

You probably noticed that I use the word the default serializer, although in the code there is no call to any kind of serializer. This is because of the beautiful .ToJson extension. As you can see by using the Visual Studio peek definition or by pressing ALT + F12, ToJson is an extension of the object type defined in the MongoDB.Bson namespace. So, it should be safely used on any type.

There is one thing about ToJson I have to write about at this spot. As you probably noticed, the extension optionally receives a parameter of type JsonWriterSettings.

If you examine this class by using the Visual Studio peek definition, you will notice that the class has properties defined as shown in the image below.

You can pay special attention to the property outputMode, as shown surrounded by red on the image above. This is an enumerator with two possible values. The default one is JsonOutputMode.Shell and the second one is JsonOutputMode.Strict. So, what is the difference? According to the MongoDB documentation :

  • Strict mode. Strict mode representations of BSON types conform to the JSON RFC. Any JSON parser can parse these strict mode representations as key/value pairs; however, only the MongoDB internal JSON parser recognizes the type of information conveyed by the format.
  • mongo Shell mode. The MongoDB internal JSON parser and the mongo shell can parse this mode.

My experience is that the mongoimport utility does not understand shell mode, so, I had to change the default serialization behavior to be mode strict in order to generate a JSON file that could be accepted by the mongoimport utility. You can find more about the differences between strict and shell mode here.

Besides decorating classes with attributes, there is an option to use so-called ClassMap. For example, if you want to keep serialization details out of their domain classes and do not want to play with attributes, you will use the ClassMap approach instead. This is not the only scenario in which you might use ClassMap. You can combine attribute decoration with ClassMap as well.

 

To practice working with ClassMap, switch the startup object in the article’s solution to ClassMap. In this example, I’m using the same type as before, an object of type TestSerializerWithOutAttributes, to produce the same output as in the previous example.

The class should be registered only once like is shown in the snippet below.

An exception will be thrown if you try to register the same class more than once. Internally, BsonClassMap holds information about registered types in a dictionary like shown in the snippet below.

So, registering a class twice means adding a key to the dictionary that exists, which is, on the other hand, an exception. Usually, you call RegisterClassMap from some code path that is known to execute only once (the Main method, the Application_Start event handler, etc.).

The most common way when working with ClassMap is to use so-called AutoMap and after that perform some add-on coding as shown in the snippet below.

In the article’s solution, there is an example showing how to use AutoMap named ClassMapAutoMap, located in Serialization folder. However, I will not use this option in the article. For example, to specify that the decimal type should be rendered (NOTE: I use term rendered and serialized interchangeably) as a decimal rather than strings, following code is used. Notice that there is a predefined serializer for a decimal type.

In order to specify the element name and change the order that the element is rendered, the following snippet is used:

To serialize the DateTime type with only the date part, or to use Local Time, the following snippet is used. Notice that there is predefined serializer for the datetime type.

Similarly, to specify that SalesOrderId should be treated as a BsonId, the following snippet is used

Finally, to specify only a significant part of the digits to be serialized when working with the double type, the following code snippet is used

Also, by pressing F5, you should get the same result as in the previous example.

That is not the end of the possibilities. There is an option to use the so-called Convention Pack. In short, to be able to follow the section from this point, change the startup object to be TestConventionPack.

When working with ClassMap and ‘automapping,’ many decisions should be made. What property should be BsonId, how should the decimal type be serialized, what will be the element name, and so on?

Answers to these questions are represented by a set of conventions. For each convention, there is a default convention that is the most likely one you will be using, but you can override individual conventions and/or write your own convention.

If you want to use your own conventions that differ from the defaults, simply create an instance of ConventionPack, add in the conventions you want to use, and then register that pack (in other words, tell the default serializer when your special conventions should be used).

For example, to instantiate an object of the type of ConventionPack, you might use the following snippet:

The first convention, CamelCaseElementNameConvention, will tell default serializer to put all elements name in CamelCase. This is a predefined convention. All other conventions are defined in the example and represent the custom conventions. DecimalRepresentationConvention will tell the default serializer to serialize all decimal properties as a decimal, rather than a string, which is the default option. Similar to this is the DateOnlyRepresentation and LocalDateReporesentation. When working with ClassMap, you should connect your class with convention pack. This is usually accomplished like is shown in the code snippet:

In the Serialization folder, there is a class named TestTypes as well. Please, change the startup object to be TestTypes. This example shows how complex types are transformed into JSON(BSON). This includes .NET native types like generic collections, as well as classes that inherit from other classes. When a class inherits from other class, a special field _t, that represents the type, is rendered as shown in the image below, surrounded with red.

As a take away from this section, it’s possible to decorate the class with attributes, work with ClassMap, and, finally, work with ConventionPack.

API is very easy to use and very hard to misuse when working with serialization. Now it is time to talk about collections. How are they designed? The next section is about schema design.

Schema design

One of the documents in the adventureWorks2016 collection looks as shown in the image below.

As you’ll notice, a detail array that represents details of an order is embedded in the document. In the SQL world, details of an order are put in a separate table known as Sales.SalesDetail. You could do the same thing in MongoDB, e.g., put details in a separate collection, but as you may recall from the introduction, in the NoSQL world, the shape of the data is determined by the application itself. There’s a good chance that when you are working with the sales data, you probably need sales details. The decision about what to put in the document is pretty much determined by how the data is used by the application. The data that is used together as sales documents is a good candidate to be pre-joined or embedded.

One of the limitations of this approach is the size of the document. It should be a maximum of 16 MB.

Another approach is to split data between multiple collections which is also used in the article solution. For example, details about products and special offers are separated into another collection. One of the limitations of this approach is that there is no constraint in MongoDB, so there are no foreign key constraints as well. The database does not guarantee consistency of the data. Is it up to you as a programmer to take care that your data has no orphans.

Data from multiple collections could be joined by applying the lookup operator, as I’ll show in the section that talks about aggregations. But, a collection is a separate file on disk, so seeking on multiple collections means seeking from multiple files, and that is, as you are probably guessing, slow. Generally speaking, embedded data is the preferable approach.

The underlying CLR class to work with the adventureWorks2016 collection is SalesHeader. Its definition is located in the Pocos solution folder.

In the class definition, the detail is represented as shown in the following snippet.

It is a generic List of SalesDetail. The SalesDetail class mimics the Sales.SalesDetail table. The CLR class to work with the product collection is Product and, to work with the spetialOffer collection, a class SpetialOffer is designed. The source code that shows you how these collections are generated is located in the Loaders folder.

The image below displays the content of the spetailOffer collection.

The spetialOffer collection will be used in the next section, which talks about C(reate), R(ead), U(pdate) & D(elete) operations.

CRUD Operations

To follow this section, change the startup object of the article’s solution to CrudDemo. This example demonstrates how to add, modify and, finally, delete a couple of documents in the spetialOffer collection.

As you may recall from the previous section, the spetailOffer collection has IDs from 1 to 16. If you try to add an ID that already exists, a run-time exception will occur. For example, you can define an object of type SpetialOffer, as shown in the following snippet:

Running this will return an exception with the message shown in the snippet. Finally, try inserting a document that has an ID of 20.

In order to get the document, the Find extension of IMongoCollection is used as shown in the following snippet.

To insert many documents into a collection, you have to pass an enumerable collection of the document to the InsertMany method. InsertMany is an extension of IMongoCollection. To insert in the collection documents with ID 30 and 31, you could execute the code snippet shown below.

One interesting thing to notice is the second parameter of InsertManyAsync. It is an object of type InsertManyOptions. Particularly, its property IsOrdered is interesting. When set to false, the insertion process will continue on error.

There are two kinds of updates. There is a replace extension which replaces the entire document, and there is an update extension that updates just the particular field or fields in the document. When replacing a document, first you have to specify a filter function to find the document(s) to be replaced. It is not possible to change ID during replacement. If you try to do something like that an exception will be thrown. If you specify a condition that does not match any documents, the default behavior is to do nothing.

Usually, if you want to replace one document, a snippet like following is used:

ReplaceOneAsync is an extension of the IMongoCollection and represents a high order function which takes as a parameter another function – lambda expression. One of the parameters that is not provided in the above snippet is UpdateOptions. In my opinion, a better name should be ReplaceOptions because it’s in the context of Replacing. Particularly, in that class, a property IsUpsert is interesting. When specified to be true, an insert is made if the filter condition did not match any document.

 

When updating the document, usually you will execute a snippet like that shown below

In the above snippet, I’m using an example with Builders. Builders are classes that make it easier to use the CRUD API. Builders help define an update condition. This time, only part of a document has been changed. Similar to another CRUD extension of IMongoCollection is the Delete extension. For example, the following snippet will delete the document with ID 20.

 

There are more extensions like FindOneAndUpdate, FindOneAndReplace, FindOneAndDelete, and so on.

There is no transaction in MongoDB, but there is the so-called atomic operation. Any write operation on the particular document is guaranteed to be atomic – not breakable. Starting with MongoDB 4.0 there is a transaction, but they are limited only for replica sets. (NOTE: A replica set in MongoDB is a group of mongod processes that maintain the same data set ).

In the article’s solution, there is an example that uses a transaction, named InsertOneWithTransaction. It is commented, but in short, in order to use the transaction, a session object should be created in the context of the MongoDB client. Then the session object starts the transaction, and the session object is passed to CRUD methods (extensions) as shown in the snippet below.

All examples in this section use the async stack and TPL library (task parallel library). There is also a sync stack as well, but the first one should be considered as a modern way of programming and was introduced with driver version 2.X.

In this section, I just briefly mentioned how to read and filter data. The next section will talk more about how to filter data.

Filtering

To follow this section, set up FindWithCursor as a startup object for the article solution. In this example, I query documents where

  • TerritoryId equal to 1,
  • SalesPersonId equal to 283,
  • Total Due greater than 90000
  • and limiting the number of documents to be returned to 5,
  • Sorting the result ascending by Due Date.

The task is accomplished by utilizing Find, an extension of IMongoCollection. Find is defined as shown in the image below.

It returns an IFindFluent, which is a context type for method chaining in searching documents. Other methods like Sort, Limit, Project in context of IMongoCollection return IFindFluent.

Find takes two parameters, a lambda expression and an object of type FindOptions. Besides other properties, FindOptions defines the batch size. By using Find you could get a result in chunks. If you limit the number of the document to 5, a total of 3 batches is returned to the client. The complete code is shown in the image below.

Getting the next batch is accomplished by invoking cursor.Result.MoveNextAsync. Inside that batch, you can iterate through the documents by processing cursor.Result.Current. The benefit of that approach is that if you get a large number of documents as a result, you can process them in batches, which will use less memory.

There is an option to get all results by invoking ToListAsync(). In that case, all returned documents live in memory, and there is an option to process the cursor using ForEachAsync as shown in the image below.

Invoking cursor.ToString()will return a MongoDB shell query. I did not find a proper way to get the query plan in the code. It was possible before the 2.7 release of the .NET driver. In the article’s solution, there is an example of how to get the query plan from the code. The example is named Explain and is located in the Filtering folder.

To get the execution plan, you could save the query text and execute in the context of the Mongo shell. I saved the query in a file QueryUsingCursor.js.

So, if you append explain() before find, you will be able to see the execution plan.

Explain receives a parameter that describes what the type of output should be. The parameter specifies the verbosity mode for the explain output. The mode affects the behavior of explain() and determines the amount of information to return. The possible modes are queryPlanner, executionStats, and allPlansExecution.

The example uses executionStats. The plan is shown in the image below.

As you can see, to return five documents, you have to process all the documents in the collection. That is the part when the index comes to play. To create an index, you should execute a command shown in the image below, in the context of the simpleTalk database.

The rule in index creation requires a field that participates in filtering first and then fields that are included in sorting.

An index could be created in the foreground which is the default. What does it mean? MongoDB documentation states: ‘By default, creating an index on a populated collection blocks all other operations on a database. When building an index on a populated collection, the database that holds the collection is unavailable for reading or write operations until the index build completes. Any operation that requires a read or writes lock on all databases will wait for the foreground index build to complete’. This does not sound good.

For potentially long-running index building operations on standalone deployments, the background option should be used. In that case, the MongoDB database remains available during the index building operation.

To create an index in the background, the following snippet should be used. There is no need to create the index again. This index is small, and it’s creation takes a few seconds.

In MongoDB, there is no need to rebuild indexes. The database itself takes care of everything. Great!

Let’s get back to the main topic about the query plan. After the index is created, examine the execution plan again.

As you can see, highlighted with yellow, the execution plan looks much better now. Only five documents are examined.

Another example located in the Filter folder named FilterHeader uses the MongoDB aggregation framework to query the document, which is the next section. To explore this example, change the startup object to be FilterHeader, and you will receive the output as shown in the image below.

Aggregation

Aggregation operations process some input data, in the case of MongoDB, documents and return computed results. Aggregation operations group values from multiple documents together and can perform a variety of operations on the grouped data to return a single result. MongoDB provides three ways to perform aggregation:

I will write mostly about the first by introducing a common problem in the SQL world that is called TOP n per group. The single purpose aggregation methods are briefly mentioned in the first example that connects to MongoDB when Count was introduced. Map-reduce was the only way to aggregate in prior versions of MongoDB.

In MongoDB, there is a difference when querying embedded documents such as sales details, or the main document. In the first case, the unwind operator must be introduced. So, I decided to include the same problem twice.

The first example finds the top N (one) customer per territory that has the greatest Total Due and then limits the result to those territories and customers where the sum of Total Due is greater than the Limit (a defined number). The result should be sorted on Sum of Total Due descending order. The second example does something similar with special offer and products.

There are a couple of ways to accomplish this task in SQL. In the article’s solution, I include two T-SQL scripts, located in TSQL folder. One for querying the header table named QueryingHeader.sql, and one for querying the detail table named QueryingDetail.sql.

I provided three possible ways to solve the problem in T-SQL, by using T-SQL window functions and the APPLY operator.

Both results for querying the header and the detail table, are displayed in the image below. In the first case, the limit is 950.000, and in the second case, the limit is 200.000.

To see how is the problem solved in MongoDB, switch the startup object of the article’s solution to be AggregationSales.

In the class, there are three ways how to accomplish the same task.

  • Using IAggregateFluent
  • Using LINQ
  • Parsing BsonDocument ( the MongoDB’s shell-like syntax )

The first way is shown in the snippet below.

First, you have to notice that everything in the snippet above is strongly typed!

Then, you should notice that Aggregate is an extension of the IMongoCollection. Aggregate returns a type of IAggregateFluent<TDocument>. All other operators like SortBy, Group, Match, Project, etc. do the same thing. They are extensions of IAggregateFluent and return IAggregateFluent<TDocument>. This is the how methods chaining is accomplished which is really one of the characteristics of the fluent API. So, again, I will repeat the API is easy to use and difficult to misuse, and every method name is self-documenting.

Take a look at IAggregateFluent, in the image below. Surrounded with red is a read-only property named Stages.

It represents every operation performed on the IMongoCollection. If you set the breakpoint on the line

And try to examine the query variable, you will notice that there are six stages defined as shown in the image below.

Every stage has the type of IPipelineStageDefinition and has the Input and Output type as well as the shell operator name as shown in the image below. The .NET driver is perfectly mapped to MongoDB shell!

To capture the whole query, execute query.toString();. The query is saved in the article’s solution in a file Query.js, located in Aggregation folder.

The second way to accomplish the same task is to use LINQ like syntax. To do that, the MongoDB.Driver.Linq namespace must be included. This time the query looks similar like is shown in the snippet below.

As you see this time, the AsQueryable extension is used. This is an extension of the IMongoCollection and returns an instance of IMongoQueryable. Every method also returns IMongoQueryable, and this is the how method chaining is accomplished, this time with LINQ style syntax.

Method names are similar to IAggregateFluent. Instead of Group, there is GroupBy; instead of Sort; there is OrderBy; instead of Project, there is Select; instead of Match, there is Where, etc.

You can also use SQL like syntax to accomplish the same task, as shown in the snippet below.

There is a method to pass BsonDocument to the pipeline as well. In the solution the method name is UsingMongoShellLikeSyntax.

The result from all three methods is displayed on the console output and shown in the image below.

 

At this point, I have to make two digressions. First, I briefly describe the lookup operator and then the unwind operator.

As you recall when I write about schema design, the preferable way to work with data is embedded documents, but, there is a situation when you have to join data from a different collection. In such a situation, the lookup operator will help you.

Switch the solution’s startup object to be SampleLookup located in the Aggregation folder. This example creates two collections, one that contains names and the other that contains the meaning of the names. Of course, the only reason I am doing like this is to show how the lookup operator works.

If you use the peek definition of Visual Studio, it will show what Lookup to expect, as shown in the image below, surrounded with red.

The first Lookup is an extension of the IMongoCollection. It executes in the context of the collection and requires a foreign key collection as the first parameter, then the local field on which relation is established, the foreign field that composes the relation, and finally the result type. Unfortunately, the result type cannot be an anonymous type (or I did not discover how to be anonymous?). As always is expected that from the foreign collection will be returned more than one element, so, the operator always expects an array to be returned.

In this case that looks like the code shown in the snippet below.

And the result is displayed on the console window as shown in the image below.

The MongoDB documentation states following for unwind: Deconstructs an array field from the input documents to output a document for each element. Each output document is the input document with the value of the array field replaced by the element.

That means if you want to aggregate on embedded documents, you have to promote them to the top level. Here’s an example. If you would like to aggregate sales details, an array of documents embedded in the main document, you have to use unwind. The example of the basic usage of unwind operator could be found in the SampleUnwind solution file.

If you take just one document in the adventureWorks2016 collection by applying Limit(1) and after that apply Unwind on the Details field, the output will produce as many rows as the number of embedded documents that are in the Details fields. The source code to accomplish such a task is displayed in the snippet below.

And finally, make AggregationUnwind the startup object of the article’s solution. There are five tests in this class. Two of them are accomplished by using LINQ, two of them by using IAggregateFluent and one by parsing BsonDocument as shown in the image below. The result is always the same and equal what is returned by executing the T-SQL script in SSMS.

I found that using LINQ style coding is extremely easy. Unwind in LINQ style syntax is simple SelectMany for joining, as shown in the image below highlighted with yellow.

Using IAggregateFluent is a little bit difficult, especially when working with Unwind. Unwind has three different overloads. One is obsolete, one uses BsonDocument as output, and that means breaks strongly typed writing, and the last one is a little bit confusing.

I expected that the result of the UnWind operation should be an anonymous type. Unfortunately, it should be a concrete type of class that is defined. It is shown surrounded with red on the image below.

And SalesDetailHelper is defined as shown in the image below.

Pretty simple, but on the other hand still a little bit annoying. It would be nice if you could use an anonymous type.

The aggregation framework has limitations. The stages have a limit of 100 MB of RAM, per stage. If a stage exceeds the limit, MongoDB produces an error.

Funny thing about this limitation is that I was unable to find how much memory is used per stage. Thanks to, the moderator in MongoDB forums, I finally found out that there are no possibilities to get that information. It will be available in the next release. As a drawback for such a situation, you could pass an optional parameter allowDiskUse to true to enable writing data to temporary files, as in the following example:

Another limitation is the ability to use indexes. I was trying to avoid the collection scan when using the group operator. Unfortunately using indexes with the group operator does not work. The operator always performs the scan. Furthermore, if in the aggregation pipeline, any stage cannot benefit from index usage, all stages that follow also will not use indexes.

Besides that, everything seems to be excellent.

Summary

There is no doubt, MongoDB is a great product. Although I still search for some features that exist in the SQL world, my general impression is excellent. I can say the same for the .NET driver. It is easy to use, every method (extension) emphasizes the intent of the code, enables you to write less code, and everything is where you expect to be.

Although the article moves at the speed of light, I hope that the article with article’s solution would be enough to encourage readers to start exploring MongoDB and the .NET driver.

May 14, 2021 three hours walking | 7 KM walk from office to home | my favorite art

 








Thursday, May 13, 2021

Stock investment: How to avoid gambling?

May 13, 2021

I like to figure out more about stock investment. How to tell if I invest on a stock or gamble on a stock? 

一个赌博的故事告诉你:股市能赚钱的只有一种人

为什么把股市当成赌局你就输了

赌博的每个人都会输钱

以旁观者眼光看,按照概率来讲,赌钱输赢几率应该是一半,为什么只要看到赌徒都是身负巨额债务,只见输不见赢、输是输在一个 贪 字N多戒赌吧或者论坛有句至理名言:赢只是个过程,输是结果。很多赌徒把数钱原因归结为黑、假、骗、只能说还没有吃够教训、总觉得赌局公平就能赢。事实上只要有数学常识就能明白,无论赌局公平与否,最终结果永远是输光的、 庄家:金钱有限、时间有限、你心态不行怎么跟我玩。

一个很简单的数学原理,甚至连数学原理都称不上。你一块钱我一块钱,赌一局,你我输赢概率各50%。你一块钱我十块钱,我赢光你签的几率比你大得多。你一块钱,我本钱无限、你无论如何都不可能赢我的、参与赌博的人,只要你一直在赌,赌场桌上的钱相当于你就是等于无限,赌局结束只有一个结果——你输光你身上的全部。

明白了吗?这里的关键在哪儿?在“一直在赌”这四个字。你参加的赌局越多,你离输光就越近。

为什么赌徒会一直深陷,越输越多?

为什么明明输了钱,不能吸取教训、非要再一次把钱丢进去,甚至借钱丢进去?被火烧了知道不再靠近火,被小偷偷了知道要看紧钱包,为什么赌钱输了,还要一而再再而三的输?

首先,是因为刺激。赌博带来的生理上的兴奋感,输赢只见巨大的心理落差造成的大喜大悲给身体带来的刺激,如同过山车和高空跳伞,是别的娱乐活动很难提供的、对这种刺激感的追求,是“赌瘾”的直接来源。

当然,这只是最直接、最简单也是最原始的原因。

第二原因,是情绪。世界上百分之九十九的人无法真正控制自己的情绪,做到完全理性的人少之又少。赌桌是一个能无限放打人的负面情绪和心理弱点的地方、所谓“输急眼”是最可怕的心理因数。而人类有一个普遍的心理弱点是:人因为“失去”而产生的负面情绪远超过“获得”而产生的正面情绪。

这就造成了一个常见的现象。赌徒是难以分辨“自己赢的钱”和“本钱”的无论赢了多少钱,只要输一把,就一定要再赢回来,无数人无数次在“再赢一把就收手”的情况下,因为输了一把非要扳回来而导致无法挽回。

第三层原因,是压力。当你输了一万,你觉得没关系。几把就能翻回来了、当你输了十万,你开始有压力这时只要停止,还有挽回的余地,但你不敢面对过去十年的辛苦付诸东流的结果,不敢面对家人的责难,于是你开始举债。但残酷的数据规律会告诉你,把几十万块钱输光很容易,但想从零开始重新赢回十万块却成了几乎不可能完成的任务、你只有一次又一次的重复“借债——输光——再借债”的恶性循环,直至无法掩饰全面爆发。

这是有人就会想,输钱会有压力。会控制不住情绪,如果我运气稍好点,一开始赢了钱。或者输了些钱很快翻回来了,不就能有好的结果吗?

这就是我要说的三点,也就是最重要的一点。

输钱不可怕,赢钱才可怕!

赢钱了你会收手吗?不可能的。

每个人根深蒂估的劣根性——成功是我个人强大、失败因为环境掣肘。

翻译到赌桌上:赢钱=我是赌神。数钱=今天运气不好,明天就会转运。

这种毫无理由的自负一定会遭到概率之神的惩罚,告诉你一切都是运气,你并不比任何人聪明。

万一,概率之神眷顾了你,你赢了更多的钱呢?

真正可怕的事情可能会发生——你在金钱观充值被摧毁了。

你一年能挣多少钱?步行街均薪三十万。算你一年能挣一百万吧!你靠工作能在一个晚上挣十万吗?能在一场球的时间里挣几万块吗?但是靠赌博,你挣到了。

那么,你愿意去工作吗?如果你经历过一晚上几十万的输赢,你在精神上已经很难再回到靠工作,靠努力去挣回曾经失去的一切。

杰出的心性,才是高手的刀!

人的聪明绝不在投资取巧的技能方面,而在于脚踏实地,在于善良、勤奋、认真、守信和对正义、真理的执著追求和坚守。

你可以踏上投资的道路,但绝不会因为善于投资而成功。

投资就像攀登珠穆朗玛峰,选择它是不理智的,但是选择之后任何不理智的想法和做法都会带来灭顶之灾,拥有任何侥幸心理只能说明受到的教训还不够深刻。

当你看到一个游刃有余的高手,他的背后一定隐藏无数血淋淋的伤口,九死一生,概莫能外。

投资之道的探索是没有止境的,不存在最高境界,只有更高境界。我们只能在方向正确的前提下一步一个脚印地向前走,成功之路没有任何捷径,苦难是我们的良师。

市场犹如一个魅力无比却又凶残成性的魔鬼。她不仅能吞掉你的金钱,甚至还能吞噬人的生命。可悲的是,人在绝对意义上讲是无法战胜股市的,因为判断力和想象无法战胜事实,你的胜利与否不是由你决定而是由市场决定。我想,最终意义上的高手是那些能够作到从容进退的人。李佛摩尔之所以失败是因为他迷信自己的力量,不能拒绝股市的诱惑。换句话说,他拿得起来却放不下,在《股票作手回忆录》中可以多次看到他放弃了原定的假期急急忙忙杀回到股市的例子。其实拿得起来容易,放得下才是最难作到的。在股市成功,你必须全身心地关注它,必须付出相当的代价,但你又不能因此就放弃一切。股市没有比生活重要,你必须注意在两者之间找到平衡点,否则很容易既失去了钱,又失去了生活。

在股市竞争获胜的难度要远远大于其它行业。原因在于:1、炒股获利快速。2、工作不辛苦。3、进入门槛低。这些因素导致了股市里竞争的强度比在其它行业激烈和残酷得多。由于人性的缺陷和对竞争规律认识的不足,绝大多数人必定将成为牺牲品。我们必须正视这个严峻的现实:股市挣钱很不容易!但的确着存在着一些长期以来一直比较成功的人。

那么,成功之路到底在何方呢?

大家可以看到,股市中有中线高手、有长线高手;有人利用技术分析获得成功,有人用价值投资理念获得成功,还有人竟然利用随即漫步理论取得成功。我们自己已经投入了大量时间、金钱和感情在股海苦苦挣扎却依然屡战屡败,伤痕累累,到底是什么使他们成为那少数的幸运儿呢?我现在的观点是:方法、理论甚至理念都不是关键的诀窍,高手之所以从碌碌众生中脱颖而出,是因为他们的个人本身具备与众不同的特质,他们这些特殊的人找到了适合自己的方法、理论或理念,并善于利用时机。方法、理论甚至理念可以被人学到,但他们的特质你学不到。

高手们成功方法不尽相同,只有探询他们的心路历程,我们才能发现他们取胜的关键。股市是人性的竞争,在股市能够长期生存下来的人无疑是强者,高手们真正意义上的强大是他们的心灵,所有高手的共同特点是他们能够在心性(类似持久的精神状态)层次上建立起自己独特的核心竞争力进而在市场上持续获得竞争优势,就是说,高手都有属于他自己的心之刀。

股市的成功在于能否突破人性的藩篱,超越自我,在心性层次为自己打造出一把刚柔并济的好刀。这是所有的高手的共同点,如果你没有自己的刀,即使你花费了再多的时间研究那些高级的刀招,也只有落得等死或投降的结局。

高手的刀,其实就是杰出的心性。对我们低手来说,这需要长时间相当痛苦的积累过程才能达到,和一个人智商的高低、理论知识掌握的多寡无关。我相信,只有你到了拥有高手之心的那一天,你自然同时拥有了高手的眼界。K线虽然还是那些K线,波浪虽然还是那些波浪,但从这一刻起才会真正变得鲜活、生动起来;招数虽然还是那些招数,也只有从这一刻起才能对你有实际意义。在质变的临界点没有达到之前,即便你满腹经纶、绝顶聪明,被人性弱点蒙蔽的双眼又怎能看见遍地的黄金呢?股道精华,实在股外!

有了自己的刀,你才真正踏上了属于自己的成功路,但锻造这把刀却是任何大师和前辈都无法教你的事情,就好比想学会游泳,你必须亲自下水了解水性,精通流体力学和运动生理学对你不会有太大帮助一样!通往高手之路是一条漫长而痛苦的蜕变之路,孤独是你唯一的伙伴,失败是你最好的老师!百忍成钢,当心性修炼得有如镜子般明彻,如流水般圆韧时;当你切切实实生活在不以物喜、不以己悲的宁静中时;当你发觉胸中不断流动着“虽千万人而吾往矣”般的勇气时,历经千锤百炼,你的刀就练成了。拔刀四顾、纵横天下的那一天还会远吗?

Stock investment: 赌徒的五种阶段

May 13, 2021

I like to figure out gambling and how to identify the idea of gambling in stock investment. 

赌徒的五种阶段

第一阶段:懵懵懂懂

往往刚接触赌博的人最开始只是为了娱乐,或者纯粹的出于好奇的心理,偶尔只是小打小闹,这一阶段持续的时间很短,赢了钱,高兴一阵,就会收手,偶尔想起来了就会玩一下,不管输赢都能及时收手,对赌博谈不上兴趣。不过,慢慢的接触得多了之后,如果反复受到赌博的刺激,尝到了赌博的甜头,就会令人不自觉的增加了赌博的行为,甚至会错误的认为自己已经掌握了赌博的技巧,总幻想着不劳而获,无形中,他们赌博的次数越来越频密,赌注也会越来越大,自然而然的踏入了第二阶段。

第二阶段:贪得无厌

赌博的行为,源于内心的的贪婪、懒惰,经常抱有侥幸的心理,企图通过不劳而获来获取利益,因贪婪而痴迷,就像是一团火花,遇上干柴就会熊熊燃烧,一发不可收拾。时刻幻想着一个月能赢多少钱好过去打十年工,当成功赢得一部分后,欲望就会开始膨胀,还想着拥有更多,永无止境。这种强烈的欲望久而久之转为可怕的贪念,成为了赌徒痴迷赌博最坚不可摧的动力,然而“贪”字,到最后的结局必定是输的,为了拿回自己的失去的一切,赌徒会踏入第三阶段。

第三阶段:企图翻身

到了这个地步,你完全放不下,心里充满了不甘,越想越觉得不值,不断的加大筹码,希望把以前输掉的拿回来,因为常常花费大量的金钱用于赌博,输了之后就更加觉得不服气,日益沉迷,心中只要想着还有钱,就一定会扭转乾坤,在上岸与洗白之间轮回。于是乎债务越来越多,窟窿越来越大,为了填补这个缺口,不断的编织谎言去借钱,借遍了家人、亲人、朋友,希望能够短时间内赢回所有,不计后果,后来当谎言被揭发后,情况开始变得恶化,于是,进入了第四阶段。

第四阶段:妄想补天

面对血浓于水的家人,深爱自己的亲人,称兄道弟的朋友,心里充满了愧疚,想起那些失去的血汗钱,充满了不甘,越是愧疚,越是不甘心,越是想拿回自己失去的钱财,可是没有“赌本”了,开始小打小闹的试图“补天”,继续沉迷着赌博,可悲的是每日千日砍柴一日烧,此刻心情更加的烦躁不安,不断的研究和变换投注规律,好不容易赢回来一点点,然后又是一天输光。就这样周而复始的轮回着,到头来输的钱也越来越多了,宣告补天失败,但是还未能醒悟,还妄想着继续靠赌去解决所有问题,此时,身边已经借无可借了,开始从信用卡和各种高息网贷上套取金钱继续下去,更有甚者不惜用非法的手段铤而走险获取钱财,继续在这条道路上形影单只的走着。

第五阶段:丧心病狂

当你走到这个阶段,已经完全属于失控的状态,满脑子想的都是赌,可是一旦下了海,想上岸哪有那么简单?到了这个地步的你,走投无路,负债累累,家庭支离破碎。在绝望中,赌徒会陷入疯狂,会走上犯罪的深渊,开始危害到社会的稳定性。外面的债务,家人花费一辈子辛苦打拼来的积蓄帮你偿还了,你跪在地上声泪俱下的答应着他们说不赌了,可是一转头又忘的一干二净,直到最后,贷无可贷,亲朋好友疏远,家人不管,债务爆发,到了这个阶段的赌徒,吃一天,饿三天,不是在赌桌,就是在到处想办法筹钱的路上,直到最后甚至连别人一丝的同情都得不到,只能行尸走肉般的活着。

Walking study: Understanding Weight Loss: How to Lose 20 Pounds by Walking | Lose 20 pounds in 20 weeks

May 13, 2021

Here is the article. 

Walking is a great way to lose 20 pounds for many reasons, and knowing how to do it effectively will help you reach your goal weight in no time. Walking is enjoyable for most people, easy on your joints, and one of the safest forms of exercise. Many people find they can stick to a walking program long term which is essential for weight maintenance. The key to losing 20 pounds by walking is to set appropriate goals and understand the fundamentals of weight loss.

How Long Will it Take Me to Lose 20 Pounds?

At a weight loss rate of ½ -1 pound per week, it will likely take you at least 20 weeks to lose 20 pounds. Losing weight at this pace is safe and will help you keep the weight off long term. To accomplish a weight loss of ½ - 1 pound per week, try to burn an extra 250-500 calories per day by walking. If you find you're not burning this many calories by walking alone, simply reduce your calorie intake through diet in addition to walking.

How Often Should I Walk?

If you're a beginner, start by walking 3 days per week for at least 15-20 minutes. Gradually increase the frequency and duration of your walks until you are walking 30-60 minutes per day, most days of the week. To help keep your walks enjoyable try alternate walking indoors with walking outdoors, watching television during your walks (using a treadmill), or listening to music or a book on tape with headphones. For most people it's not walking they dislike, but becoming bored during the walk. Work walking into your regular routine and make it a priority.

How Many Calories Can I Burn By Walking?

The number of calories per minute you can burn by walking is determined by your body weight and walking pace. If you walk at a pace of 4 miles per hour (a common pace) you can burn the following amount of calories per minute: 120 lb. person = 4.7 calories; 140 lb. person = 5.5 calories; 160 lb. person = 6.3 calories; 180 lb. person = 7.1 calories; 200 lb. person = 7.8 calories; and 220 lb. person = 8.6 calories. If you plan to lose 20 pounds by walking alone, try to burn at least 250 extra calories during your walk per day. For example, if you weigh 160 pounds you'd have to walk at least 40 minutes per day at a pace of 4 miles per hour to lose ½ pound per week. If you're unsure of your pace, try walking on a treadmill to give you a better idea.

How Can I Lose Weight and Stay Toned?

Walking alone will definitely help you lose weight, however adding resistance exercise to your routine will help keep you tight and toned during your weight loss. Try walking with arm or ankle weights some days or interval train a few days per week (alternate power walking with moderately paced walks). On the days you don't walk, try lifting weights, Pilates or strength band training to stay toned while losing 20 pounds.