Sunday, October 10, 2021

Chevron CEO Warns of High Energy Prices for the Foreseeable Future

Oct. 10, 2021

Here is the link. 

Sep.16 -- The world is facing high energy prices for the foreseeable future as oil and natural gas producers resist the urge to drill again, according to Chevron CEO Mike Wirth. He spoke with Bloomberg's Alix Steel about challenges ahead for the energy sector and consumer.


Open transcript

let's start with the tough spot that oil
00:02
companies find themselves in in the
00:03
energy transition like you're going to
00:05
have the critics that are like you're
00:06
not going far enough you're not going
00:07
fast enough then you have the other guys
00:09
that are like why are you spending more
00:10
money where is my return how do you
00:12
think about that as a ceo of a major oil
00:14
company well yesterday we actually
00:16
talked to investors about what what i'd
00:18
call a winning combination
00:21
and what that is is delivering
00:24
a high return
00:26
lower carbon traditional business which
00:28
generates the cash today
00:31
but we also announced a tripling of our
00:34
capital commitment to over 10 billion
00:35
dollars
00:37
in new energies so these are faster
00:39
growing lower carbon new energies that
00:42
leverage the strengths that our company
00:44
has built up so we laid out ambitious
00:46
growth targets in renewable fuels
00:48
in hydrogen in carbon capture and
00:50
offsets these are lower carbon forms of
00:53
energy where we can create competitive
00:56
advantage for shareholders and it also
00:58
helps us
00:59
solve problems for customers in the more
01:01
difficult to electrify
01:03
sectors of the economy like aviation or
01:05
marine transportation or heavy duty
01:07
transportation
01:08
when you were working on the plan what
01:10
kind of assumptions
01:12
in terms of the oil play price in terms
01:14
of government support policy support did
01:17
you model
01:19
well we we start with what we know today
01:22
and so we have models for
01:24
population growth economic growth
01:27
oil supply demand
01:30
the the new energies are tougher because
01:32
these are newer businesses i think the
01:34
range of uncertainty is is wider on
01:36
technology on cost on rate of market
01:40
development and so the error bars are
01:43
larger but we've really focused into
01:45
geographies where we believe
01:48
the right combination exists so these
01:50
are places where we have the
01:51
capabilities the assets and the
01:53
customers to start these businesses up
01:55
california's a great example that's
01:57
where our headquarters are
01:59
and california has
02:01
carbon pricing uh policies that exist to
02:05
incentivize investment and so we're
02:07
building a carbon negative power plant
02:11
in california we're working on hydrogen
02:13
bringing hydrogen out to retail in the
02:16
transport sector and our renewable
02:18
natural gas
02:19
sustainable aviation fuel business in
02:21
california is where it will begin so the
02:23
conditions there allow us to invest in
02:25
projects that should earn good returns
02:27
and then we believe that those uh
02:29
learnings will help us extend into other
02:31
markets other geographies and policy
02:33
will evolve so so how price sensitive is
02:35
it like if oil all of a sudden goes 75
02:38
to 40. do you rethink that capex plan
02:41
we really are a long-term allocator of
02:44
capital other than extremes so with
02:46
covid where we saw negative oil prices
02:49
and real questions we pulled capital
02:51
spending down in that kind of an
02:53
environment but
02:54
by and large we take a long-term view on
02:57
these things and so commodity cycles are
02:59
part of our business and our plans need
03:01
to be
03:02
robust through a range of prices for
03:04
both the traditional energy products and
03:06
these new energy products so they are uh
03:08
price dependent to deliver returns but
03:11
they're not in the short term going to
03:13
be um
03:16
our plans won't change if we see prices
03:18
trend above or below unless we think
03:20
there's a structural change so if the
03:22
returns are depend are obviously
03:24
dependent on price
03:25
what kind of return what kind of price
03:27
did you model your returns at if you're
03:28
looking for double digit yeah modest oil
03:30
prices i mean oil prices that look uh no

Saturday, October 9, 2021

Chevron Live Tour of the Archive with John Harper, Episode 1

Oct. 9, 2021

Here is the link. 

05/04/2017 - Did you know we have an amazing archive facility that showcases our company's rich history spanning 135 years? You'll have the chance to tour the facility with our historian John Harper and his team during a four-part video series featuring the Chevron Archive. Learn more about Chevron’s history: https://www.chevron.com/about/history


Day in the Life: energy trader

 Oct. 9, 2021

Here is the link. 

Nicole Oakes, a crude supply and trading manager for Chevron’s Latin America team, shows us what life is like on the trading room floor. https://www.chevron.com/stories/i-am-...

Nicole Oakes

Here is the linkedin profile. 


Friday, October 8, 2021

Stock investment: 2020 birthday rookie bet | 2021 birthday sweet bet

Oct. 8, 2021

2020 birthday | GME stock | $30 dollar loss | One day | Panic sale | 54 year old




Highest price of GMO was $483/ share, if I do have confidence and choose to keep those 30 shares of GME, then I can make profit around $15,000 US dollars. 


2021 birthday | Review stocks | Candidates | Sweet bet | 55 year old 




Storage area network | wiki article

 A storage area network (SAN) or storage network is a computer network which provides access to consolidated, block-level data storage. SANs are primarily used to access data storage devices, such as disk arrays and tape libraries from servers so that the devices appear to the operating system as direct-attached storage. A SAN typically is a dedicated network of storage devices not accessible through the local area network (LAN).

Although a SAN provides only block-level access, file systems built on top of SANs do provide file-level access and are known as shared-disk file systems.


Storage architectures

Storage area networks (SANs) are sometimes referred to as network behind the servers[1]: 11  and historically developed out of a centralized data storage model, but with its own data network. A SAN is, at its simplest, a dedicated network for data storage. In addition to storing data, SANs allow for the automatic backup of data, and the monitoring of the storage as well as the backup process.[2]: 16–17  A SAN is a combination of hardware and software.[2]: 9  It grew out of data-centric mainframe architectures, where clients in a network can connect to several servers that store different types of data.[2]: 11  To scale storage capacities as the volumes of data grew, direct-attached storage (DAS) was developed, where disk arrays or just a bunch of disks (JBODs) were attached to servers. In this architecture, storage devices can be added to increase storage capacity. However, the server through which the storage devices are accessed is a single point of failure, and a large part of the LAN network bandwidth is used for accessing, storing and backing up data. To solve the single point of failure issue, a direct-attached shared storage architecture was implemented, where several servers could access the same storage device.[2]: 16–17 

DAS was the first network storage system and is still widely used where data storage requirements are not very high. Out of it developed the network-attached storage (NAS) architecture, where one or more dedicated file server or storage devices are made available in a LAN.[2]: 18  Therefore, the transfer of data, particularly for backup, still takes place over the existing LAN. If more than a terabyte of data was stored at any one time, LAN bandwidth became a bottleneck.[2]: 21–22  Therefore, SANs were developed, where a dedicated storage network was attached to the LAN, and terabytes of data are transferred over a dedicated high speed and bandwidth network. Within the SAN, storage devices are interconnected. Transfer of data between storage devices, such as for backup, happens behind the servers and is meant to be transparent.[2]: 22  In a NAS architecture data is transferred using the TCP and IP protocols over Ethernet. Distinct protocols were developed for SANs, such as Fibre Channel, iSCSI, Infiniband. Therefore, SANs often have their own network and storage devices, which have to be bought, installed, and configured. This makes SANs inherently more expensive than NAS architectures.[2]: 29 

BigTable > Related work > Shared nothing architecture > The Case for Shared Nothing | Ricardo Jimenez-Peris | Linkedin article

 Oct. 8,  2021

My notes:

  1. Learn three architecture options: Shared-nothing, shared disk, and shared-memory
  2. Definition of shared-nothing? Think about it in a minute. 
  3. Definition: A shared-nothing architecture is simply a distributed architecture. This means the server nodes do not share either memory or disk and have their own copy of the operating system, hence they share “nothing” other than the network to communicate with other nodes
  4. Google the following statement - SAN - storage area network, block-based protocol
  5. Shared-disk requires disks to be globally accessible by the nodes, which requires a storage area network (SAN) that uses a block-based protocol. 
  6. block-based protocol - learn three things in 20 minutes
  7. SAN - a storage area network (SAN) - in 20 minutes
  8. SAN - focus on a block-based protocol - in 10 minutes

The Case for Shared Nothing

Ricardo Jimenez-Peris

Shared-nothing has become the dominant parallel architecture for big data systems, such as MapReduce and Spark, analytics platforms, NoSQL databases and search engines [Özsu & Valduriez 2020]. The reason is simple: it is the only architecture that can provide scalability at reasonable cost, typically within a cluster of servers. In the context of cluster computing, scalability can be further characterized by the terms scale-up versus scale-out. Scale-up (also called vertical scaling) refers to adding more power (processor, memory, IO devices) to a server and thus gets limited by the maximum size of the server, e.g. 32 processors. Scale-out (also called horizontal scaling) refers to adding more servers, called “scale-out servers”, in a loosely coupled fashion, to scale almost infinitely.

In a parallel database system, the real challenge is to scale linearly with increasing workloads (see our blog post on Scalability), including more users and more data. For instance, if you double the size of your cluster, you would expect to support a workload that is twice as big. For an OLAP workload, this may mean dividing the response time of a large analytical query by two, whereas for an OLTP workload this may mean doubling the system throughput (e.g., number of transactions per minute). Note that these are very different objectives, which can be achieved with different architectures and techniques. The level of difficulty to implement basic functions (concurrency control, fault-tolerance, availability, database design and tuning, load balancing, etc.) varies from one architecture to the other.

Stonebraker proposed the term “shared-nothing” to better contrast with two other popular architectures, shared-memory and shared-disk. A shared-nothing architecture is simply a distributed architecture. This means the server nodes do not share either memory or disk and have their own copy of the operating system, hence they share “nothing” other than the network to communicate with other nodes [Valduriez 2018]. In the 1980s, shared-nothing was just emerging (with pioneers like Teradata and Tandem NonStopSQL) and shared-memory was the dominant architecture.

In shared-memory (see above), any processor (P) has access to any memory module (M) or disk unit through some interconnect. All the processors are under the control of a single operating system. One major advantage is simplicity of the programming model, which is based on shared virtual memory. Metadata (e.g. directory) and control data (e.g., lock tables) can be shared by all processors, which means that writing database software is not very different than for single processor computers. In particular, load balancing is easy since it can be achieved at runtime by allocating each new task to the processor with least load. However, the major problem with shared-memory is limited scalability and availability.

With increasingly quick processors (even with larger caches), conflicting accesses to the shared-memory increase rapidly and degrade performance. Furthermore, since the memory space is shared by all processors, a memory fault may affect most of them, thereby hurting data availability. Depending on whether physical memory is shared, two approaches are possible: Uniform Memory Access (UMA) and Non-Uniform Memory Access (NUMA) (see Figure 2). UMA is the architecture of multicore processors while NUMA is used for tightly-coupled multiprocessors. Interestingly, today’s servers are NUMA, which introduces many difficulties for database managers since they were built for UMA. In particular, seeing as all threads can access memory from an CPU, this results in a high fraction of accesses to remote NUMA memory. This results in exhausting the memory bandwidth to remote memories that is significantly smaller than the memory bandwidth of the local memory. Most database servers just do not scale up linearly in NUMA and a lot of work is needed for them to become NUMA-aware.

In a shared-disk architecture, any processor has access to any disk unit, but exclusive (nonshared) access to its main memory. Each processor–memory node is under the control of its own copy of the operating system and can communicate with other nodes through the network. Then, each node can access database pages on the shared-disk and cache them into its own memory. Shared-disk requires disks to be globally accessible by the nodes, which requires a storage area network (SAN) that uses a block-based protocol. Since different processors can access the same page in conflicting update modes, global cache consistency is needed. This is typically achieved using a distributed lock manager, which is complex to implement and introduces significant contention. Shared-disk has three main advantages: simple and cheap administration, high availability, and good load balance. Database administrators do not need to deal with complex data partitioning, and the failure of a node only affects its cached data, while the data on disk is still available to the other nodes. Furthermore, load balancing is easy because any request can be processed by any node. The main disadvantages are cost (because of the cost of the SAN) and limited scalability, caused by the potential bottleneck and overhead of cache coherence protocols for large databases. In the case of OLTP workloads, shared-disk has remained the preferred option as it makes load balancing easy and efficient. However, their scalability is heavily limited by the distributed locking that causes severe contention and limits the scalability to a few nodes.

Today, a shared-nothing architecture (see Figure 4) is cost-effective as servers can be off-the-shelf components (multicore processor, memory and disk) connected by a regular Ethernet network or a low-latency network such as Infiniband or Myrinet. To maximize performance (yet at additional cost), nodes could also be NUMA multiprocessors, thus leading to a hierarchical parallel database system [Bouganim et al. 1996]. By favoring the smooth incremental growth of the system by the addition of new nodes, shared-nothing provides excellent scalability, unlike shared-memory and shared-disk. However, it requires careful partitioning of the data on multiple disk nodes. Furthermore, the addition of new nodes in the system presumably requires reorganizing and repartitioning of the database to deal with the load balancing issues. Finally, fault-tolerance is more difficult than with shared-disk, seeing as a failed node will make its data on disk unavailable, thus requiring data replication. It is due to its scalability advantage that shared-nothing has been first adopted for OLAP workloads, in particular data warehousing, as it is easier to parallelize read-only queries. For the same reason, it has been adopted for big data systems, which are typically read-intensive.

A final question is: can shared-nothing be used to support big write-intensive transactional workloads as well? NoSQL (see our blog post on NoSQL), in particular key-value systems, have excellent horizontal scalability, i.e., scaling over a cluster of nodes. However, since they did not manage to scale transactional management, they gave up on transactional consistency; choosing to focus instead on scalability. However, it is possible to provide both scalability for data management and transactional management (see our blog post on the CAP theorem). Yet, it is a hardcore problem that only a few systems in the NewSQL category have managed to solve (see our blog post on NewSQL).

MAIN TAKEAWAYS

There are three main parallel database architectures: shared-memory, shared-disk and shared nothing. Shared-memory is used for in-memory databases, and shared-disk for small clusters. Shared nothing is becoming the dominant technology since it works anywhere, from on-premise cluster to private or public cloud. The main challenge is to attain linear scalability and transactional (ACID) consistency. Initially, key-value NoSQL systems have been able to achieve linear scalability, but by sacrificing transactional consistency. More recently, a few NewSQL systems have been able to achieve both linear scalability and transactional consistency.





Thursday, October 7, 2021

能源期市数据

「能源期市数据」经济走弱拖累原油需求,WTI原油期货价格震荡为主 

截至10月30日北京时间17时40分,NYMEX11月原油期货(CONX)上涨1.22美元,涨幅1.56%,报79.54美元/桶。国际衍生品智库分析师认为,油价突破新高主要是两方面原因,一是欧佩克+相关国家决定欧佩克继续按原计划每月增产40万桶/日而不是市场所预计的增产稳价。而后WTI原油期货成功站上78美元/桶,为2014年11月以来首次,布伦特原油上破81美元/桶,续刷2018年10月以来新高。二是天然气价格走高提振油价,且美国加利福尼亚州南部奥兰治县海岸日前发生严重原油泄漏事故,利多市场。而俄罗斯普京发生抑制天然气价格大幅走高,欧盟也在应对能源危机方案,油价在假期最后回跌。原油库存连续两周增加也在抑制油价过热的上涨情绪,综合来看,短期油价预计高位震荡,而中期天然气价格高企促使消费转向石油,寒冷冬季造成原油消费激增,以及美国重新开放边境导致航空需求上升,中期油价并无下行的驱动,仍然是偏强走势。NYMEX 11月原油期价短期关注75-76美元附近支撑位,建议背靠支撑位逢低入多为主。

全球性能源短缺

 对于当前全球性能源短缺的"可持续性",陈洪斌分析认为,在短期之内都很难缓解,可能还要持续很长一段时间,预计今年四季度北半球将迎来一个"昂贵的冬天"。

东亚前海首席策略分析师易斌则向记者坦言,展望四季度,随着北半球即将迎来冬季用煤和用电高峰,石化品短缺的局面短期难以化解。经济重启带来的供需缺口与强势能源价格将继续推动海外通胀处于高位,与此同时能源供给短缺与价格高企也将对工业生产与经济增长构成拖累,全球经济面临类滞胀压力。与此同时,他指出,全球流动性预期仍然继续收紧,"10月6日新西兰央行宣布将基准利率上调25个基点至0.5%,并表示明年将进一步加息,以抑制通货膨胀和不断上涨的房价。此前韩国、挪威央行也进行了加息,欧央行和美联储均释放出了偏鹰派的信号,物价上涨压力将进一步推动主要经济体货币政策正常化的进程。对于权益市场而言,类滞胀环境叠加利率水平的上升与去年末的市场环境存在本质差异,市场很难复制当时的再通胀交易,海外市场的风险仍有待进一步释放。"

至于说,能源短缺会不会影响到国内市场,是不是会影响我国的经济,陈洪斌谈了自己的三点思考:

"第一,我们首先要考虑到能源价格高企对于每个板块都是有影响的。第二,我们要考虑的是,我们国家的政策一定会保经济、保民生,一定会对抗全球性的问题,但是我们有多少张牌,哪些政策会怎么出台,这是我们做宏观研究必须要提前做预判的,不仅仅是现在眼前‘双控’的问题。第三,虽然我们处在全球产业的中游,但由于现在全球产业链非常稀缺,所以我们在转移通胀上是有话语权的,因此,我们可以选择把通胀转移向下端,甚至也有可能把这个通胀转移向上端。如果我们能有效的把通胀压力转移出去,将有利于我们国家维护自己经济的正常运转。我们看,现在整体的中央政策,无论是从央行的货币政策,还是相应的贸易政策、财政政策,其实都做出了很多的调整,预计在接下来的这一个季度,我们相信会有更多的政策出台。"

对于近期全球范围的油、气、电等价格上涨,引发的广泛的"能源危机"与"滞胀"担忧,一些机构认为,虽然短期内危机恐无法快速解除,但市场对此也无需过度解读,另外需要将国内外市场做一定的区别对待。

中金公司策略团队日前指出,这些能源类价格的上涨背后的原因较为综合,可能受疫情、天气、地缘关系以及减碳举措等多方面因素的综合影响,这些因素主要来自供给侧,与七十年代需求旺盛、供应中断造成持续的"能源危机"以及长期"滞胀"还是有根本的不同。从中国的角度来看,三季度愈演愈烈的缺电缺煤,也更多是供给侧的因素导致,但在国庆前夕国家发改委已经开始组织会议部署能源电力保供工作,国资委也强调将保供作为能源企业的考核要求,同时媒体报道煤炭进口也有所松动,这些都对四季度及明年上半年中国煤炭及电力等供应提供了保障。

易斌向记者表示,由于国内市场,无论是商品价格还是货币政策的预期修正都要早于海外市场,因此相对而言直接冲击要小于海外市场。在四季度宏观经济增速回归常态的背景下,今年货币政策推进将呈现明显前置,四季度信贷投放和地方债发行都有望呈现超越季节性扩张。继7月15日全面降准后宽信用政策预期逐步发酵,8月23日央行货币信贷形势分析座谈会指出"增强信贷总量增长的稳定性"、"要促进实际贷款利率下行";央行货币政策委员会三季度例会再提"增强信贷总量增长的稳定性"。另外一方面,因为疫情扰动带来盈利增长的异常波动逐步消退,对于未来盈利的可预测性大幅增强,这也将推动市场在经济正常化后的第一个财报季后,对于未来的长期增长作出新的预测。此外,随着中美贸易谈判逐步回归正轨,叠加10月底的G20峰会和全球气候大会,也会进一步提升国内市场风险偏好。

对于这轮全球滞胀的压力还会持续多久?招商基金研究部首席经济学家李湛接受记者采访表示,短期来看,四季度的滞胀压力仍然比较大,但预计,在北半球冷冬旺季需求过后,国内外的能源危机将得到部分缓解,滞胀格局的阶段性终结可能会出现在明年一季度后,但完全打消恐慌需等待两个信号,一是可再生能源供电占比出现快速提升,以及碳中和相关政策节奏边际调整。二是发达经济体步入货币政策收紧通道,全球流动性对大宗商品价格带来抑制。

李湛提醒道,中长期看,碳中和主导下,能源体系变革之中的能源供需矛盾依然大,能源危机出现频率或更高,同时,新能源相关金属供需矛盾加大,滞胀的压力可能会延续较长时间。其一,传统能源受制于碳中和政策,主动收缩产能或进行转型,供需矛盾加大。2021年以来雪佛龙(Chevron)、埃克森美孚(EXXON)、壳牌(Shell)等能源公司都加大了清洁能源生产的投入或将石油资产出售,这将导致全球的能源供应能力受到影响。即使保障产能,高碳能源要实现碳中和也要付出高额的溢价。其二,新能源发电受天气等因素影响较大,放大了能源供应链的脆弱性。新旧能源转型的背景下,"能源危机"出现的频率我们认为未来依然会明显增加。其三,构建新能源体系及其他行业实现碳中和,将带动大宗商品的需求,特别是相关金属。根据国际能源署(IEA)的估计,铜,锂、钴、镍、稀土等新能源相关金属如果在2050年实现碳中和的目标下,其整体需求将扩张6倍。而中长期矿山面临产能不足,供给缺口放大。

Crude oil price: Energy Aspects Ltd.首席石油分析师Amrita Sen

 【油价止跌回升,因美国称目前没有投放战略石油储备计划】① 纽约原油期货周四止跌回升,收涨1.1%,因美国能源部周四表示,目前没有释放战略石油储备以遏制汽油价格上涨的计划;② 英国《金融时报》 周三的一篇报道称,美国能源部长提出了释放战略石油储备的可能性,原油价格一度下跌2.7%;对此美国能源部周四发布声明称:“能源部继续监控全球能源市场供应,并将与我们的合作机构一同确定是否以及何时需要采取行动。工具箱中的所有工具都在考虑之列,但目前还没有采取行动的计划”;美国能源部发言人称,没有寻求禁止原油出口;③ 市场焦点现在回到全球天然气供应短缺,这势必会提高今冬发电对原油的需求;拜登政府越来越多地公开表达对能源价格高企的担忧;Energy Aspects Ltd.首席石油分析师Amrita Sen表示,要记住的关键一点是,拜登政府非常迫切希望给消费者低廉的汽油,因此,如果油价继续上涨和过热,美国就会对OPEC施压;④ 花旗称,OPEC+加速增产“只是时间问题”,特别是如果油价超过每桶80美元的话;⑤ 西德克萨斯中质油11月交割原油期货上涨87美分,结算价报每桶78.30美元;布伦特12月交割原油期货上涨87美分,结算价报每桶81.95美元。

Amrita Sen: Analyst Sen Sees Energy Prices Staying High for Years

Oct. 7, 2021

Here is the link. 

  1. Northern hemisphere winter - a lot of uncertainties 
  2. Inventories are everywhere crude particularly but even products very very low - inventory dynamics
  3. 12 - 18 months guess - high frequency data 


How investors can hunt for opportunities in a volatile market

Oct. 7, 2021

Here is the link. 

Josh Brown of Ritholtz Wealth Management and Pete Najarian of MarketRebellion.com join 'Halftime Report' to discuss where they're finding opportunities to invest in a downturn market.

Goldman Sachs Damien Courvalin: Global oil supply-demand deficit larger than expected: Goldman Sachs' Courvalin

Oct. 7, 2021

Here is the link. 

Damien Courvalin, Goldman Sachs head of energy research, joins 'Power Lunch' to discuss why he raised his year-end price target for oil and walks through the state of the energy market. For access to live and exclusive video from CNBC subscribe to CNBC PRO: https://cnb.cx/2NGeIvi

Goldman Sachs hikes Brent crude oil prices to $90 by years end

Oct. 7, 2021

Here is the link. 

#oilprices #Brentcrude #gasprices Yahoo Finance's Brian Sozzi and Julie Hyman spoke with Goldman Sachs Head of Energy Research Damien Courvalin about the outlook for oil and gas prices. Don't Miss: Valley of Hype: The Culture That Built Elizabeth Holmes WATCH HERE: https://youtu.be/Sb179GLPNYE


Goldman Raises Year-End Brent Forecast to $90

Oct. 7, 2021

Here is the link. 

Oct.04 -- Goldman Sachs Head of Energy Research Damien Courvalin says declining oil inventories, stalled U.S.- Iran negotiations, and OPEC+ are behind the bank's bullish $90 a barrel year-end forecast. "Inventories are about to fall to their lowest in 10 years. And in our view that requires another leg higher" Courvalin says on "Bloomberg Markets."

OPEC+ to Stick With Plan for Oil Output: Sen

Oct. 7, 2021

Here is the link.

Oct.04 -- Amrita Sen, director and founder of Energy Aspects, discusses the current energy crunch across the globe, looks ahead to the OPEC+ meeting later and gives her outlook for energy prices. She speaks on “Bloomberg Daybreak: Europe.”

Natural Gas, Oil to Stay Structurally Higher: Energy Aspects’ Sen

Oct. 7, 2021

Here is the link.

Sep.27 -- Amrita Sen, director of research at Energy Aspects, discusses the factors pushing natural gas and oil prices higher. She speaks on "Bloomberg Markets."

Review: Hackerrank: Kindergarten Adventures | Binary index tree | Segment tree

 https://codereview.stackexchange.com/questions/158242/hackerrank-kindergarten-adventures

Log Structured Merge Tree (LSM-tree) Implementations (a Demo and LevelDB)

Here is the link.  


Paper reading: The Log-Structured Merge-Tree (LSM-Tree)

 ABSTRACT. 

High-performance transaction system applications typically insert rows in a History table to provide an activity trace; at the same time the transaction system generates log records for purposes of system recovery. Both types of generated information can benefit from efficient indexing. An example in a well-known setting is the TPC-A benchmark application, modified to support efficient queries on the History for account activity for specific accounts. This requires an index by account-id on the fast-growing History table. Unfortunately, standard disk-based index structures such as the B-tree will effectively double the I/O cost of the transaction to maintain an index such as this in real time, increasing the total system cost up to fifty percent. Clearly a method for maintaining a real-time index at low cost is desirable. The Log-Structured Merge-tree (LSM-tree) is a disk-based data structure designed to provide low-cost indexing for a file experiencing a high rate of record inserts (and deletes) over an extended period. The LSM-tree uses an algorithm that defers and batches index changes, cascading the changes from a memory-based component through one or more disk components in an efficient manner reminiscent of merge sort. During this process all index values are continuously accessible to retrievals (aside from very short locking periods), either through the memory component or one of the disk components. The algorithm has greatly reduced disk arm movements compared to a traditional access methods such as B-trees, and will improve cost performance in domains where disk arm costs for inserts with traditional access methods overwhelm storage media costs. The LSM-tree approach also generalizes to operations other than insert and delete. However, indexed finds requiring immediate response will lose I/O efficiency in some cases, so the LSM-tree is most useful in applications where index inserts are more common than finds that retrieve the entries. This seems to be a common property for History tables and log files, for example. The conclusions of Section 6 compare the hybrid use of memory and disk components in the LSM-tree access method with the commonly understood advantage of the hybrid method to buffer disk pages in memory.

Bigtable: A Distributed Storage System for Structured Data | Related work

Oct. 7, 2021

Introduction

It is a good idea to work on the paper reading again and again. I like to work on related work this time. 

My notes

  1. Google Boxwood project - distributed agreement, locking, distributed chunk storage, and distributed B-tree storage
  2. Google topics: distributed agreement, locking, distributed chunk storage, and distributed B-tree storage
  3. distributed agreement - 
  4. locking 
  5. distributed chunk storage 
  6. distributed B-tree storage
  7. Bigtable - the goal of Bigtable is to directly support client applications that wish to store data
  8. The Boxwood project - provide infrastructure for building higher-level services such as file systems or databases
  9. Unrelated topics - distributed hash tables, CAN, Chord, Tapestry
  10. Statement: the key-value pair model provided by distributed B-trees or distributed hash tables is too limiting
  11. a shared-nothing [33] architecture - https://www.linkedin.com/pulse/case-shared-nothing-ricardo-jimenez-peris/
  12. the Log-Structured Merge Tree [26] stores updates to index data
  13. Read the paper related to the Log-Structured Merge Tree

Related work

The Boxwood project [24] has components that overlap in some ways with Chubby, GFS, and Bigtable,  since it provides for distributed agreement, locking, distributed chunk storage, and distributed B-tree storage. In each case where there is overlap, it appears that the Boxwood’s component is targeted at a somewhat lower level than the corresponding Google service. The Boxwood project’s goal is to provide infrastructure for building higher-level services such as file systems or databases, while the goal of Bigtable is to directly support client applications that wish to store data.

Many recent projects have tackled the problem of providing distributed storage or higher-level services over wide area networks, often at “Internet scale.” This includes work on distributed hash tables that began with projects such as CAN [29], Chord [32], Tapestry [37], and Pastry [30]. These systems address concerns that do not arise for Bigtable, such as highly variable bandwidth, untrusted participants, or frequent reconfiguration; decentralized control and Byzantine fault tolerance are not Bigtable goals.

In terms of the distributed data storage model that one might provide to application developers, we believe the key-value pair model provided by distributed B-trees or distributed hash tables is too limiting. Key-value pairs are a useful building block, but they should not be the only building block one provides to developers. The model we chose is richer than simple key-value pairs, and supports sparse semi-structured data. Nonetheless, it is still simple enough that it lends itself to a very efficient flat-file representation, and it is transparent enough (via locality groups) to allow our users to tune important behaviors of the system.

Several database vendors have developed parallel databases that can store large volumes of data. Oracle’s Real Application Cluster database [27] uses shared disks to store data (Bigtable uses GFS) and a distributed lock manager (Bigtable uses Chubby). IBM’s DB2 Parallel Edition [4] is based on a shared-nothing [33] architecture similar to Bigtable. Each DB2 server is responsible for a subset of the rows in a table which it stores in a local relational database. Both products provide a complete relational model with transactions.

Bigtable locality groups realize similar compression and disk read performance benefits observed for other systems that organize data on disk using column-based rather than row-based storage, including C-Store [1, 34] and commercial products such as Sybase IQ [15, 36], SenSage [31], KDB+ [22], and the ColumnBM storage layer in MonetDB/X100 [38]. Another system that does vertical and horizontal data partioning into flat files and achieves good data compression ratios is AT&T’s Daytona database [19]. Locality groups do not support CPU cache-level optimizations, such as those described by Ailamaki [2].

The manner in which Bigtable uses memtables and SSTables to store updates to tablets is analogous to the way that the Log-Structured Merge Tree [26] stores updates to index data. In both systems, sorted data is buffered in memory before being written to disk, and reads must merge data from memory and disk.

C-Store and Bigtable share many characteristics: both systems use a shared-nothing architecture and have two different data structures, one for recent writes, and one for storing long-lived data, with a mechanism for moving data from one form to the other. The systems differ significantly in their API: C-Store behaves like a relational database, whereas Bigtable provides a lower level read and write interface and is designed to support many thousands of such operations per second per server. C-Store is also a “read-optimized relational DBMS”, whereas Bigtable provides good performance on both read-intensive and write-intensive applications.

Bigtable’s load balancer has to solve some of the same kinds of load and memory balancing problems faced by shared-nothing databases (e.g., [11, 35]). Our problem is somewhat simpler: (1) we do not consider the possibility of multiple copies of the same data, possibly in alternate forms due to views or indices; (2) we let the user tell us what data belongs in memory and what data should stay on disk, rather than trying to determine this dynamically; (3) we have no complex queries to execute or optimize.

 

Wednesday, October 6, 2021

System design: Web crawler | BigTable

I like to learn from the following writing in a book: 

In 2003, Google published a paper titled “The Google File System”. This scalable distributed file system, abbreviated as GFS, uses a cluster of commodity hardware to store huge amounts of data. The filesystem handled data replication between nodes so that losing a storage server would have no effect on data availability. It was also optimized for streaming reads so that data could be read for processing later on. 

Shortly afterward, another paper by Google was published, titled “MapReduce: Simplified Data Processing on Large Clusters”. MapReduce was the missing piece to the GFS architecture, as it made use of the vast number of CPUs each commodity server in the GFS cluster provides. MapReduce plus GFS forms the backbone for processing massive amounts of data, including the entire search index Google owns. 

What is missing, though, is the ability to access data randomly and in close to real-time (meaning good enough to drive a web service, for example). Another drawback of the GFS design is that it is good with a few very, very large files, but not as good with millions of tiny files, because the data retained in memory by the master node is ultimately bound to the number of files. The more files, the higher the pressure on the memory of the master.

So, Google was trying to find a solution that could drive interactive applications, such as Mail or Analytics, while making use of the same infrastructure and relying on GFS for replication and data availability. The data stored should be composed of much smaller entities, and the system would transparently take care of aggregating the small records into very large storage files and offer some sort of indexing that allows the user to retrieve data with a minimal number of disk seeks. Finally, it should be able to store the entire web crawl and work with MapReduce to build the entire search index in a timely manner.

Being aware of the shortcomings of RDBMSes at scale (see “Seek Versus Transfer” on page 315 for a discussion of one fundamental issue), the engineers approached this problem differently: forfeit relational features and use a simple API that has basic create, read, update, and delete (or CRUD) operations, plus a scan function to iterate over larger key ranges or entire tables. The culmination of these efforts was published in 2006 in a paper titled “Bigtable: A Distributed Storage System for Structured Data”, two excerpts from which follow:

Bigtable is a distributed storage system for managing structured data that is designed to scale to a very large size: petabytes of data across thousands of commodity servers. 

…a sparse, distributed, persistent multi-dimensional sorted map.

It is highly recommended that everyone interested in HBase read that paper. It describes a lot of reasoning behind the design of Bigtable and, ultimately, HBase. We will, however, go through the basic concepts, since they apply directly to the rest of this book. 

HBase is implementing the Bigtable storage architecture very faithfully so that we can explain everything using HBase. Appendix F provides an overview of where the two systems differ.


Leetcode algorithm: 630 solved | Instagram | Equity research

Oct. 6, 2021

Introduction

It is a lonely journey to build a successful career, I have to work on my advancement continuously. With great learning experience, I have over 10 onsite interviews from Facebook (3), Amazon (4), Microsoft (2), Google (1) from 2008 to 2021. I learned so much from those interview experience. Last two years, I started to learn how to invest on stock market, it is tough since I did not make any gain on my Questrade.com TFSA accounts. Last 11 years, I did not purchase any real estate in Canada. 

Case study | Lack of confidence | My impulsive decision making | Six months efforts

Recently I sold my 20,000 shares of GTE stock at price 0.91/ share, after Amazon onsite, in less than 2 weeks, the price went up to 1.10/ share. I guessed that I must be lack of confidence on my own research.

I have to watch out myself to check my confidence level. I have to learn how to work with myself better. Train myself better to handle stress of investment. 

Work on something else | Read more large distributed system books | Invest and work on equity research 

I have a wonderful life in the city of Vancouver, stay healthy and enjoy simple life. Being frugal, I choose to stay at a tiny bedroom rented, drive a 2001 Honda accord, in July I chose to ask my neighbor to help maintenance jobs in order to cut cost, a series of jobs including coolant flush, air conditioning fix, gasket replacement, climate control unit order in ebay and replacement, fix of front body dent etc., in total of $660, and work on equity research, and being a lonely investor on stock market. 

I just keep learning. So much fun in terms of 55 year old, I did run 10 K race first time less than a month ago, and also take care of my frozen shoulder, try to warm up my shoulder every day before going to work. 

I also like to solve more algorithms on Leetcode.com. And also I like to post more videos on Instagram, and learn how to use my Google 4a 5G more often to record good videos and I like to be creative, learn more on Instagram.com.