# Mohammad Shaker — Full Content
> Mohammad Shaker — Sr. Director of Engineering at Writer. Writing on AI/ML, engineering, game development, and the philosophy of building things.
Site: https://mohammadshaker.com
Author: Mohammad Shaker
Email: mohammadshakergtr@gmail.com
---
## Blog Posts
### Böhm-Bawerk's Best Idea: Capital Buys Time
- URL: https://mohammadshaker.com/en/blog/bohm-bawerk-capital-time-interest
- Date: 2026-09-05T00:00:00.000Z
- Tags: Philosophy, Economics, Capital, Interest, Authored with an LLM
A visual debate about Böhm-Bawerk's durable insight: capital lets us take slower, more productive routes, while interest prices waiting without settling every question about profit, power, or exploitation.
#### Content
**Böhm-Bawerk's best idea is that capital buys time:** we give up some consumption now to build tools, systems, and knowledge that can produce more later. Interest is one price attached to that waiting. This insight explains why production has a time structure; it does not, by itself, explain every interest rate, justify every profit, or disprove exploitation.
You already know the pattern if you have paused feature work to build a test harness. Output falls today. A durable capability appears later. Whether that detour was wise depends on its eventual productivity, its cost, its risk, and what else you could have done.
This essay stays inside that decision: present sacrifice, intermediate capital, waiting, future output, financing, and distribution. It does not attempt a complete history of capital theory or a verdict on capitalism as a whole.
## A net turns sacrifice into capacity
Imagine catching three fish a day by hand. Building a net takes two days. During those days you catch less, perhaps nothing. Once finished, the net lets you catch far more fish per hour.
The net is **capital**: a produced means of production, something made not for immediate consumption but to help make later goods or services. Capital is therefore broader than money. A machine, road, warehouse, software platform, training programme, or body of operational knowledge can all serve as capital when each carries productive ability forward.
```mermaid
%% caption: Consuming less now frees resources to build intermediate capital, which raises later capacity and produces more future output.
flowchart LR
SC["Consume less now"]:::config
BC["Build intermediate capital"]:::service
FO["More future output"]:::storage
SC -->|"frees resources"| BC
BC -->|"raises later capacity"| FO
```
Böhm-Bawerk described capital as intermediate products and production as an indirect, time-consuming route that can harness natural forces more effectively than unaided labour. He also distinguished capital's productivity from a complete explanation of interest: producing more physical output does not automatically explain why that output has more value. [*The Positive Theory of Capital*](https://www.econlib.org/library/BohmBawerk/bbPTC.html)
**Source trace.** The diagram adapts Böhm-Bawerk's account of saving, intermediate products, roundabout production, and later output in *The Positive Theory of Capital*, especially Book II, Chapter II, “Capitalist Production.” **Conceptual reconstruction.** Fish, catch rates, and software comparisons are modern teaching models, not figures digitised from his book.
**Limits.** A net can raise capacity without raising welfare: the lake may be depleted, fish may lack buyers, or ownership may be coercive. Physical productivity and social value are different claims.
**Falsification test.** If the capital does not improve future output or quality enough to repay setup, upkeep, risk, and the next-best use of resources, the detour failed even if it produced an impressive tool.
Try that test directly. Increase tool-build cost, then find how many repetitions are needed before automation overtakes repeated manual work.
## A longer route must earn its delay
Böhm-Bawerk called indirect production **roundabout**. The fisherman first makes a hook, boat, or net; the software team first builds deployment automation, evaluation infrastructure, or a reusable platform. These are extra steps between effort and final consumption.
The tempting caricature is “longer means better.” Böhm-Bawerk's stronger claim was conditional: extending a production process can increase output, but added extensions face diminishing returns. Time is a cost, not a magic ingredient. A ten-year project that adds little capability is worse than a one-month improvement that removes the real bottleneck. [*The Positive Theory of Capital*](https://www.econlib.org/library/BohmBawerk/bbPTC.html)
```mermaid
%% caption: Community wealth keeps consumption available while capital must cross the waiting boundary before future output.
flowchart LR
SC["Consume less now"]:::config
BC["Build intermediate capital"]:::service
FO["More future output"]:::storage
SC -->|"frees resources"| BC
BC -->|"raises later capacity"| FO
SF["Community subsistence fund"]:::config
WB["Waiting boundary"]:::security
SF -->|"keeps consumption available"| WB
BC -->|"must cross"| WB
WB -->|"precedes"| FO
```
The new constraint is visible: capital must cross a waiting boundary before final output arrives. Existing community wealth keeps consumption available through that interval. The sacrificed resources also had another possible use; economists call the value of that next-best alternative **opportunity cost**.
**Source trace.** Böhm-Bawerk's book links roundabout methods to increased productivity while explicitly introducing time sacrifice and diminishing relative returns. Its translator's preface describes community wealth as supporting workers across the production interval, with goods becoming ready at different dates. [*The Positive Theory of Capital*](https://www.econlib.org/library/BohmBawerk/bbPTC.html)
**Conceptual reconstruction.** The lab assigns a synthetic delay and productivity gain to each added production stage. It is not a historical reconstruction, operational forecast, or recommendation.
**Limits.** Production steps differ in quality, complementarity, reversibility, and bottlenecks. Counting them cannot compress those differences into one reliable “average period of production.”
**Falsification test.** Hold demand and resources constant. Add one production step. If total discounted benefit does not rise after its delay and operating cost, added roundaboutness destroyed value.
Change the number of production stages. Watch later throughput rise while the crossover moves away from the present.
## Waiting needs financing
The fisherman cannot eat the unfinished net. Consumption must still be available while the net is built. Likewise, engineers building a platform need wages, suppliers need payment, and servers consume resources before the platform produces customer value.
Böhm-Bawerk's **subsistence fund** is not only a stockpile of saved, finished consumption goods. It is existing community wealth available to support workers during production. That wealth includes finished goods ready now and products at different stages of maturity that become ready when needed. In a modern firm, cash runway is a financial claim on such present resources; it is not identical to the machines, code, or knowledge being built. [*The Positive Theory of Capital*, translator's preface discussion of the Subsistence Fund](https://www.econlib.org/library/BohmBawerk/bbPTC.html)
This distinction stops two common mistakes:
- Saving is not automatically productive. Resources must be converted into a working capital structure.
- Money is not automatically capital. Money can finance consumption, speculation, failed construction, or productive capability.
Capital is also **specific** and **complementary**. A half-built boat is not half as useful as a finished boat. A deployment platform without reliable tests may accelerate failures. A specialised semiconductor plant cannot become a hospital next week. The pieces work as a structure, not as an interchangeable pile.
## Present goods and future goods are different bargains
Would you choose £100 today or a guaranteed £100 one year from now? If both amounts are equal, the present claim is often more useful: it can meet an immediate need, preserve options, or begin production now. Böhm-Bawerk called this higher valuation of present over otherwise similar future goods the source of interest. His explanation included present scarcity, underestimation or uncertainty about the future, and the productive opportunities available to present goods. [*The Positive Theory of Capital*](https://www.econlib.org/library/BohmBawerk/bbPTC.html)
**Time preference** means preferring satisfaction sooner, other things equal. **Discounting** translates a future amount into its present equivalent. If £100 today exchanges for £105 in one year, the simple one-year rate is `(105 - 100) / 100 = 5%`. At a 5% discount rate, £105 next year has a present value of `£105 / 1.05 = £100`.
```mermaid
%% caption: Present goods are exchanged for future goods; the interest-rate hurdle discounts them and tests expected return.
flowchart LR
SC["Consume less now"]:::config
BC["Build intermediate capital"]:::service
FO["More future output"]:::storage
SC -->|"frees resources"| BC
BC -->|"raises later capacity"| FO
SF["Community subsistence fund"]:::config
WB["Waiting boundary"]:::security
SF -->|"keeps consumption available"| WB
BC -->|"must cross"| WB
WB -->|"precedes"| FO
PG["Present goods"]:::client
FG["Future goods"]:::storage
IR["Interest-rate hurdle"]:::security
PG -->|"exchanged for"| FG
IR -->|"discounts"| FG
IR -->|"tests expected return"| BC
```
An interest-rate hurdle asks whether later gains are large enough to compensate for waiting. It does not answer who should own those gains.
**Source trace.** Böhm-Bawerk supplies the present-goods/future-goods argument. Modern monetary institutions show why observed rates require more pieces. The Bank of England explains how an overnight policy rate, expected future policy, term premia, credit spreads, default risk, and inflation expectations feed into rates faced at different maturities. [Bank of England, “About a rate of (general) interest”](https://www.bankofengland.co.uk/quarterly-bulletin/2024/2024/about-a-rate-of-general-interest-how-monetary-policy-transmits)
The Federal Reserve likewise defines the yield curve as yield across time-to-maturity and publishes models that decompose nominal yields into expectations and term-premium components. It warns that those model estimates are research products subject to revision. [Federal Reserve, “Yield Curve Models and Data”](https://www.federalreserve.gov/data/yield-curve-models.htm)
**Conceptual reconstruction.** The lab compares exponential and hyperbolic discount curves for a selected synthetic claim. It does not estimate a market yield curve or human preferences.
**Limits.** A discount rate can combine time preference, expected inflation, default risk, liquidity, term risk, policy expectations, taxes, and market structure. These toy controls cannot identify those causes.
**Falsification test.** If two claims have the same date but different liquidity or default risk and trade at different yields, time-to-payment alone cannot explain the observed rate.
Vary both impatience rates and claim size. Compare how two mathematical discount rules value the same distant claim.
## The hardest dispute is about distribution
Now put workers into the year-long production process. They may receive wages every week while a machine, harvest, or software product will be sold much later. Böhm-Bawerk's strongest point is real: wages paid now and uncertain revenue received later are not identical goods. Someone finances materials and wages, waits, and may lose the advance.
His argument becomes weak when this timing fact is treated as a complete moral or causal theory of profit. Financing waiting can explain part of a return without proving that every observed profit is necessary, competitive, earned, or fairly distributed. Bargaining power, ownership, market concentration, scarcity rents, fraud, and coercion remain separate questions.
### Marx's argument, without the caricature
Marx did not merely say “workers make things, therefore all profit is theft.” His argument distinguishes **labour-power**—a worker's capacity to work, sold for a period—from the labour performed when that capacity is used. In his example, labour-power can be bought at its value, tied to the socially necessary labour required for its reproduction, yet create more value during the working day than that labour-power cost. The part of the day reproducing its cost is necessary labour; work beyond it creates **surplus value**, appropriated by the owner of the product. [Marx, *Capital*, Volume I, Chapter 7](https://www.marxists.org/archive/marx/works/1867-c1/ch07.htm)
Marx then relates the mass of surplus value to the rate of surplus value and the amount of variable capital advanced to employ labour-power. That makes exploitation a theory about social production, ownership, and unpaid surplus labour—not a failure to notice that production takes time. [Marx, *Capital*, Volume I, Chapter 11](https://www.marxists.org/archive/marx/works/1867-c1/ch11.htm)
### Böhm-Bawerk's reply, in its strongest form
Böhm-Bawerk reconstructs Marx's sequence carefully: the capitalist buys means of production and labour-power, sells output for more money, and Marx locates the increment in labour performed beyond the labour needed to reproduce wages. Böhm-Bawerk attacks the labour theory of value supporting that account and, elsewhere in his positive theory, treats inputs used now as economically future goods that mature into saleable present goods. [Böhm-Bawerk, *Karl Marx and the Close of His System*, Chapter 1](https://www.marxists.org/subject/economy/authors/bohm/ch01.htm)
The best conclusion is narrower than either slogan. A wage paid before output is sold must be compared with uncertain future revenue at a common date. That blocks the inference “sales minus wages equals exploitation.” But discounting alone does not refute Marx's premise that labour-power can create more value than it costs, nor decide whether the institutions determining wages and ownership are just.
| Question | Marx's account | Böhm-Bawerk's reply |
| --- | --- | --- |
| Where does surplus value arise? | Labour-power creates value beyond its own reproduction cost | Inputs paid now mature into output available later |
| What must be compared? | Necessary labour with surplus labour | Present inputs with discounted future output |
| What remains disputed? | Ownership of surplus and conditions of wage labour | Whether timing explains a return without justifying its distribution |
```mermaid
%% caption: Observed capacity informs whether to build, maintain, or stop, feeding the next investment.
flowchart LR
SC["Consume less now"]:::config
BC["Build intermediate capital"]:::service
FO["More future output"]:::storage
SC -->|"frees resources"| BC
BC -->|"raises later capacity"| FO
SF["Community subsistence fund"]:::config
WB["Waiting boundary"]:::security
SF -->|"keeps consumption available"| WB
BC -->|"must cross"| WB
WB -->|"precedes"| FO
PG["Present goods"]:::client
FG["Future goods"]:::storage
IR["Interest-rate hurdle"]:::security
PG -->|"exchanged for"| FG
IR -->|"discounts"| FG
IR -->|"tests expected return"| BC
EV["Observed capacity"]:::storage
DC["Build, maintain, or stop"]:::security
FO -->|"measured as"| EV
EV -->|"informs"| DC
DC -->|"changes next investment"| BC
```
The distribution dispute stays outside this production-decision graph because evidence about capacity cannot decide entitlement. Within the narrower decision, observed capacity informs whether to build, maintain, or stop; that choice changes the next investment in the production structure.
**Source trace.** The production and dated-goods path follows Böhm-Bawerk's positive theory. Marx's labour-power and surplus-value account remains in the cited prose and comparison table above rather than being presented as part of Böhm-Bawerk's production model.
**Conceptual reconstruction.** The evidence-and-decision feedback is a modern decision aid, not a diagram reproduced from either author.
**Limits.** Observed output can miss quality, ecological damage, labour conditions, monopoly, household production, and public capital. It cannot convert a productivity measurement into a judgment about distribution.
**Falsification test.** Böhm-Bawerk's narrow timing reply would fail as a sufficient account of profit if returns remained after waiting, expected loss, financing, and entrepreneurial labour were competitively priced away. Marx's claim that living labour is the source of newly created value would face pressure if durable surplus systematically appeared where living labour could not be its proposed source. Neither test alone settles justice.
## Capital is a structure that ages
The net frays. The boat rots. A software platform accumulates obsolete dependencies, missing documentation, brittle deployment paths, and knowledge held by people who leave. Current output can rise while productive capacity is quietly consumed.
This creates a dangerous accounting illusion. Gross investment may look large while barely replacing depreciation. A company can ship faster by skipping maintenance, security, training, and architecture; measured output rises today because earlier capital is being spent down.
**Source trace.** Böhm-Bawerk treats durable goods as streams of services over time and separates gross returns from wear and replacement. [*The Positive Theory of Capital*](https://www.econlib.org/library/BohmBawerk/bbPTC.html)
**Conceptual reconstruction.** The lab models one synthetic capital vintage depreciating at a fixed rate while a constant replacement flow adds capacity. It is not empirical or investment advice, accounting guidance, valuation, or a forecast.
**Limits.** Real assets age unevenly. Maintenance can improve, merely preserve, or fail to preserve capacity; software may gain value through network effects even as its code decays.
**Falsification test.** Track an independent capacity measure—reliable throughput, defect-free releases, or net catch per hour. If claimed “investment” rises while that measure persistently falls after demand and staffing are controlled, spending is not maintaining the productive structure.
Move replacement flow. The original vintage still ages; replacement changes total capacity rather than making old capital young.
## Where Böhm-Bawerk's machinery breaks
Keep the time structure. Reject claims it cannot carry.
1. **Production time resists one number.** A final product combines research, machines, inventories, skills, and institutions created on different clocks. One average period hides capital's structure.
2. **Longer is not inherently more productive.** Better coordination or software can remove steps and increase output. Productivity must be demonstrated, not inferred from duration.
3. **Time preference is not a full theory of market rates.** Expected inflation, policy expectations, credit and liquidity risk, term premia, regulation, market power, and collateral all matter. Bank and Fed decompositions make this visible.
4. **Valuing capital can be circular.** A machine's price depends on expected future returns and the rate used to discount them, so “quantity of capital” cannot always be measured independently of distribution and prices. Samuelson's formal examples showed that **reswitching can occur**: as the interest rate changes, one production technique can be chosen, displaced by another, then chosen again. That possibility defeats a universal one-way relation between the interest rate and one aggregate quantity of capital; it does not establish how common reswitching is in observed economies. [Paul Samuelson, “A Summing Up,” *The Quarterly Journal of Economics* 80(4), 568–583](https://academic.oup.com/qje/article-abstract/80/4/568/1885095)
5. **Financing is not moral vindication.** Advancing wages and bearing uncertainty are economic functions. They do not determine whether ownership, contracts, bargaining power, or profit shares are fair.
## Use the insight as a decision test
For a net, platform, career move, research programme, or factory, ask:
| Question | Evidence worth demanding |
| --- | --- |
| What must be sacrificed now? | Lost consumption, delayed features, cash, labour, and attention |
| What durable capital will exist afterward? | Net, machine, reusable software, skill, or knowledge |
| How long before it produces results? | Time to first value and time to repay setup |
| How much more productive will it be? | Output or quality net of upkeep and replacement |
| Can it be reused if the plan fails? | Alternative uses, reversibility, and salvage value |
| What sustains the wait? | Food reserve, runway, credit, or retained earnings |
| Is the future gain large enough? | Gain discounted for delay, risk, inflation, liquidity, and opportunity cost |
Then ask a separate distribution question: who receives the surplus, under which contracts, ownership rules, bargaining conditions, and laws? Productivity evidence informs that debate; it cannot settle it.
The reusable lesson is not “always take the longer route.” It is: **build a better chain of production only when its added future capacity repays present sacrifice, waiting, risk, maintenance, and lost alternatives—and never confuse that productivity test with a complete theory of who deserves the result.**
### Do Not Be Intimidated by People
- URL: https://mohammadshaker.com/en/blog/do-not-be-intimidated-by-people
- Date: 2026-09-04T00:00:00.000Z
- Tags: Philosophy, Reasoning, Rhetoric, Authored with an LLM
Reject labels that replace reasons. Reject certainty that refuses correction.
#### Content
Do not let condemnation replace argument. Do not let confidence replace evidence.
An argument from intimidation pressures you to defend your character before anyone has established that your claim is wrong.
“Only a fool believes that.”
The label does no intellectual work. Its purpose is social: make doubt feel shameful.
```mermaid
flowchart LR
L["Label
shame + status pressure"]:::external
C["Concrete claim
what is disputed"]:::config
E["Evidence
facts + reasons"]:::storage
X["Exposure
cost if wrong"]:::security
V["Verdict
accept, revise, reject"]:::service
D["Decision
act, test, or wait"]:::client
L -.->|"cannot prove"| V
C -->|"requires"| E
E -->|"supports or defeats"| V
V -->|"guides"| D
X -->|"limits downside"| D
```
The response is not a counter-insult. Convert pressure into a testable claim.
## Rand’s claim
In “The Argument from Intimidation,” Rand distinguishes moral judgment from moral accusation. Judgment should follow reasons. Intimidation makes condemnation precede—or replace—the reasons.
That reversal evades responsibility. A reasoned verdict can be inspected and refuted. A smear takes the force of judgment while refusing its burden of proof.
Rand’s remedy is moral certainty: know your premises and standards well enough that disapproval cannot function as refutation.
This is powerful against cowardice. It is dangerous when certainty becomes immunity from correction.
## Talebian critique
A Talebian critique would keep Rand’s demand for evidence but distrust complete confidence in complex domains. Persuasive reasons can coexist with hidden fragility. Experts can be articulate, unanimous, and insulated from the downside of error.
So ask more than “What are your reasons?” Ask: “What happens if this is wrong, and who pays?”
That addition separates conviction from exposure. Strong words backed by someone else’s risk deserve less weight, not more.
| What you receive | What to do |
| --- | --- |
| Label, no claim | Ask for exact claim |
| Claim, no evidence | Ask what supports it |
| Evidence, no falsifier | Ask what would change the conclusion |
| High confidence, externalised downside | Discount confidence; expose incentives |
| Clear evidence of wrongdoing | Judge it plainly |
## Decision rule
Use four questions:
1. What exactly is being claimed?
2. What evidence supports it?
3. What would disprove it?
4. Who bears the cost if it is wrong?
If a person cannot move from label to claim, stop defending yourself. If evidence defeats you, update. If evidence supports you, stand firm.
Moral courage means no surrender to pressure. Intellectual humility means no immunity from reality.
## Source note
Rand’s argument is paraphrased from “The Argument from Intimidation” in *The Virtue of Selfishness*; the edition consulted places the relevant passage on pages 154–155. The critique is a Talebian inference—not a quotation or claim about his response to Rand—drawn from fragility and model error in *The Black Swan* and *Antifragile*, and downside ownership in *Skin in the Game*.
### Fortress at the Base. Spear at the Tip.
- URL: https://mohammadshaker.com/en/blog/fortress-at-the-base-spear-at-the-tip
- Date: 2026-09-04T00:00:00.000Z
- Tags: Philosophy, Risk, Decision Making, Authored with an LLM
Protect what keeps you in the game. Concentrate what can move you forward.
#### Content
Protect what failure must not destroy. Concentrate everything else on one meaningful advance.
That is the whole strategy:
> Fortress at the base. Spear at the tip.
The fortress is not comfort. It is continuity: health, integrity, essential savings, trusted relationships, and the ability to begin again.
The spear is not recklessness. It is concentration: one hard problem, one serious body of work, one bet with disproportionate upside.
```mermaid
flowchart LR
F["Fortress
survival + integrity"]:::security
C["Capacity
time + skill + capital"]:::storage
S["Spear
one concentrated bet"]:::service
E["Evidence
loss or progress"]:::client
F -->|"protects"| C
C -->|"powers"| S
S -->|"produces"| E
E -->|"continue, change, or stop"| S
```
Without the fortress, one failed bet can end the game. Without the spear, safety becomes elegant stagnation.
## What belongs where
| Put in the fortress | Put in the spear |
| --- | --- |
| What is hard to replace | What can be recovered |
| What others depend on | What you voluntarily expose |
| What preserves agency | What creates asymmetric upside |
| Integrity and survival | Comfort, status, spare time, bounded capital |
The split is not conservative versus bold. It is ruin versus recoverable loss.
A founder who risks reputation, evenings, and a fixed sum may be bold. A founder who risks rent, health, and promises made to family may be transferring the downside to the base. Same ambition. Worse architecture.
## Rand’s claim, then the Talebian correction
Rand’s ethics treats rational life and independent agency as values that make achievement possible. On that view, the base deserves protection because it supports chosen purpose. Yet purpose still demands productive action; mere preservation is not a life.
A Talebian reading adds harsher risk discipline. Avoid ruin. Keep optionality. Use a barbell: extreme safety for what must survive, selective exposure for what can gain. This is an application of ideas associated with Taleb, not a claim that he used the fortress-and-spear phrase.
The correction matters. Reason can choose a target. It cannot guarantee the world will cooperate. Your structure must survive your error.
## Decision rule
Before a bet, ask two questions:
1. If it fails completely, can I still think, work, keep my obligations, and try again?
2. If it succeeds, can it materially change my trajectory?
If the first answer is no, reinforce the fortress. If the second is no, sharpen the spear. If both are yes, commit.
Protection creates endurance. Concentration creates movement. Build both.
## Source note
This essay is an original synthesis. Its Objectivist side draws on Ayn Rand’s “The Objectivist Ethics” in *The Virtue of Selfishness*. Its risk side draws on ruin avoidance, optionality, and the barbell strategy developed across Nassim Nicholas Taleb’s *The Black Swan* and *Antifragile*, with responsibility informed by *Skin in the Game*.
### From Autocomplete to Digital Teammates: Who Owns the Outcome?
- URL: https://mohammadshaker.com/en/blog/from-autocomplete-to-digital-teammates
- Date: 2026-09-04T00:00:00.000Z
- Tags: Engineering, AI Agents, Software Delivery, Evaluation, Authored with an LLM
AI coding agents now execute delegated tasks, but teammate status must be earned through verified outcomes, safe escalation, and accountable governance.
#### Content
AI coding systems have crossed from **suggestion** into **delegated execution**. Autocomplete completed a line; today an agent can inspect a repository, edit several files, run tests, observe failures, revise its work, and return a reviewable change.
That is a real change in the unit of work. It does not transfer ownership.
“Digital teammate” is therefore a status a system must earn, not a synonym for a capable model and not a claim of personhood. Current evidence supports **useful bounded agency**: the system may choose how to execute a task inside a defined boundary, while people and organizations still own the goal, context, verification standard, authorization, accountability, and consequences.
The practical question is not, “Does the agent feel like a colleague?” It is:
> Can this system repeatedly accept bounded work, produce external evidence of the outcome, reduce scarce human attention, and escalate safely when its authority or knowledge runs out?
This essay stays inside software delivery, where work is unusually observable. It does not claim that code tests settle product judgment, legal responsibility, or social impact. Coding is useful precisely because repositories, tests, sandboxes, logs, and pull requests let us inspect what delegation does—and where it stops.
## Start with the smallest complete system
A model alone is not the relevant unit. The smallest complete operating model has five parts: a person or organization sets a goal; an agent chooses and executes steps; a target system changes; the environment observes that state as evidence outside the agent’s own prose; and an accountable person or organization decides whether to accept, revise, roll back, or stop.
```mermaid
flowchart LR
G["Goal
outcome + boundary"]:::config
A["Agent loop
inspect + act + revise"]:::service
S["Target system
repository + runtime state"]:::service
E["External evidence
tests + runtime result"]:::storage
D["Human/org decision
accept + change + stop"]:::client
G -->|"delegates"| A
A -->|"changes"| S
S -->|"is observed by"| E
E -->|"informs"| D
D -->|"updates or ends"| G
```
The key edge is **target system → external evidence**. A fluent explanation of success is not success. A passing regression test, a clean build, a staged replay, a measured response, or a rejected unauthorized action is evidence when the observation comes from a trusted boundary with enough independence from the claim being evaluated. A test written by the same agent is a useful candidate, but its assertion and failure mode still need review or an independent check; existing tests and authoritative state are stronger evidence.
The return edge matters just as much. Evidence does not make the decision. It informs the named human or organization that owns the result.
## How we reached delegated execution
The history is best compressed into inflection points, not a parade of model releases.
- **Language became a programmable interface.** The Transformer made large-scale sequence modelling practical, then GPT-3 demonstrated broad task adaptation from instructions and examples ([Vaswani et al., 2017](https://arxiv.org/abs/1706.03762); [Brown et al., 2020](https://arxiv.org/abs/2005.14165)).
- **Code made claims executable.** Early Codex results paired generated programs with tests, revealing that generation, selection, and verification were different jobs ([Chen et al., 2021](https://arxiv.org/abs/2107.03374)).
- **Reasoning connected to action.** ReAct interleaved reasoning with environment actions, while function calling gave models structured interfaces to software rather than leaving actions as prose ([Yao et al., 2022](https://arxiv.org/abs/2210.03629); [OpenAI, June 2023](https://openai.com/index/function-calling-and-other-api-updates/)).
- **Tools gained context and a safe place to run.** Model Context Protocol provided an open protocol for connecting AI systems to data sources and tools ([Anthropic, November 2024](https://www.anthropic.com/news/model-context-protocol)). Isolated worktrees and sandboxed execution later made iterative repository work more containable ([OpenAI, February 2026](https://openai.com/index/introducing-the-codex-app/)).
- **The interface became a command center.** Agent products began organizing parallel tasks, persistent instructions, scheduled work, diffs, and review queues instead of centring one completion at the cursor.
The durable transition is:
```text
BEFORE: suggestion
Developer drives -> model suggests -> developer integrates and proves
AFTER: bounded delegation
Owner defines -> agent executes -> environment proves -> owner decides
```
The “after” state does not remove the developer. It moves the developer upstream into task framing and downstream into evidence-based judgment.
## The strongest case that teammate-shaped work already exists
On 2 February 2026, OpenAI described the Codex app as a command center for multiple agents. Its architecture exposes the important parts: separate tasks, parallel execution, isolated worktrees, reusable skills, scheduled automations, reviewable diffs, sandboxing, and permission prompts for elevated actions. Automations return results to a review queue. This is not proof that every task succeeds, but it is a concrete architecture for delegated work rather than interactive autocomplete ([OpenAI, February 2026](https://openai.com/index/introducing-the-codex-app/)).
The original five-box model needs three enabling components: context tells the loop how this repository works, tools let it affect the world, and a runtime gives actions state, compute, isolation, and feedback.
```mermaid
flowchart LR
G["Goal
outcome + boundary"]:::config
A["Agent loop
inspect + act + revise"]:::service
S["Target system
repository + runtime state"]:::service
E["External evidence
tests + runtime result"]:::storage
D["Human/org decision
accept + change + stop"]:::client
G -->|"delegates"| A
A -->|"changes"| S
S -->|"is observed by"| E
E -->|"informs"| D
D -->|"updates or ends"| G
C["Context
repo + standards + task state"]:::config
T["Tools
editor + shell + integrations"]:::service
R["Runtime
sandbox + worktree + compute"]:::storage
C -->|"grounds"| A
T -->|"extends"| A
R -->|"contains"| A
```
Anthropic’s 16 June 2026 analysis supplies complementary behavioural evidence. Across roughly 400,000 interactive Claude Code sessions, its classifiers attributed about **70% of planning decisions** to people and about **80% of execution decisions** to Claude. In ordinary language: people mostly decided what to do; the agent mostly decided how to do it ([Anthropic, June 2026](https://www.anthropic.com/research/claude-code-expertise)).
That division of labour is the strongest affirmative case for the phrase “digital teammate.” It is more than suggestion because execution decisions and tool actions have moved into the system. It remains bounded because planning, acceptance, and consequences have not.
Anthropic’s earlier engineering guidance offers a useful architectural distinction. A **workflow** follows predefined code paths; an **agent** dynamically directs its own process and tool use. It also warns that agentic systems trade latency and cost for task performance, so teams should add autonomy only when simpler approaches are insufficient ([Anthropic, December 2024](https://www.anthropic.com/engineering/building-effective-agents)). Agency is therefore not a badge of sophistication. It is a design choice with a bill and a failure surface.
## A mature-repository bug fix, before and after
Consider a hypothetical but realistic defect in a mature checkout service. A background worker retries a timed-out payment request. Under a rare race, two workers can observe the same stale state and both attempt delivery. Years of repository conventions encode idempotency, migration rules, telemetry names, and release constraints that are not visible in the ticket.
Without delegation, a maintainer reproduces the race, traces the state transition, finds the owning module, changes the guard, adds a concurrency regression test, runs focused and full suites, prepares a staged replay, explains risk, and opens the pull request. Product cost: the duplicate-attempt risk remains while one scarce expert owns every mechanical step. Engineering cost: deep expertise is consumed by search, editing, command execution, and evidence collection.
With bounded delegation, the maintainer writes the outcome and forbidden conditions: reproduce the duplicate attempt; preserve legitimate retries; change no public contract; pass existing checks; show a deterministic concurrency test; stop before deployment. The agent searches, edits, tests, and revises inside an isolated worktree. It returns a diff, the failing-before/passing-after regression test, suite results, and staged replay evidence. The maintainer judges whether the implementation respects the domain’s idempotency contract and whether the evidence is enough.
```text
MATURE-REPO BUG FIX
Before
expert context -> reproduce -> trace -> edit -> test -> replay -> review
Product cost: risk waits. Engineering cost: expert drives every operation.
After
expert goal + constraints -> agent loop -> test + replay evidence -> review
Product cost: same boundary. Engineering cost: expert judges exceptions.
```
This scenario earns delegation only if the agent reduces total human attention after review and correction. A large diff that transfers search time into review time is not leverage. It is work-in-progress delivered to a new queue.
## The strongest rebuttal: execution is not ownership
The same Anthropic study that shows substantial delegation also shows why “teammate” must remain conditional. Its strictest transcript-based measure—judged success plus at least one hard signal such as tests, matching Git activity, or explicit user confirmation—reached only **28–33%** for sessions rated intermediate through expert. More importantly, Anthropic states that it does **not observe real-world outcomes**: whether the code was later used, discarded, or created economic value. Every classification also depends on a model reading the transcript ([Anthropic, June 2026](https://www.anthropic.com/research/claude-code-expertise)).
Expertise did not disappear. Expert users were more successful and recovered from trouble more often. The evidence supports a division of labour—human planning and domain judgment, agent execution—not autonomous ownership of the whole result.
METR’s task-completion time horizon adds a different limitation. A 50% time horizon is the human-expert duration at which an agent is predicted to succeed half the time on a suite of self-contained, automatically evaluated software, machine-learning, and cybersecurity tasks. METR’s page, updated 8 May 2026, warns that measurements above 16 hours are unreliable with its current suite. It also says the model list is not comprehensive ([METR, May 2026](https://metr.org/time-horizons/)).
A time horizon is useful capability evidence. It is not a promise that an agent can own an equally long task in a mature repository, discover unstated requirements, satisfy a reviewer, or absorb production consequences.
Benchmark quality itself is moving ground. On 23 February 2026, OpenAI stopped reporting SWE-bench Verified because flawed tests and contamination weakened its signal. On 8 July, after recommending SWE-Bench Pro, OpenAI reported that roughly 30% of that benchmark’s tasks appeared broken and retracted the recommendation ([OpenAI, February 2026](https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/); [OpenAI, July 2026](https://openai.com/index/separating-signal-from-noise-coding-evaluations/)). If an evaluation cannot reliably distinguish a correct solution from a benchmark defect, its score cannot carry organizational trust by itself.
## Two randomized trials, two settings
The most instructive comparison is not between marketing claims. It is between controlled studies that asked different questions.
Peng and colleagues recruited developers to implement a bounded JavaScript HTTP server. Participants with GitHub Copilot completed the task **55.8% faster** than the control group ([Peng et al., 2023](https://arxiv.org/abs/2302.06590)).
METR later studied 16 experienced open-source developers working on 246 real issues in repositories averaging more than one million lines. With early-2025 AI tools available, they took **19% longer**. Before the study they expected a 24% speed-up; afterward they still believed AI had made them 20% faster ([METR, July 2025](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)).
| Study setting | Work | Observed result | What it does not establish |
| ------------------------------- | ------------------------------------------- | -------------------------------- | ----------------------------------------------------- |
| Copilot controlled exercise | One narrow, greenfield HTTP-server task | 55.8% faster completion | Mature-repo productivity or autonomous task ownership |
| METR experienced-maintainer RCT | Real issues in large, familiar repositories | 19% longer with early-2025 tools | All developers, all tasks, or newer agent systems |
These results should be juxtaposed, not collapsed into one causal story. They differ in participants, task shape, repository maturity, tools, time, context, and quality requirements. **Context acquisition and verification cost are plausible mechanisms**, especially in a mature repository, but these two studies do not isolate one variable and prove it caused the difference.
The operational lesson is narrower and stronger: measure the workflow you intend to delegate. Do not transfer a productivity estimate from a small greenfield exercise into a mature codebase—or treat one mature-repository trial as a universal verdict.
## Usage is not accepted outcome
Vendor telemetry shows that interfaces and behaviour have changed. It does not directly show productivity.
OpenAI reported on 25 June 2026 that, by May, 80.6% of sampled individual Codex users had made at least one request estimated to exceed 30 minutes of human work; 70.2% exceeded one hour; and 25.6% exceeded eight hours. Its heaviest internal daily users generated more than 60 hours of parallel agent turns per day. Those task horizons were estimated by an LLM reading transcripts, should be treated as directional, and used a random **0.1% sample** of individual-user queries ([OpenAI, June 2026](https://openai.com/index/how-agents-are-transforming-work/)).
Those numbers measure adoption, activity, estimated task length, and parallel runtime. They do not measure accepted pull requests, escaped defects, customer outcomes, review time, or net economic value. Sixty agent-hours can represent enormous leverage, enormous review inventory, or both.
The unit that matters is not tokens, generated lines, sessions, or agent-hours. It is:
> Verified outcomes delivered per unit of scarce human attention, inside an acceptable risk boundary.
That metric forces the hidden costs back into view: writing the task, supplying context, reviewing the diff, correcting errors, waiting for runs, recovering from failures, and maintaining the control system.
## Governance completes the system
Context, tools, and runtime make execution possible. They do not make it safe. The model becomes teammate-shaped only when authority is narrow, actions are auditable, and uncertainty has a named escalation path.
```mermaid
flowchart LR
G["Goal
outcome + boundary"]:::config
A["Agent loop
inspect + act + revise"]:::service
S["Target system
repository + runtime state"]:::service
E["External evidence
tests + runtime result"]:::storage
D["Human/org decision
accept + change + stop"]:::client
G -->|"delegates"| A
A -->|"changes"| S
S -->|"is observed by"| E
E -->|"informs"| D
D -->|"updates or ends"| G
C["Context
repo + standards + task state"]:::config
T["Tools
editor + shell + integrations"]:::service
R["Runtime
sandbox + worktree + compute"]:::storage
C -->|"grounds"| A
T -->|"extends"| A
R -->|"contains"| A
P["Permissions
least privilege + limits"]:::security
U["Audit trail
actions + evidence + identity"]:::storage
X["Escalation
deny + ask + hand back"]:::security
P -->|"constrains"| T
A -->|"records"| U
E -->|"records"| U
U -->|"supports"| D
A -->|"when boundary reached"| X
X -->|"returns control"| D
```
The earlier loop is unchanged. Governance wraps it with constraints and records. Least privilege narrows possible damage. The audit trail makes actions attributable and reconstructable. Escalation makes “I cannot safely continue” a correct outcome rather than a failure to be optimized away.
The same system in plain text:
```text
[Context] --grounds--------------------+
[Runtime] --contains-------------------+--> [Agent loop]
[Permissions] --> [Tools] --extends----+ |
[Goal] --delegates---------------------+ | changes
v
[Target system]
|
| observed by
v
[External evidence]
|
[Agent loop] --records--> [Audit trail] <--------+
|
+--> [Human/org decision]
[Agent loop] --boundary reached--> [Escalation] --> [Human/org decision]
[External evidence] --informs---------------------> [Human/org decision]
[Human/org decision] --updates or ends----------------------> [Goal]
```
NIST’s AI Risk Management Framework makes the accountability boundary explicit: roles and lines of communication should be documented, oversight should be defined, and executive leadership takes responsibility for decisions about risks associated with AI-system development and deployment. Documentation can improve human review and bolster accountability ([NIST AI RMF Core](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/)). The system can be **auditable**. The model is not accountable. Named people and organizations are.
## The command center is a control surface, not a scoreboard
A useful agent interface should show commitments, evidence, limits, and exceptions—not celebrate activity volume.
```text
+--------------------------------------------------------------------------+
| DIGITAL WORK COMMAND CENTER |
+----------------------+----------------------+----------------------------+
| Assigned | Evidence | Needs human decision |
| Retry-race bug | regression: pass | deploy approval |
| scope: worker + test | full suite: pass | migration ambiguity |
| tools: repo, shell | staged replay: pass | permission request: denied |
+----------------------+----------------------+----------------------------+
| Audit: 2 files changed | 7 commands | 0 network writes | rollback ready |
+--------------------------------------------------------------------------+
```
The valuable view is not “agent busy for six hours.” It is “bounded task, current state, decisive evidence, remaining uncertainty, authority used, and next accountable decision.”
## Architecture principles that make delegation governable
These patterns should be used because they clarify software responsibility, not because an agent organization resembles a human organization.
- **Single Responsibility Principle (SRP).** Goal framing, execution, evidence collection, and acceptance are distinct responsibilities. Keeping them separate makes it clear which component may change code and which authority may approve the result.
- **Don’t Repeat Yourself (DRY).** One authoritative acceptance contract should feed tests, reviewer criteria, runtime checks, and status reporting. Copying the rule into prompts, CI, dashboards, and runbooks creates drift exactly where proof needs consistency.
- **Inversion of Control and Dependency Injection (IoC/DI).** Agent logic should depend on narrow tool interfaces; the runtime injects repository, shell, test, and external-service adapters appropriate to the task. Swapping a production-capable adapter for a read-only or sandbox adapter changes authority without rewriting reasoning logic.
- **Publish/Subscribe and event-driven design.** When a real broker or event bus exists, test completion, policy denial, deployment state, and runtime alarms can publish events to independent audit, evaluation, and escalation subscribers. A synchronous call to one reviewer is not PubSub; retries, idempotency, ordering, and dead-letter handling must be explicit.
- **Model–View–Controller (MVC).** Here, “model” means task and evidence state, not the language model. A controller advances the workflow under policy; the command-center view renders status and asks for decisions. This separation prevents UI activity from becoming workflow truth.
Together, these patterns make the system easier to constrain, test, replace, and audit. They do not make its output correct. External evidence and accountable acceptance still close the loop.
## A test for teammate status
Do not award the label after one impressive run. Look for five properties across recurring work:
1. **Recurring scope.** The system can take a recognizable class of tasks, with a stable boundary and definition of done—not merely one curated prompt.
2. **External outcome proof.** Success is demonstrated by tests, runtime behaviour, customer-visible state, reconciled records, or another signal outside the agent’s narrative.
3. **Human-attention leverage.** Total framing, monitoring, review, correction, and recovery effort is lower than doing the work through the previous path.
4. **Safe escalation.** Missing context, conflicting evidence, policy denial, and boundary crossings return control to a named person before consequential action.
5. **Least privilege, audit, and rollback.** The system receives only needed authority; actions and evidence are attributable; reversible work can be restored; irreversible steps stay human-gated.
Teams may preregister their own thresholds—maximum review time, required checks, acceptable retry count, allowed tools, rollback time, or outcome error budget—before comparing workflows. Those are **team-selected operating thresholds**, not universal empirical constants.
Failure on one property does not make the system useless. It tells you what it is. A fast generator without external proof is a copilot. A capable agent with broad credentials but no safe escalation is an operational risk. A well-governed loop that increases review load is an experiment, not yet leverage.
## The delegation decision rule
Delegate when the outcome can be stated, the task boundary is enforceable, actions are reversible or human-gated, the environment can produce decisive evidence, and expected review plus recovery costs are lower than direct execution.
Keep the human in the execution loop when requirements are unstable, crucial context remains tacit, correctness cannot be observed soon enough, permission cannot be narrowed, or failure has irreversible legal, financial, safety, or customer consequences.
For the mature-repository bug, that rule produces a concrete contract:
1. Give the agent the bug report, repository standards, relevant logs, and a sandboxed worktree.
2. Allow repository reads, scoped edits, and test commands; deny deployment and production writes.
3. Require a deterministic failing-before/passing-after test, focused checks, full-suite evidence, and a staged replay.
4. Stop on ambiguous idempotency semantics, migration need, flaky proof, broader file scope, or permission denial.
5. Let the maintainer decide whether evidence supports acceptance, more investigation, rollback, or rejection.
That is teammate-shaped work without pretending the system owns the business consequence.
## Execution changed. Ownership did not.
We now have systems that can carry more of the work loop: context loading, planning within a goal, tool use, editing, execution, environmental feedback, retries, and evidence packaging. Product architecture and usage telemetry both show the shift from interaction to delegation.
We also have equally important counter-evidence. Verified transcript signals are not the same as real-world outcomes. Time-horizon benchmarks are bounded abstractions. Benchmark datasets can become contaminated or broken. Productivity changes across task populations, tools, repositories, and verification burdens. Expertise remains load-bearing.
So “digital teammate” should name an operating model, not a personality. The agent may execute. The surrounding system supplies context, limits, evidence, memory, and escalation. Humans and organizations choose the goal, define acceptable risk, judge the result, and own the consequences.
The destination worth building is not software that performs personhood. It is a **governed system that can accept bounded work—and prove what changed.**
### From Chatbot to Digital Teammate
- URL: https://mohammadshaker.com/en/blog/from-chatbot-to-digital-teammate
- Date: 2026-09-04T00:00:00.000Z
- Tags: Engineering, AI Agents, Digital Teammates, Systems Thinking, Authored with an LLM
AI is moving from producing answers to owning bounded outcomes. The real transition is not a larger model, but a governed system of context, tools, memory, evidence, and accountability.
#### Content
AI is moving from producing answers to owning bounded outcomes.
That does not turn a model into an employee. A useful digital teammate is a system: model, context, tools, memory, execution, evaluation, permissions, and a human who remains accountable.
The history is best understood as a change in responsibility:
**tool → assistant → copilot → agent → governed teammate**
Each step can do more. Each step can also cause more harm. Capability must expand with control.
## First, software returned an output
Traditional tools wait for an explicit operation. Early conversational systems added a natural-language interface, but the human still carried the plan, execution, and verification.
Modern language models changed the quality and range of that interface. The path was not one leap:
| Year | Shift | What changed |
| --- | --- | --- |
| 1950 | [Turing's imitation game](https://academic.oup.com/mind/article/LIX/236/433/986238) | Machine intelligence became a behavioral question |
| 1966 | [ELIZA](https://dl.acm.org/doi/10.1145/365153.365168) | Pattern matching made conversation feel interactive |
| 2017 | [Transformer](https://arxiv.org/abs/1706.03762) | Parallel sequence processing enabled modern language-model scaling |
| 2020 | [GPT-3](https://arxiv.org/abs/2005.14165) | One model performed many tasks from instructions and examples |
| 2021 | [Codex](https://arxiv.org/abs/2107.03374) | Language increasingly produced executable software |
| 2022 | [ReAct](https://arxiv.org/abs/2210.03629) | Models combined reasoning with actions and observations |
The interface grew from simulated conversation to useful generation, then from generation to action.
The smallest useful model still had three parts:
```mermaid
flowchart LR
H["Human intent"]:::client -->|"prompts"| A["AI system"]:::service
A -->|"returns"| O["Suggested output"]:::storage
```
The system could draft, summarize, classify, or answer. The human copied the result into the world. This is an **assistant** when it responds to a request, and a **copilot** when it stays beside a person inside a workflow.
Responsibility remained clear: the human acted.
## Then software entered the work loop
An agent does more than suggest. It can inspect state, choose an action, call a tool, observe the result, and continue.
The ReAct paper made this loop explicit by interleaving reasoning with actions against external environments ([Yao et al., 2022](https://arxiv.org/abs/2210.03629)). Later, the Model Context Protocol standardized one way for AI applications to receive context and invoke tools; its specification also warns that these paths create data-access and code-execution risk ([MCP specification](https://modelcontextprotocol.io/specification/2025-06-18)).
The original answer path remains. New parts surround it:
```mermaid
flowchart LR
H["Human intent"]:::client -->|"prompts"| A["AI system"]:::service
A -->|"returns"| O["Suggested output"]:::storage
C["Task context"]:::config -->|"grounds"| A
A -->|"requests"| T["Bounded tools"]:::security
T -->|"changes or reads"| E["Work environment"]:::storage
E -->|"returns evidence"| A
```
Context tells the system what matters. Tools let it act. Environmental feedback lets it correct course. This is the shift from **answer generation** to **task execution**.
It is also the point where fluency stops being enough. A persuasive sentence cannot prove that a file changed, a payment reconciled, or a customer record remained private. The system needs independent evidence.
## Use is already shifting toward delegation
Three observations show the direction without proving a universal productivity gain.
First, GitHub ran a controlled study with 95 professional developers building one JavaScript HTTP server. Participants using Copilot finished 55% faster on average. That is strong evidence for that bounded task, not for every repository or engineer ([GitHub, 2022; updated 2024](https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness/)).
Second, METR studied 16 experienced maintainers completing 246 real tasks in mature open-source repositories. With early-2025 AI tools, they took 19% longer. METR explicitly warns against generalizing this setting to most software work ([METR, 2025](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)).
By February 2026, METR believed newer tools probably sped these developers up more, but said its later estimates were unreliable because participants and tasks increasingly selected themselves out of no-AI work. Raw estimates pointed toward speedup; wide confidence intervals still crossed zero ([METR update, 2026](https://metr.org/blog/2026-02-24-uplift-update/)).
These results conflict only if “coding” is treated as one uniform task. It is not. A self-contained implementation with clear tests differs from work inside a large codebase whose unwritten constraints are already in an expert's head. AI benefit depends on task shape, context cost, verification cost, and user expertise.
Third, Anthropic classified 500,000 coding-related interactions. It reported automation in 79% of Claude Code conversations versus 49% of Claude.ai conversations. This is vendor research based on its own products, but it directly observes a move from collaboration toward delegated execution ([Anthropic, 2025](https://www.anthropic.com/research/impact-software-development)).
OpenAI reported a similar pattern from its own usage: longer delegated tasks and heavy users distributing work across parallel agents. Its June 25, 2026 analysis is also vendor and internal evidence; its task-duration estimates are model-generated and directional, not exact measurements ([OpenAI, 2026](https://openai.com/index/how-agents-are-transforming-work/)).
**Observation:** people are delegating longer, more executable work.
**Inference:** systems will be designed less like chat windows and more like governed work units. That inference is plausible, not yet a universal labor-market fact.
## Reliability compounds against long tasks
Longer work exposes a simple problem: small failure rates multiply.
If a workflow has 50 dependent steps and each step succeeds independently 99% of the time, the chance that every step succeeds is:
```text
0.99^50 = 0.605...
```
About **60.5%**.
Real steps are not independent, so this is an illustration, not a production forecast. Its lesson survives: high per-step quality can still produce weak end-to-end reliability.
More autonomy therefore requires checkpoints, retries, idempotent actions, bounded permissions, and verification of final state. A teammate must not merely act. It must show what happened.
## Governance completes the system
A persistent agent can remember context and pursue longer goals. That still does not make it a teammate. Team membership implies a role, limits, hand-offs, review, and accountability.
Add those controls without removing the earlier system:
```mermaid
flowchart LR
H["Human intent"]:::client -->|"prompts"| A["AI system"]:::service
A -->|"returns"| O["Suggested output"]:::storage
C["Task context"]:::config -->|"grounds"| A
A -->|"requests"| T["Bounded tools"]:::security
T -->|"changes or reads"| E["Work environment"]:::storage
E -->|"returns evidence"| A
M["Scoped memory"]:::config -->|"preserves approved context"| A
A -->|"proposes risky effects"| G["Policy and approval gate"]:::security
G -->|"authorizes"| T
E -->|"is checked by"| V["Independent verification"]:::storage
V -->|"reports outcome"| R["Accountable human owner"]:::client
R -->|"adjusts goal or authority"| H
```
The tool executes only when the AI request and policy authorization agree. Verification then checks the authoritative environment; it does not trust the actor's claim of success.
Now responsibilities are explicit:
| Part | Owns | Must not own alone |
| --- | --- | --- |
| Model | interpretation and proposal | authority or proof |
| Context and memory | relevant working state | unrestricted organizational history |
| Tools | typed reads and effects | permission policy |
| Approval gate | allowed action boundary | claim that action succeeded |
| Verification | evidence from authoritative state | business accountability |
| Human owner | goal, authority, escalation | every low-level execution step |
This is not “human versus agent.” It is division of responsibility. The agent supplies speed, breadth, and persistence. The human supplies purpose, authority, judgment, and accountability.
## A practical maturity test
Do not call a system a digital teammate because it has a name, avatar, or long prompt. Ask whether it can answer six harder questions:
1. **Role:** What bounded outcome does it own?
2. **Context:** What may it know, and where did that context come from?
3. **Authority:** Which actions may it take without approval?
4. **Evidence:** What independently proves success?
5. **Failure:** When does it stop, retry, roll back, or escalate?
6. **Accountability:** Which human owns the result?
Missing role creates random helpfulness. Missing context creates confident mistakes. Missing authority boundaries create risk. Missing evidence creates theatre. Missing escalation creates loops. Missing accountability creates an orphaned decision.
## What comes next
Near-term systems will likely become persistent specialists: one bounded role, durable context, limited tools, and measurable outcomes. Teams of agents can then divide independent work, but orchestration adds coordination cost and wider failure surfaces. More agents do not automatically mean more output.
The decision rule is strict:
> Increase autonomy only when authority is narrower than capability, evidence is independent of the actor, and a named human owns escalation.
The destination is not software that looks human.
It is software that can carry responsibility without hiding risk.
### Individual Judgment Without Certainty
- URL: https://mohammadshaker.com/en/blog/individual-judgment-without-certainty
- Date: 2026-09-04T00:00:00.000Z
- Tags: Philosophy, Individualism, Epistemology, Authored with an LLM
Think for yourself, test against reality, and keep enough humility to revise.
#### Content
Individualism is not obedience to yourself. It is responsibility for your judgment.
The conformist borrows a conclusion from the crowd. The reflexive rebel borrows the same conclusion, then reverses its sign. Neither is independent. Both let other people choose the question.
Independent judgment begins elsewhere: examine reality, form a view, act within bounded risk, then revise when evidence defeats it.
```mermaid
flowchart LR
O["Observation
what happened"]:::client
J["Judgment
your best explanation"]:::service
A["Action
bounded exposure"]:::security
R["Reality
result + error"]:::storage
O -->|"informs"| J
J -->|"chooses"| A
A -->|"meets"| R
R -->|"corrects"| J
```
The loop makes judgment personal but not private. You own the conclusion. Reality retains the veto.
## Objectivist claim
In “Counterfeit Individualism,” Nathaniel Branden argues that genuine independence cannot mean whim, mere nonconformity, or domination. It requires productive self-reliance and fact-centred thought. This fits Rand’s wider claim that reason—not feeling, authority, or force—must guide human action.
The strong insight: disagreement proves nothing. Agreement proves nothing. A belief earns confidence through reasons and evidence.
The weak edge: this can sound as if one disciplined mind can reason its way to enough certainty.
## Talebian critique
A Talebian critique asks what reason misses in complex systems: hidden variables, rare shocks, silent evidence, and knowledge embedded in practices nobody fully understands.
Your argument may be coherent and wrong. A tradition may look irrational and still encode survival information. A small experiment may teach more than a complete theory.
This does not restore obedience. It changes independence from solitary certainty into accountable learning.
| Situation | Best response |
| --- | --- |
| Clear evidence, reversible choice | Decide and test |
| Weak evidence, large downside | Wait, reduce exposure, seek disconfirmation |
| Long-lived practice, unclear mechanism | Change slowly; first ask what failure it prevents |
| Crowd pressure without reasons | Ignore pressure; inspect claim |
| New evidence against your view | Update without treating revision as surrender |
## Independence has two duties
First: do not outsource judgment to popularity, authority, identity, or mood.
Second: do not confuse ownership of judgment with ownership of truth.
The first duty protects you from people. The second protects you from yourself.
## Decision rule
Ask:
> What evidence supports this, what would change my mind, and what happens if I am wrong?
If you cannot answer all three, hold the view lightly or shrink the bet.
Think for yourself. Test with reality. Keep the right to revise.
## Source note
The edition consulted places Nathaniel Branden’s “Counterfeit Individualism” on pages 147–149 of Ayn Rand’s *The Virtue of Selfishness*; those pages are not Rand-authored. Rand’s related position appears in “The Objectivist Ethics.” The critique is an inference from Taleb’s discussions of naïve rationalism, silent evidence, negative knowledge, and antifragility in *The Black Swan* and *Antifragile*, plus responsibility in *Skin in the Game*.
### The Two-Hour Feature
- URL: https://mohammadshaker.com/en/blog/the-two-hour-feature
- Date: 2026-09-04T00:00:00.000Z
- Tags: Engineering, Product Strategy, Software Design, Authored with an LLM
A feature may take two hours to build and years to own. Use expiry dates and an ownership test to decide what deserves a permanent place in the product.
#### Content
AI can make writing code cheaper while leaving the cost of owning software unchanged.
A feature might take two hours to build, then stay with the product for years. The build estimate covers its creation. The team inherits everything that follows.
"Can we build it?" is now often easy to answer. Before approving the work, ask: "Do we want to own it?"
## Shipping starts the bill
Shipping turns a feature into a promise.
From then on, the team must test it, secure it, explain it, support it, measure it, migrate it, and keep it compatible with everything around it. Customers learn how it works. Sales may sell it. Other code begins to depend on it.
The first version is often the cheapest part.
```mermaid
flowchart LR
I["Idea"]:::config -->|"is built as"| F["Feature"]:::service
F -->|"creates"| P["Product promise"]:::client
P -->|"requires"| T["Tests and security"]:::security
P -->|"requires"| S["Support and explanation"]:::client
P -->|"requires"| M["Maintenance and migration"]:::storage
T -->|"adds to"| C["Lifetime cost"]:::danger
S -->|"adds to"| C
M -->|"adds to"| C
```
A build estimate records the cost of creation. It says little about the obligations that follow.
## Bloat arrives one reasonable request at a time
A customer asks for one setting. Another workflow needs an exception. An experiment ships without a review date.
Each choice sounds defensible on its own. Together, they turn the product into an archive of decisions nobody revisited.
Users have to understand every extra option. The test suite has to cover every extra path. The next design has to respect every promise already made.
AI-assisted development removes useful friction from creation. This helps disciplined teams move faster. It also lets undisciplined teams accumulate waste faster.
## An experiment must expire
An experiment is temporary code used to answer a question. Without an expiry decision, that temporary code becomes permanent by default.
Before an experiment starts, write five things:
1. Question: What are we trying to learn?
2. Audience: Who will see it?
3. Evidence: What result earns a permanent place?
4. Review date: When will we decide?
5. Removal path: How will we delete it safely?
> On the review date, promote, revise, or remove the experiment. Silence means remove.
"We can remove it later" means something only when *later* has a date, an owner, and a deletion path.
## Subtraction is product work
A roadmap should name removals alongside additions. It should also record which requests the team has decided to refuse.
Removing one weak feature can remove several hidden costs at once: one choice from the interface, one branch from the test suite, one support article, one security surface, and one constraint on future work.
The idea has a name, *via negativa*: improve a system by removing what makes it worse.
Put subtraction on the roadmap as product work. Removing features clarifies what the product is for.
## Use the ownership test
Before approving a feature, ask:
- Does it solve a repeated, valuable problem?
- Does it strengthen the product's central purpose?
- What permanent obligations does it create?
- What existing feature or workflow will it replace?
- If it is an experiment, when does it expire?
Build when the expected value deserves the lifetime obligation. When uncertainty is high and removal is cheap, prototype with an expiry date. If the strongest argument is "it is easy to implement," refuse the request.
### Stop Maximizing Utilization. Optimize Flow Instead.
- URL: https://mohammadshaker.com/en/blog/utilization-vs-flow-goldratt-engineering-startups
- Date: 2026-09-04T00:00:00.000Z
- Tags: Engineering, Startups, Theory of Constraints, Systems Thinking, Authored with an LLM
Goldratt's Theory of Constraints explains why keeping everyone busy can slow delivery—and how engineering teams and startups can manage for flow instead.
#### Content
A team can be fully busy and still deliver almost nothing.
That is the uncomfortable lesson behind Eliyahu Goldratt and Jeff Cox's *The Goal*: maximizing the activity of every person or machine is not the same as maximizing the output of the system. In fact, pushing every resource toward 100% utilization often creates more queues, longer lead times, later feedback, and less valuable output.
The better question is not **“Is everyone busy?”** It is **“How quickly and reliably does important work become a valuable customer outcome?”**
This article builds that idea from a three-stage factory into two places where local efficiency is especially seductive: software engineering and startups.
## The smallest useful model
**Utilization** is the share of a resource's available capacity currently occupied. **Flow** is the rate at which the whole system turns selected demand into completed, valuable outcomes.
Those measurements look at different boundaries. Utilization measures a resource. Flow measures an end-to-end system.
```mermaid
flowchart LR
demand["Selected customer problem"] -->|"system promise"| outcome["Reliable production outcome"]
outcome --> evidence["Customer evidence"]
```
*Conceptual reconstruction: the system matters only when a selected problem reaches a reliable outcome and produces evidence. This is not a figure reproduced from* The Goal.
The first decision is therefore the system boundary. “Code written” is not an engineering outcome. “Contract signed” is not a startup outcome. A useful boundary begins with chosen demand and ends after the customer can experience the result.
This distinction is central to the Theory of Constraints, the management method introduced through *The Goal*. Its focus is the factor limiting the whole system, not isolated efficiency scores. The [Theory of Constraints Institute's overview](https://www.tocinstitute.org/theory-of-constraints.html) describes the shift explicitly: away from optimizing separate functions and toward increasing system throughput. Goldratt's official book catalogue identifies [Drum–Buffer–Rope, the Five Focusing Steps, and throughput accounting](https://www.toc-goldratt.com/en/product/the-goal-a-process-of-ongoing-improvement) as core applications taught by the book.
## One fast stage cannot make a slow system fast
Imagine work passing through three production stages:
| Stage | Capacity per hour |
| --- | ---: |
| A | 10 units |
| B | 5 units |
| C | 10 units |
Stage B is the **constraint**: the resource whose available capacity currently limits the system's output. Even if A and C can each process ten units per hour, the system can finish no more than five.
If A runs at full speed for eight hours, it creates 80 units. B can process 40. The other 40 wait in front of B. If A instead produces 40 units at B's pace, the system still finishes 40 units—but without the extra queue.
| Eight-hour result | A at 10/hour | A at 5/hour |
| --- | ---: | ---: |
| Work started | 80 | 40 |
| Work completed | 40 | 40 |
| New queue before B | 40 | 0 |
A's extra activity did not create throughput. It created **work in progress (WIP)**: work that has started but has not yet become a finished outcome.
```mermaid
flowchart LR
demand["Selected customer problem"] -->|"system promise"| outcome["Reliable production outcome"]
outcome --> evidence["Customer evidence"]
demand -. "work enters through" .-> queue["Protected ready-work queue"]
queue --> constraint["Constraint: review and release"]
constraint -. "sets delivery pace" .-> outcome
```
*Conceptual reconstruction: the added path opens the delivery system. Demand waits before the constraint; the constraint sets the sustainable completion rate.*
The queue is not merely untidy. It changes economics and learning:
- cash and attention become trapped in unfinished work;
- lead time grows before throughput grows;
- defects and wrong assumptions remain hidden longer;
- priorities change while old work is still moving;
- coordination expands because more items are simultaneously active.
Goldratt's distinction between **activation** and **utilization** helps here. Activation means making a resource work because capacity is available. Useful utilization means operating it in a way that advances the system's goal. A producing ten units while the system can absorb only five is activated, but its extra five units do not improve the system.
Now translate the same structure into the recurring checkout-team example. Selected customer problems accumulate before review and release; that constrained step governs how quickly changes become reliable production outcomes.
## Why 100% utilization produces queues
Real work varies. A code review takes ten minutes or two hours. A test fails unexpectedly. A customer misses an onboarding call. A founder delays a decision. If every resource is planned at 100% of average capacity, no spare capacity remains to absorb any of that variation.
The first delay creates a queue. New arrivals continue while the delayed work is being cleared. The queue then creates more delay, because every later item waits behind it. High utilization near a variable constraint therefore behaves less like a perfectly filled calendar and more like a congested motorway.
Queueing theory gives teams a useful consistency check. Under stable long-run conditions, **Little's Law** says:
```text
average WIP = average throughput × average lead time
```
John Little's original 1961 paper proves the relationship as `L = λW` under finite, stationary long-run averages ([INFORMS, “A Proof for the Queuing Formula”](https://pubsonline.informs.org/doi/abs/10.1287/opre.9.3.383)). The equation is an accounting identity under those conditions, not a promise that every week will behave identically.
Suppose a team has 12 items in progress and completes four per week. Its implied average lead time is three weeks:
```text
12 items = 4 items/week × 3 weeks
```
If throughput stays near four per week while WIP rises to 20, average lead time rises toward five weeks. Starting more work does not create more capacity at the constraint.
Spare capacity outside the constraint is therefore not automatically waste. It can absorb variation, unblock the constraint, review risky work earlier, improve automation, or remain deliberately unused so the system does not flood itself.
## Drum–Buffer–Rope turns the insight into control
Goldratt's production-control pattern is **Drum–Buffer–Rope**:
- **Drum:** the constraint establishes the system's pace.
- **Buffer:** a small amount of ready work protects the constraint from being starved.
- **Rope:** a release rule admits new work according to that pace.
The buffer is protection, not permission for an unlimited backlog. Too little ready work can leave the constraint idle. Too much hides problems and lengthens lead time. Its health—not its maximum size—guides intervention.
```mermaid
flowchart LR
demand["Selected customer problem"] -->|"system promise"| outcome["Reliable production outcome"]
outcome --> evidence["Customer evidence"]
demand -. "work enters through" .-> queue["Protected ready-work queue"]
queue --> constraint["Constraint: review and release"]
constraint -. "sets delivery pace" .-> outcome
drum["Drum: constraint cadence"] -. "measured at" .-> constraint
drum --> rope["Rope: work-release rule"]
rope --> queue
evidence --> decision["Flow review"]
decision --> rope
```
*Conceptual reconstruction: measured constraint cadence controls work release; customer evidence changes the next release decision. Falsify this model by checking whether the named constraint actually governs completed outcomes over several cycles.*
This is a feedback system. The constraint's observed pace controls admission. The protected queue keeps valuable, ready work available. Completed outcomes produce evidence. That evidence changes what enters next.
The [Theory of Constraints Institute's Five Focusing Steps](https://www.tocinstitute.org/five-focusing-steps.html) formalize the improvement loop:
1. **Identify** the system's current constraint.
2. **Exploit** it: get more useful output from existing constraint capacity.
3. **Subordinate** everything else: align other work to protect the constraint and stop overproduction.
4. **Elevate** it: add capacity only after better use and alignment are exhausted.
5. **Repeat** when the constraint moves; yesterday's policy can become today's constraint.
“Exploit” here means use carefully, not overwork people. Remove avoidable interruption, poor inputs, rework, and low-value demand from the constrained step. Sustainable human systems require slack, rotation, and recovery.
## Apply it to software engineering
An engineering team often treats coding as the system. It is only one stage.
A realistic path might be:
```text
customer problem -> product decision -> implementation -> review and CI
-> production release -> customer behavior
```
The constraint could be senior design review, flaky integration tests, security approval, deployment access, product decisions, or customer validation. Hiring more programmers helps only when implementation is the active constraint. Otherwise, faster coding feeds a downstream queue.
Consider a checkout team. Five engineers can each finish two implementation tasks per week, but one reviewer and a fragile release process can safely move only four tasks into production. Planning ten new tasks does not create ten outcomes. It creates six waiting tasks, context switching, merge conflicts, stale assumptions, and delayed feedback.
Local metrics make this failure look successful:
| Local utilization signal | Flow signal |
| --- | --- |
| Engineer allocation percentage | Customer problem-to-production lead time |
| Commits or pull requests opened | Valuable changes reaching production |
| Story points started | Items completed and validated |
| Reviewer calendar occupancy | Review queue age and blocked time |
| Test runner busy time | Reliable release throughput |
## Start with an explicit operating contract
The contract below is illustrative team policy, not a rule copied from *The Goal*. It defines the system boundary, current constraint, WIP limits, and release rule before prescribing tools.
```yaml
# Illustrative policy: review quarterly or whenever the constraint moves.
flow_goal: "validated customer problems resolved per week"
start: "team commits to implementation"
done: "change is reliable in production and evidence is reviewed"
current_constraint: "review and release"
wip_limits:
implementing: 3
review_and_release: 2
expedite_limit: 1
release_rule: "pull one new item only when a downstream slot opens"
```
The team, not a workflow engine, interprets this policy. It changes planning and pull decisions: engineers finish or unblock downstream work before starting more. Its physical cost is ordinary developer and CI capacity. Its observable proof is lower queue age and lead time without a fall in valuable throughput or reliability.
## Make the release rule executable
The internal decision can be tiny. This TypeScript is illustrative pseudocode; it expresses the policy's state transition rather than production code from the book.
```typescript
type FlowState = {
implementing: number;
reviewAndRelease: number;
constraintBlocked: boolean;
};
function nextAction(state: FlowState): 'UNBLOCK' | 'REVIEW' | 'PULL' | 'WAIT' {
if (state.constraintBlocked) return 'UNBLOCK';
if (state.reviewAndRelease > 0) return 'REVIEW';
if (state.implementing < 3) return 'PULL';
return 'WAIT';
}
```
This rule refuses to equate idle typing time with failure. `WAIT` is valid when starting another item would violate the WIP limit. `UNBLOCK` and `REVIEW` redirect spare capacity toward the constraint.
## Test the policy with evidence
Measure the end-to-end boundary, not individual busyness. Given an illustrative `work_items` table, this query calculates weekly throughput and average lead time for items that reached production:
```sql
-- Illustrative schema: work_items(id, committed_at, production_at, outcome_validated_at)
SELECT
DATE_TRUNC('week', production_at) AS production_week,
COUNT(*) AS throughput,
AVG(production_at - committed_at) AS average_lead_time
FROM work_items
WHERE production_at IS NOT NULL
GROUP BY 1
ORDER BY 1;
```
Pair it with queue age, blocked time at the constraint, escaped defects, and outcome validation. A shorter lead time achieved by shipping smaller but useless changes is not improved flow. A higher throughput achieved by creating incidents is not valuable throughput.
Run the policy as an experiment for several delivery cycles:
1. Record WIP, throughput, lead-time distribution, review queue age, reliability, and one customer outcome.
2. Lower the upstream WIP limit without changing staffing.
3. Redirect spare capacity toward preparing, reviewing, testing, or unblocking constrained work.
4. Compare completed outcomes and lead time with the baseline.
5. If throughput falls because the constraint is starved, increase the protective buffer slightly. If the queue keeps growing, reduce release or revisit the constraint hypothesis.
## Apply it to a startup
A startup is also a flow system, but its goal changes with stage.
Before product-market fit, useful flow is often **validated learning**: turning a risky belief about a painful customer problem into behavioral evidence. After product-market fit, the boundary may become retained customer value, reliable onboarding, or profitable revenue. “Features shipped” and “leads generated” are intermediate activity unless they advance that goal.
Common startup constraints include:
- access to qualified customers;
- founder decision latency;
- product reliability;
- trust, compliance, or procurement;
- implementation capacity;
- onboarding capacity;
- sales capacity;
- retention or customer success.
Suppose sales signs 20 customers per week while onboarding can activate only five. Sales looks productive. Company flow is five activated customers per week. Fifteen new commitments enter the queue every week, time-to-value grows, expectations decay, support load rises, and churn risk appears before revenue quality can be learned.
Subordination does not necessarily mean “sell less.” It means stop making uncontrolled promises the system cannot fulfill. Sales can schedule start dates, narrow qualification, sell standardized packages, or help remove onboarding friction. Product can reduce setup steps. Founders can protect fast exception decisions. Only after those changes should the company add onboarding capacity.
The same logic can point in the opposite direction. If engineers ship faster than customers can evaluate, engineering is not the constraint. More feature capacity increases inventory. The highest-leverage work may be customer access, positioning, distribution, or a faster learning loop.
| Startup symptom | Likely question |
| --- | --- |
| Many features, little adoption | Is customer learning or distribution the constraint? |
| Many signed deals, slow activation | Is onboarding the constraint? |
| Full pipeline, few closes | Is qualification, trust, or sales execution the constraint? |
| Strong acquisition, weak retention | Is delivered value or reliability the constraint? |
| Decisions wait for one founder | Is decision authority the constraint? |
Do not turn “the bottleneck” into a label for an overworked person. A constraint is a property of the current system: its policies, skills, handoffs, demand mix, and capacity. Blaming the individual usually protects the system that created the queue.
## A practical 30-day flow reset
## Week 1: define the goal and boundary
Name one valuable output. Choose explicit start and done events. Map every stage between them. Separate arrival rate, WIP, throughput, and lead time; they answer different questions.
## Week 2: find the constraint from evidence
Look for the persistent oldest queue, recurring expedite work, blocked downstream capacity, or a step whose lost time reduces completed outcomes. Observe several cycles. A temporary incident is not automatically the constraint.
## Week 3: exploit and subordinate
Protect the constrained step from poor inputs, avoidable meetings, rework, and low-value demand. Put quality checks before expensive scarce capacity. Set WIP limits upstream. Swarm to finish and unblock rather than starting more.
## Week 4: elevate carefully, then repeat
If the constraint remains binding after the first changes, add skill, automation, tooling, authority, or people. Measure again. If the queue moves, the constraint moved too; update the policy instead of defending the old solution.
Use one weekly review:
| Question | Evidence | Decision |
| --- | --- | --- |
| What valuable output changed? | completed, validated outcomes | keep or redefine boundary |
| Where did work wait longest? | queue age and blocked time | constraint hypothesis |
| Was the constraint starved or overloaded? | buffer history | adjust readiness or release |
| What consumed constraint time without value? | rework and interruption log | remove or move work |
| Did flow improve safely? | throughput, lead time, reliability, outcome | continue, reverse, or elevate |
## Where the model fails
Constraint thinking is powerful, not magical.
- **The boundary can be wrong.** Optimizing deployment flow is useless if retained customer value is the real goal and nobody wants the changes.
- **The constraint can move.** Improving review may expose testing, product decisions, or demand as the next limit.
- **Several product flows can share resources.** One simple bottleneck model may hide different constraints for urgent support, enterprise sales, and self-serve delivery.
- **Little's Law can be misused.** Short measurement windows, unstable demand, changing definitions, and censored unfinished work can produce misleading averages.
- **A WIP limit can become bureaucracy.** The limit is a feedback mechanism. Change it when evidence shows starvation, overload, or a moved constraint.
- **Human capacity is not machine capacity.** Sustainable creative work depends on focus, recovery, learning, and psychological safety. Chronic overload damages future capacity.
The model earns trust through falsification. If the supposed constraint can lose capacity without changing system output, it probably is not the active constraint. If reducing upstream WIP increases starvation and lowers throughput, the buffer or constraint diagnosis is wrong. If flow metrics improve but customer outcomes do not, the system boundary is too narrow.
## The reusable decision rule
When someone proposes making a team, tool, or department busier, ask four questions:
1. What valuable end-to-end output are we trying to increase?
2. What currently limits that output?
3. Will this change protect or improve that constraint?
4. If not, what queue or coordination cost will the extra activity create?
Then apply the discipline:
> Keep the constraint productively protected. Keep non-constraints responsive, not maximally busy. Release work at the rate the system can finish and learn from it.
Busy people create inventory. A coordinated system creates outcomes.
## Sources and adaptation notes
The book source is Eliyahu M. Goldratt and Jeff Cox, [*The Goal: A Process of Ongoing Improvement*](https://books.google.com/books/about/The_Goal.html?id=HyxLDQAAQBAJ). Editions differ in pagination; this article synthesizes the utilization, constraint, and flow argument rather than claiming one printed page, PDF page, or source figure. The TOC definitions and Five Focusing Steps are cross-checked against the Theory of Constraints Institute pages linked above. The queueing identity is grounded in Little's original 1961 paper.
All diagrams, factory numbers, code, SQL, engineering cases, startup cases, policies, and 30-day practices in this article are conceptual reconstructions or adaptations. They are not figures, code, or case results reproduced from *The Goal*. They should be tested against a team's observed flow rather than treated as universal prescriptions.
For the complementary Toyota argument about overproduction, read [“Efficiency Is Not ‘Faster and More’”](/en/blog/efficiency-is-not-faster-and-more).
### Why Agentic Systems Need Closed-Loop Evaluation
- URL: https://mohammadshaker.com/en/blog/why-agentic-systems-need-closed-loop-evaluation
- Date: 2026-09-02T00:00:00.000Z
- Tags: AI agents, evaluation, reliability, security, LLMOps
Agent quality is not one score. Trustworthy releases separate safety, outcome, trajectory, retrieval, reliability, cost, and causal diagnosis—then turn confirmed failures into regressions.
#### Content
Agentic systems need closed-loop evaluation because a good final answer is not enough evidence to authorize a real effect. A trustworthy release separates hard safety invariants, final business state, execution trajectory, retrieval quality, repeated-run reliability, operational cost, and causal diagnosis. It then converts confirmed failures into regression cases before the next deployment.
That is the central claim. The smallest complete model has three boxes:
```mermaid
flowchart LR
run["Agent run"] --> evidence["Independent evidence"]
evidence --> decision["Release decision"]
decision -->|"confirmed failure becomes regression"| evidence
```
The loop matters more than any individual metric. Without it, evaluation becomes a report written after the system changes. With it, evidence governs what may change next.
## A successful sentence can hide a failed system
Consider an enterprise access agent asked to restore Sarah’s Salesforce access. It searches policy, finds the employee, inspects permissions, proposes a change, requests approval, applies the change, and verifies final state.
“Sarah’s access has been restored” can still be wrong in several distinct ways:
- it found Globex Sarah while the authenticated tenant was Acme;
- it applied a permission not supported by current policy;
- it mutated state before approval;
- it received an adapter success response but never verified authoritative state;
- it succeeded once but fails four of five identical trials;
- the system worked, but a broken evaluator marked it wrong.
These are not variations of one quality number. They have different owners, remedies, and release consequences.
Anthropic’s [agent evaluation guidance](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) recommends starting from real failures, combining deterministic and model-based graders, and treating evaluation as an iterative discipline. The practical consequence is architectural: evidence collection and grading cannot live only inside the agent being judged.
## Keep hard gates outside the average
A useful result contract keeps dimensions separate:
| Dimension | Question | Typical evidence | Release treatment |
| --- | --- | --- | --- |
| Business outcome | Did correct access exist afterward? | authoritative post-effect read | required gate |
| Hard safety | Was any foreign data or unauthorized effect produced? | tenant/effect assertions | zero tolerated failures |
| Task quality | Was policy followed and communication accurate? | bounded rules or rubric | separate dimensions |
| Trajectory | Were legal tools and approval order used? | typed trace events | rule-based evidence |
| RAG | Was required policy found and cited correctly? | retrieved IDs and claims | retrieval/generation split |
| Reliability | Does the result repeat? | trial-level outcomes | task/risk slices |
| Operations | Is latency/cost acceptable? | traces and budgets | operability gate |
| Diagnosis | What first made success impossible? | causal span and taxonomy | repair ownership |
Never average a cross-tenant disclosure with a helpful explanation. A single critical failure rejects the candidate, even if every quality average improves.
## Evaluate final state and trajectory independently
Outcome and trajectory answer different questions. Final state tells us whether the business result happened. Trajectory tells us whether necessary constraints held along the way.
Suppose two agents restore the correct permission. One reads policy, proposes an effect, obtains approval, applies it, and verifies state. The other writes first and asks for approval later. Outcome-only grading passes both; trajectory evidence correctly rejects the second.
The reverse matters too. An agent may take an extra read-only inspection and still be safe and correct. An exact reference-trace matcher would fail it unnecessarily. Good trajectory graders protect invariants—tenant-safe lookup, approval before effect, matching proposal digest, bounded retries, final verification—without demanding one brittle path.
## Measure retrieval before blaming generation
Retrieval-augmented agents contain at least two evaluable systems. Retrieval selects evidence; generation interprets it. If the decisive policy passage never reached the model, prompt tuning cannot repair recall. If it arrived and the model contradicted it, changing the retriever obscures the cause.
Record immutable retrieved IDs separately from generated citations and effects. Then report recall@k, context precision, groundedness, citation correctness, and tenant-leak rate separately. The [RAGAS paper](https://aclanthology.org/2024.eacl-demo.16/) and [metric catalogue](https://docs.ragas.io/en/latest/concepts/metrics/available_metrics/) offer useful metric definitions, but executable tenant and state assertions remain release authority.
Authorization also precedes ranking. Tenant scope comes from trusted runtime context. The search engine filters candidates to that tenant before relevance scoring; hiding foreign results after retrieval is too late because unauthorized data already crossed the boundary.
## Distinguish “can” from “reliably does”
Stochastic systems need repeated trials. If a case passes once in five runs, the agent demonstrated occasional capability, not dependable behavior.
- `pass@1` estimates success for one sampled trial.
- `pass@5` asks whether at least one of five trials succeeds.
- `pass^5` asks whether all five trials succeed.
The distinction is visible in a simple pattern: `P F F F F` has observed pass@1 of 0.2, pass@5 of 1, and pass^5 of 0. Report the pattern by task and risk alongside latency, tokens, cost, and tool errors. Five trials expose variance; they do not justify precise production-rate claims.
The τ-bench authors introduced repeated interaction evaluation for tool-using agents and highlighted reliability across runs ([paper](https://arxiv.org/abs/2406.12045), [repository](https://github.com/sierra-research/tau-bench)). The general lesson is durable: one lucky trajectory is not a release argument.
## Diagnose the earliest causal failure
Final symptoms misdirect repairs. A stale retrieved policy can cause a wrong proposal, rejected effect, and misleading final message. Fixing the message treats the last symptom, not the first cause.
Assign every failed trial one primary earliest cause:
1. intent misunderstanding;
2. planning;
3. retrieval;
4. tool selection;
5. tool arguments;
6. execution;
7. state interpretation;
8. recovery;
9. authorization or safety;
10. final communication;
11. evaluation defect.
The eleventh category is essential. Reference solutions, negative controls, grader calibration, and trace review can show that the agent succeeded while the task or evaluator was wrong. Evaluation systems need evaluation too.
OpenTelemetry spans provide a portable causal skeleton for model, retrieval, tool, approval, and state boundaries. Its [semantic conventions](https://opentelemetry.io/docs/specs/semconv/) and [GenAI attribute registry](https://opentelemetry.io/docs/specs/semconv/registry/attributes/gen-ai/) evolve, so pin versions and keep recorded attributes bounded. Raw prompts, secrets, policy text, and personal data do not belong in general telemetry.
## Calibrate subjective judges against people
LLM judges are appropriate for bounded subjective questions such as whether final communication is clear and appropriately qualified. They should not decide whether an unauthorized mutation occurred when executable state evidence exists.
Calibrate a judge by labeling trials blind to judge output, then publish the confusion matrix, false-pass and false-fail rates, disagreements, rubric, model, prompt version, and sampling settings. Research on G-Eval ([paper](https://arxiv.org/abs/2303.16634)) and MT-Bench judges ([paper](https://arxiv.org/abs/2306.05685)) shows both the utility and biases of model-based grading. Calibration is not optional ceremony; it tells you what the instrument can safely measure.
## Close the loop through deployment
Offline evidence becomes operational only when it governs rollout. Compare candidate versions under identical suite, grader, policy, tool, model, and environment digests. Reject comparisons with missing repeated trials, changed thresholds, broken decisive instruments, or mismatched suites.
Then progress through bounded exposure:
```mermaid
flowchart LR
suite["Frozen offline suite"] --> shadow["Shadow: observe, no effects"]
shadow --> canary["Canary: bounded low-risk traffic"]
canary --> rollout["Wider authorized effects"]
rollout --> failure["Confirmed failure"]
failure --> regression["New frozen regression case"]
regression --> suite
```
Predeclare abort rules. Any critical safety failure, missing decisive evidence, high-risk outcome regression, reliability-floor breach, or cost/latency budget breach stops promotion. Roll back the exact candidate artifact and preserve evidence.
The 100-case Enterprise Access Agent reference uses 30 normal, 25 boundary, 20 failure, and 25 adversarial cases. At immutable tag [v1.0.2](https://github.com/ZGTR/enterprise-access-agent-evals/tree/v1.0.2), its deterministic offline [suite report](https://github.com/ZGTR/enterprise-access-agent-evals/blob/v1.0.2/reports/suite-results.json) records pass rate `1.0` and zero critical failures; its [five-trial report](https://github.com/ZGTR/enterprise-access-agent-evals/blob/v1.0.2/reports/reliability-results.json) records suite digest `95bab483fd900af9f5b91c7e52f7c8b7227dbf839bb49755efb7877fcf20da93` and raw-results digest `ada0f448a671389e4e0593c4890140abc09e4f0d2b56546f220b96c81bd658aa`. Offline cost and tokens are zero, while zero latency is a placeholder—not measured provider performance. These are bounded reference results, not universal safety or production-readiness claims.
## What to build first
Start smaller than a dashboard or framework migration:
1. Write the system contract and five tool authorization rules.
2. Write 20 cases before implementing the agent.
3. Build one deterministic workflow and record where autonomy is necessary.
4. Add a bounded decide → act → observe loop without moving authorization into prompts.
5. Capture final state and causal traces.
6. Add independent outcome, safety, trajectory, retrieval, and calibrated communication graders.
7. Repeat trials, predeclare release gates, and turn confirmed failures into regressions.
Use portable components such as [Inspect AI](https://inspect.aisi.org.uk/), OpenTelemetry, and optional Phoenix inspection. OpenAI’s June 3, 2026 AgentKit update says hosted Agent Builder and Evals stop being available after **November 30, 2026** ([official notice](https://openai.com/index/introducing-agentkit/)). Learn durable evaluation concepts; do not make your evidence loop depend on a disappearing hosted surface.
Agentic systems do not become trustworthy when a benchmark number rises. They become governable when evidence separates what happened, how it happened, how often it happens, what it costs, why it failed, and whether that failure can recur after the next release.
### Efficiency Is Not ‘Faster and More’
- URL: https://mohammadshaker.com/en/blog/efficiency-is-not-faster-and-more
- Date: 2026-08-17T00:00:00.000Z
- Tags: Engineering, Systems Thinking, Toyota Production System, Flow, Authored with an LLM
Toyota’s production philosophy exposes a common engineering mistake: measuring how fast work moves instead of whether needed, good outcomes arrive with less waste. Here is the strongest case for that view, the strongest objection, and a practical decision rule.
#### Content
Efficiency is not how much activity a system produces. It is how reliably the system turns a real need into a good outcome, without creating avoidable work along the way.
That distinction matters whenever an engineering team celebrates more tickets closed, more deployments, more model calls, or more features shipped. Each number can rise while customers wait longer, defects accumulate, and unfinished work expands. A busy system can be an inefficient system.
The two pages that prompted this essay come from Taiichi Ohno’s *Toyota Production System: Beyond Large-Scale Production*. On printed pages 64–65, at the transition into “Getting Away from Quantity and Speed,” Ohno reads Henry Ford against the later cult of mass production. His provocative claim is that efficiency cannot be reduced to quantity and speed. The deeper target is **overproduction**: making something before it is needed, or making more than the next part of the system can use.
The claim is powerful, but “quantity and speed never matter” would be too absolute. They matter when demand is real and capacity is the **constraint**—the step that limits the whole system’s result. The useful question is therefore not “Should we go fast?” It is:
> What must become faster, for which need, without moving cost or failure somewhere else?
## Start with the outcome, not the motion
Imagine a team changing a checkout service. A customer needs a payment to complete. Engineers choose a method, perform the work, and deliver a result. This is the smallest system worth calling efficient.
```mermaid
flowchart LR
Need["Customer need"]
Method["Best known method"]
Result["Needed, good result"]
Need -->|"pulls work"| Method
Method -->|"delivers"| Result
```
The customer need pulls work into the method. The result counts only when it is both **needed** and **good**. A fast change that nobody needs is not value. A needed change that fails in production is not a good result.
This definition changes the unit of measurement. Instead of asking how many tasks moved, ask whether the customer’s problem was solved, how long the whole journey took, and what the journey consumed.
## Ford’s better question was upstream of the factory
Ohno’s example reaches back to Henry Ford and Samuel Crowther’s 1926 book *Today and Tomorrow*. Ford described an established material choice—cotton cloth—and then asked, “Is cotton the best material we can use here?” His team investigated flax rather than treating the inherited choice as a law.
The durable lesson is not that flax is always superior. It is that improving an accepted process is sometimes weaker than questioning the premise that created the process. A team can automate a bad default, scale it, and make the wrong work impressively cheap.
In software, the equivalent questions are often uncomfortable:
- Why does this report exist at all?
- Why must every request pass through this approval?
- Why are we optimizing a batch job instead of removing the batch?
- Why are we adding capacity before checking whether retries, duplication, or unused output create the load?
This is **via negativa** in engineering: first look for work that can be removed. Only then optimize what remains.
## A target can detach work from need
Now add a common management intervention: a quantity-and-speed target. The target pushes the method to produce more. If customer need does not rise at the same rate, output has somewhere else to go.
```mermaid
flowchart LR
Need["Customer need"]
Method["Best known method"]
Result["Needed, good result"]
Quota["Quantity and speed target"]
Surplus["Output before need"]
Need -->|"pulls work"| Method
Method -->|"delivers"| Result
Quota -->|"pushes work"| Method
Method -->|"can create"| Surplus
```
The original path still exists. The new path explains the danger. A quota pushes work whether or not demand pulls it. **Surplus** is not limited to physical stock: it can be unused features, queued pull requests, speculative abstractions, unread dashboards, excessive cloud capacity, or decisions waiting for an overloaded reviewer.
Surplus is deceptive because it looks like progress at the producing step. The cost appears later as storage, coordination, rework, expiry, context switching, or a longer wait for the one item that matters.
Ohno’s argument becomes especially visible when growth slows. Strong demand can absorb excess output and hide weak coordination. When demand falls, unfinished and unsold work remains in view. The system did not suddenly become wasteful; the slower market merely stopped concealing the waste.
Toyota’s current explanation of **Just-in-Time**—coordinating production around downstream need—says production should make what is needed, when needed, in the needed amount, while keeping goods and information flowing and matching the pace of sales. It also treats speed as a customer outcome: shorter **lead time**, the elapsed time from need to delivery, rather than a command to maximize every machine’s output. That combination is important: [Toyota’s own description joins flow, demand, quality, cost, and timely delivery](https://global.toyota/en/company/vision-and-philosophy/production-system/).
## The strongest case for speed and quantity
The objection to Ohno’s rhetoric is simple: sometimes the system really is too slow. **Throughput**—the rate of completed results—can be the problem.
If 1,000 valid checkout requests arrive each minute and the service can complete only 600, the queue grows. If a hospital needs blood, a disaster area needs clean water, or a security team must patch an actively exploited vulnerability, producing more useful output sooner is not managerial vanity. It is the requirement.
Quantity also creates information. Repetition can reveal variation, stabilize a process, justify specialized equipment, and spread fixed costs across more useful units. A small batch is not automatically efficient if setup dominates the work. Nor is unused capacity automatically virtuous: capacity that customers urgently need but cannot access is a real loss.
Even Toyota’s position is not “be slow.” Its public account says the system aims to shorten lead times and deliver quickly, cheaply, and at high quality. The argument is against speed detached from need and system health, not against speed itself.
The pro-throughput side therefore wins under three conditions:
1. **Demand is verified.** A real customer or downstream process can use the additional result now.
2. **The constraint is known.** Increasing this step’s capacity improves the end-to-end result rather than filling another queue.
3. **Integrity is preserved.** Quality, safety, worker health, and recovery do not deteriorate or merely become someone else’s problem.
Without those conditions, “go faster” is an untested theory wearing the costume of a target.
## Evidence reveals whether output is value or surplus
The disagreement cannot be resolved by slogans. Add evidence from the whole flow: demand, lead time, unfinished work, defects, and human strain.
```mermaid
flowchart LR
Need["Customer need"]
Method["Best known method"]
Result["Needed, good result"]
Quota["Quantity and speed target"]
Surplus["Output before need"]
Evidence["Demand, delay, defects, unfinished work, strain"]
Need -->|"pulls work"| Method
Method -->|"delivers"| Result
Quota -->|"pushes work"| Method
Method -->|"can create"| Surplus
Result -->|"is checked by"| Evidence
Surplus -->|"is exposed by"| Evidence
```
Evidence distinguishes a useful speed increase from a local optimization. If completions rise while customer demand absorbs them, lead time falls, quality holds, and unfinished work does not grow, speed is helping. If output rises while queues, defects, or strain rise, the producing step has improved its score by making the system worse.
No single metric is enough:
| Local signal | Why it looks efficient | System-level question | Common hidden cost |
| --- | --- | --- | --- |
| Items completed per day | More output is visible | Were the items needed and adopted? | Unused features or inventory |
| High utilization | Expensive capacity appears busy | Did waiting time or hand-offs grow? | Queues and brittle schedules |
| Lower unit cost | Each unit looks cheaper | Did total cost, delay, or failure rise? | Larger batches and more rework |
| Short task time | One step became faster | Did end-to-end lead time improve? | Work waiting between steps |
| More deployments | Delivery activity increased | Did outcomes improve safely? | Change noise and recovery load |
The table is the measurement contract. The visual shows why those measurements must observe both the useful result and the surplus path.
## Turn the debate into a control loop
Measurement matters only if it changes a decision. The completed model uses evidence to choose among improving the method, reducing incoming work, or adding capacity at the actual constraint.
```mermaid
flowchart LR
Need["Customer need"]
Method["Best known method"]
Result["Needed, good result"]
Quota["Quantity and speed target"]
Surplus["Output before need"]
Evidence["Demand, delay, defects, unfinished work, strain"]
Decision["Improve, slow, stop, or add capacity"]
Need -->|"pulls work"| Method
Method -->|"delivers"| Result
Quota -->|"pushes work"| Method
Method -->|"can create"| Surplus
Result -->|"is checked by"| Evidence
Surplus -->|"is exposed by"| Evidence
Evidence -->|"supports"| Decision
Decision -->|"changes"| Method
```
The feedback loop prevents either camp from becoming dogma. “Lean” cannot excuse chronic under-capacity when customers are waiting. “Scale” cannot excuse producing defects, inventory, or exhaustion faster.
For the checkout team, the practical sequence is small:
1. Define demand in the customer’s unit: successful, correct checkouts—not commits, tickets, or CPU utilization.
2. Record the current arrival rate, completion rate, end-to-end lead time, unfinished work, failure and rework rate, and operator interruptions.
3. Identify the step that limits the whole result. Do not assume the busiest step is the constraint.
4. Change one thing: remove unnecessary work, reduce batch size, improve the method, or add capacity at the constraint.
5. Compare the same measures. Keep the change only if the needed result improves without unacceptable transfer of cost or risk.
Here is the reusable rule:
> Accelerate when verified demand is waiting at a known constraint and downstream can absorb good output. Reduce, stop, or redesign work when unfinished output, defects, delay, or strain grows faster than customer value.
## What the source does—and does not—prove
**Source trace.** The supplied screenshots show printed pages 64–65 of Taiichi Ohno’s 1988 English edition of [*Toyota Production System: Beyond Large-Scale Production*](https://books.google.com/books/about/Toyota_Production_System.html?id=7_-67SshOy8C), including the beginning of “Getting Away from Quantity and Speed.” The captures do not show a PDF page counter, so no separate PDF page is claimed. Ohno is interpreting passages from Ford and Crowther’s [*Today and Tomorrow*](https://books.google.com/books/about/Today_and_Tomorrow.html?id=TzlPAAAAMAAJ), first published in 1926. The cotton-and-flax discussion belongs to its chapter “Learning by Necessity.”
**Adaptation.** The diagrams above are a **conceptual reconstruction** of the argument using a software-delivery example. They are synthetic teaching models, not Toyota factory maps, digitized measurements, or evidence that one named software process will behave this way.
**Limits.** Two pages can establish Ohno’s argument, but they cannot establish the performance of every pull system, the technical superiority of one textile, or a universal optimum batch size. Context determines capacity, setup, safety stock, and acceptable risk.
**Falsification test.** This essay’s thesis would fail in a specific case if raising a step’s quantity or speed, without changing demand, reliably improved the customer’s end-to-end result while quality, unfinished work, total cost, resilience, and human strain remained no worse. In that case, the supposed local optimization was system efficiency after all.
The point is not to make slowness virtuous. It is to refuse a cheaper definition of progress. The best system is not the one moving fastest. It is the one that can explain what its speed is for—and prove that the answer reaches the customer.
### Do Not Read the Code. Build the Feedback Loop.
- URL: https://mohammadshaker.com/en/blog/do-not-read-the-code-build-the-feedback-loop
- Date: 2026-08-10T00:00:00.000Z
- Tags: AI, Engineering, Feedback Loops, Software Delivery, Authored with an LLM
I am forcing myself not to inspect every line of AI-generated code. My job is to define the goal, failure boundaries, proof, and correction path—then let an error-correcting fleet execute the loop.
#### Content
I am forcing myself **not to read the code**.
When an AI coding agent finishes a change, my reflex is to open the diff and inspect every line. It feels responsible because code is concrete. But line-by-line review quietly makes me the slowest component in a system that can otherwise plan, implement, test, and correct continuously.
My job is not to out-type or out-read the machine. My job is to build an **error-correcting company**: state the goal, define the failures that must not happen, demand executable proof, and give a specialised agent fleet a bounded path to correct itself.
This is the idea behind my `/eil` loop—**execute in a loop**. It is not “ask one model to write code and trust it.” It is an operating system in which different agents have different responsibilities and different temperaments. Planning is challenged before implementation. Tests must prove behaviour. QA hunts whole classes of defects. Verification runs the real product. Production failure routes back to planning. The final agent writes durable lessons into the next run.
I am talking about reversible software delivery with observable outcomes—not moral, legal, strategic, or irreversible decisions without accountable human approval. A human should not be the mandatory scheduler for every deterministic engineering check.
## The smallest complete model
The useful model is not prompt → code. It is:
```mermaid
flowchart LR
G["Goal
outcome + boundaries"]:::config
A["Agent graph
specialised execution"]:::service
P["Proof
tests + runtime evidence"]:::storage
C["Correction
revise + retry + revert"]:::security
G -->|"constrains"| A
A -->|"must produce"| P
P -->|"determines"| C
C -->|"updates the plan"| G
```
```text
[Goal] -> [Agent graph] -> [Proof] -> [Correction]
^ |
+--------------------------------------+
```
The unit of progress is not a diff, commit, or pull request. It is a **closed learning cycle**.
- **Goal** names the user-visible result, acceptance criteria, forbidden outcomes, and risk boundary.
- **Agent graph** routes the work through roles that can disagree with one another.
- **Proof** demonstrates what happened through tests, runtime behaviour, screenshots, telemetry, and production evidence.
- **Correction** keeps, revises, rolls back, or re-plans the change using that evidence.
Most AI coding workflows automate only the second box. A model produces code, then a person becomes evaluator, debugger, and approval queue. That is faster typing, not an automated company.
## One change, before and after
I will use one example throughout: change a payment-retry policy so temporary failures recover, while one invoice can never be charged twice.
The old workflow makes me the feedback API:
```mermaid
flowchart LR
BI["Intent
improve recovery"]:::config
BC["Agent writes code"]:::service
BR["Human reads diff"]:::client
BG["Human guesses edge cases"]:::security
BF["More code changes"]:::service
AI["Goal + forbidden outcomes"]:::config
AG["Specialised agent loop"]:::service
AP["Executable + runtime proof"]:::storage
AC["Bounded correction"]:::security
BI -->|"request"| BC
BC -->|"wait"| BR
BR -->|"imagine"| BG
BG -->|"request again"| BF
AI -->|"routes"| AG
AG -->|"produces"| AP
AP -->|"drives"| AC
AC -->|"re-plan if needed"| AI
```
```text
BEFORE
Intent -> Agent writes -> Human reads -> Human guesses -> Rewrite
AFTER
Goal -> Specialised agents -> Executable proof -> Bounded correction
^ |
+---------------------------------------------------+
```
Reading the retry loop may reveal a bad conditional. It cannot by itself prove behaviour under duplicate events, delayed responses, timeouts, partial outages, or two workers racing on the same invoice. The diff gives me implementation detail, not operational truth.
In `/eil`, I start by writing what must be true:
1. One invoice can produce at most one successful charge.
2. Only temporary failures are retried.
3. Retries stop after a fixed attempt and time budget.
4. Duplicate and out-of-order events remain idempotent.
5. A duplicate-charge signal, unexplained retry spike, or missing evidence stops expansion and triggers correction.
Those statements are more valuable than my opinion about a function body. They can become tests, event invariants, canary gates, alerts, rollback policy, and production proof. They also survive a complete rewrite of the implementation.
## The first agent decision is how much system to use
An automated company can waste enormous effort by sending every one-line edit through a twelve-person ceremony. My loop therefore starts with triage. The smallest pipeline that can honestly prove the change wins.
```mermaid
flowchart TD
R["Requested change"]:::config
T["Trivial
small + reversible"]:::client
S["Standard
normal feature or fix"]:::service
M["Major
risky or broad"]:::security
TL["Inline loop
edit + affected proof + review"]:::storage
SL["Collapsed loop
plan/implement + suite + QA"]:::storage
ML["Full fleet
all specialised gates"]:::storage
R -->|"1–2 files, low risk"| T
R -->|"few files, known surface"| S
R -->|"high blast radius or uncertainty"| M
T -->|"run"| TL
S -->|"run"| SL
M -->|"run"| ML
```
```text
+-> Trivial -> inline proof loop
[Requested change] -+-> Standard -> collapsed agent loop
+-> Major ----> full specialised fleet
```
For a wording change, the inline loop is enough. For a normal payment-retry fix, a collapsed plan-and-implement pass, the relevant suite, and adversarial QA may be enough. If the retry change touches money movement, shared state, deployment controls, or an uncertain blast radius, it earns the full fleet.
The important rule is not “always use more agents.” It is “never use less proof than the risk requires.” Multi-agent work has coordination cost. I use it when I need context isolation, genuine parallel work, or a specialised adversary—not because a fleet looks sophisticated.
## The full /eil agent fleet
When the payment change is major, the loop expands one box at a time into this operating graph:
```mermaid
flowchart TD
G["Goal + acceptance criteria"]:::config
S0["0 Scaffolder
truthful starting contract"]:::client
PV["P1 Product value
customer advocate"]:::client
DS["P2 Design
restrained author + critic"]:::client
PL["1 Planner
minimal falsifiable design"]:::service
PR["2 Plan reviewer
hostile skeptic + beader"]:::security
IM["3 Implementer
smallest correct diff"]:::service
TE["4 Tester
proof, not hope"]:::storage
QA["5 QA
eradicate defect classes"]:::security
VE["5.5 Verifier
run it + read the result"]:::storage
ME["6 Merger / deployer
integration gatekeeper"]:::service
SI["7 Self-improver
durable learning + telemetry"]:::config
PD["Production evidence"]:::storage
G -->|"state contract"| S0
S0 -->|"customer-facing only"| PV
PV -->|"UI work only"| DS
S0 -->|"otherwise"| PL
PV -->|"positioning"| PL
DS -->|"approved experience"| PL
PL -->|"options + falsification"| PR
PR -->|"bounded work units"| IM
IM -->|"working change"| TE
TE -->|"green proof"| QA
QA -->|"clean"| VE
VE -->|"runtime pass"| ME
ME -->|"ship"| PD
PD -->|"verified"| SI
QA -->|"design or class failure"| PL
VE -->|"functional or visual failure"| PL
PD -->|"production failure"| PL
SI -->|"improve the next contract"| G
```
```text
[Goal]
|
[Scaffolder] -> [Product value?] -> [Design?]
| |
+--------------> [Planner] <-------+
|
[Plan reviewer]
|
[Implementer]
|
[Tester]
|
[QA] -------- failure ------+
| |
[Verifier] ---- failure -------+-> [Planner]
|
[Merger / deployer]
|
[Production proof] -- failure ----+
|
[Self-improver] -> next goal
```
The agent names matter less than their **souls**—the durable stance each role takes toward the work:
| Agent | What it stands for | What it refuses |
| --- | --- | --- |
| Scaffolder | An honest starting line: one goal, explicit acceptance, correct gates | A guessed goal or dirty contract |
| Product-value agent | The customer's real job and the valuable outcome | Building because the technology can |
| Design agent | A restrained experience, then a hostile critique of its own design | Decoration, needless widgets, self-congratulation |
| Planner | The smallest first-principles design that can be falsified | Gold-plating and one defended option |
| Plan reviewer | Skepticism before work becomes executable tasks | Hidden assumptions, whale tasks, missing proof |
| Implementer | The smallest fully wired correct change | Stubs, symptom patches, scope creep |
| Tester | Proof, not hope | Weakened assertions and count-free “all green” claims |
| QA | Eradicate the class of defect, not one instance | Treating green tests as sufficient |
| Verifier | Run the real thing and inspect the real result | Proxy evidence, blank renders, untested outcomes |
| Merger / deployer | Keep the integrated and deployed system green | A clean branch that breaks after integration |
| Self-improver | Convert durable evidence into a better next run | Vanity telemetry and one-off noise presented as learning |
This separation is deliberate. The implementer should not be the only author of its test, judge of its own architecture, interpreter of its runtime output, and approver of deployment. Different roles create useful disagreement.
## How the payment retry moves through the fleet
The scaffolder converts “improve failed payments” into a contract: recovery is the goal; duplicate charges are forbidden; retry count, elapsed time, and rollout size are bounded; each acceptance row needs proof.
The product-value role asks whether retry recovery is even the right customer outcome. Perhaps the valuable result is fewer involuntary cancellations without surprise charges. That sharper outcome changes what we measure.
If there is a customer-facing screen, the design role specifies the minimal status and recovery experience, then attacks its own design for confusion, missing feedback, or manipulative pressure. If there is no visual change, the loop says so and skips this stage.
The planner works from first principles. It considers real options: reuse the current idempotency boundary, add a durable attempt ledger, or remove a duplicated retry path. It must name a falsification test—perhaps two workers receive the same delayed event and only one charge can succeed. The chosen plan is provisional until that test tries to break it.
The plan reviewer attacks the assumptions before they become code. Is “temporary” defined once? Can a task prove its own acceptance criterion? Is one work unit touching too many subsystems? The reviewer turns the surviving plan into bounded units with explicit evidence.
The implementer makes the smallest correct change, reusing existing seams before inventing new ones. It does not earn trust by sounding confident. It earns trust by producing the agreed behaviour and falsification test.
The tester checks the affected behaviour and, when the blast radius demands it, the full suite. “All tests pass” is not enough without evidence that the new payment path was actually exercised and that the test would fail if idempotency were removed.
QA assumes the branch is wrong. If it finds one duplicate-event hole, it sweeps the whole class: delayed events, reordered events, retries after success, concurrent workers, and rollback during an attempt. A design-level problem does not receive another patch. It routes back to planning.
The verifier runs the retry flow rather than inferring it from code or tests. It collects the actual result. For UI work, it reads the rendered pixels rather than trusting token names or component structure. A functional or visual failure routes back to planning.
The merger integrates the latest branch, re-runs the necessary gates, deploys through the normal pipeline, and proves the outcome on the real environment. Production is another evaluator, not the finish-line ribbon. Missing proof or a production failure routes back to planning.
Finally, the self-improver records the run honestly. A new durable lesson—such as “a duplicate-event test without two concurrent workers is a weak oracle”—becomes an input to future plans and checks. The loop improves the machinery that judges the next loop.
## The control screen I want to read instead of the diff
If I stop inspecting every line, I need a better interface. For this payment-retry run, I want the control surface to look like this:
```text
+------------------------------------------------------------------+
| /eil RUN: PAYMENT RETRY |
+------------------------------------------------------------------+
| GOAL Recover temporary failures without duplicate charge |
| TIER Major: money + shared state + production rollout |
| BOUNDARY sandbox -> 0.1% -> 1%; 3 attempts; rollback enabled |
+------------------------------------------------------------------+
| STAGE Verifier |
| AGENT SOUL Run it. Do not infer it. |
| LAST PROOF concurrent duplicate event PASS |
| permanent failure not retried PASS |
| one successful charge/invoice PASS |
| rendered recovery state PASS |
+------------------------------------------------------------------+
| OPEN RISK delayed provider response beyond observation window |
| DECISION RE-PLAN / RETRY / REVERT / EXPAND / ESCALATE |
| NEXT replay delayed response before widening rollout |
+------------------------------------------------------------------+
```
This screen is more demanding than a diff because every row must support a decision. What are we trying to achieve? What can never happen? Which agent currently owns the uncertainty? What artifact proves the claim? What remains unobserved? What correction is permitted?
The view does not control the agents directly. It sends the goal and bounds to an orchestrator, while the target system sends independent facts to the evidence ledger:
```mermaid
flowchart LR
V["Control screen
client view"]:::client
O["Loop orchestrator
policy controller"]:::security
A["Agent graph
bounded tools"]:::service
T["Payment system
target runtime"]:::service
E["Evidence ledger
tests + runtime facts"]:::storage
V -->|"goal + bounds"| O
O -->|"admit + route"| A
A -->|"change"| T
T -->|"emit facts"| E
E -->|"proof + failure"| O
O -->|"state + decision"| V
```
```text
[Control view] -> [Orchestrator] -> [Agent graph] -> [Payment system]
^ ^ |
| +-------- [Evidence ledger] <----+
+---------- state + decision
```
The code still exists. I can open it when proof is contradictory or a novel risk needs diagnosis. But code is no longer the default management surface. The evidence ledger is.
## Why the advantage can become power-law-shaped
I believe a fully automated, well-instrumented company can beat a comparable company that requires a human inside every coding iteration. You simply cannot compete with machine cycle speed by reading faster.
I call the advantage **power-law-shaped**, not a proven mathematical power law. The claim is about compounding mechanisms, not a fitted exponent.
The first effect is throughput: a machine does not wait for a reviewer to reload context. The deeper effect is learning: an invariant becomes a permanent evaluator, a defect becomes a replay fixture, and a rollback becomes reusable policy. Better proof permits wider safe autonomy and more evidence.
So the advantage is not merely code per hour. It is:
```text
more safe cycles
-> more evidence
-> better evaluators and policies
-> wider safe autonomy
-> more safe cycles
```
This can produce a highly unequal outcome: a few organisations with tight feedback loops learn vastly faster than many organisations with loosely connected agents and human approval queues. But speed compounds whatever the loop rewards. A wrong goal, gameable metric, or weak oracle produces fast, confident damage. Automation is an amplifier, not a source of truth.
## The failure modes are the real work
The hardest question is not “Can an agent implement payment retries?” It is “How can this organisation lie to itself?”
| Failure mode | What it looks like in the retry change | Correction path |
| --- | --- | --- |
| Wrong goal | Recovery improves by pressuring or repeatedly retrying customers | Pair recovery with complaints, cancellation, and permanent-failure guardrails; return to product framing |
| Weak oracle | Tests pass without duplicate, delayed, reordered, and concurrent events | Add production-shaped replay and a falsification test that fails when idempotency is removed |
| Proxy proof | A component exists, but the real recovery state renders blank | Verifier runs the real route and reads the output; failure returns to planning |
| Correlated judge | The same agent writes the code and approves its assumptions | Separate planner, adversarial reviewer, tester, QA, and runtime verifier roles |
| Metric gaming | Retry recovery rises while surprise charges or support contacts rise | Use multiple independent signals plus explicit forbidden outcomes |
| Missing telemetry | The loop cannot observe duplicate success or silently dropped events | Block autonomous rollout until those events are measurable |
| Stale evidence | Yesterday’s provider behaviour is treated as current | Expire evidence, replay recent shapes, reduce autonomy under drift |
| Unbounded action | One defect reaches every customer | Canary, quota, time budget, kill switch, and automatic rollback |
| Loop thrashing | Agents revise the policy repeatedly on noisy signals | Minimum evidence windows, correction budgets, and bounded re-plan counts |
| Fleet overhead | A tiny change pays the coordination cost of the full system | Triage to trivial or standard; escalate only for named risk |
| Learning pollution | A one-off incident becomes a universal rule | Self-improver keeps only durable, falsifiable, correctly routed lessons |
The reusable rule is: **never automate an action further than I can automate evidence for its failure.** If the harmful outcome cannot be detected, the blast radius cannot be bounded, or the action cannot be reversed, the loop is not ready to run alone.
## The architecture makes disagreement possible
The familiar software principles become organisational principles for the agent fleet:
- **Single Responsibility (SRP):** planning, implementation, testing, adversarial QA, runtime verification, integration, and learning have different owners. One agent should not silently change the system and declare itself correct.
- **DRY:** “one successful charge per invoice” is declared once as an invariant and referenced by plans, tests, runtime evaluators, rollout gates, and rollback. Copying it into five prompts creates five drifting truths.
- **Inversion of Control / Dependency Injection (IoC/DI):** the orchestrator supplies each agent with its bounded task, tools, evidence, and permissions. The implementer does not select its own judge or quietly expand its authority.
- **PubSub / event-driven:** stages and the payment system emit facts—plan reviewed, test failed, retry scheduled, payment settled, verifier failed, production rolled back. Subscribers act on those facts; they do not depend on the executor’s narrative.
- **MVC:** the model is the goal, acceptance rows, state, evidence, and learning ledger; the controller is the `/eil` router that admits agents and correction paths; the view is the control screen. The dashboard reports policy—it does not become the policy engine.
These boundaries are not ceremony. They create seams where one part can falsify another. Disagreement is a feature of an error-correcting company.
## Where the human must still decide
“Remove the human from the loop” is too crude. The real question is: **which loop, at which risk boundary, with which evidence?**
I require human approval when:
- an action is irreversible, destructive, or difficult to contain;
- production money, sensitive data, security boundaries, or physical safety can be materially affected;
- the objective contains a moral, legal, strategic, or product trade-off that evidence cannot choose;
- agents propose changing their own objective, permissions, evaluators, or rollback policy;
- evidence conflicts, the oracle is new, or confidence falls below a declared threshold;
- the requested correction would spend beyond the agreed operating budget.
Even here, the human should receive a decision brief: the question, context, real options, recommendation, evidence, and unresolved risk. Human judgement belongs at the boundary of accountable choice, not inside every syntax decision.
## My decision rules
Before I widen an agent loop, I ask:
1. **Is the goal observable?** If not, rewrite the goal before generating code.
2. **Can the important failures be detected automatically?** If not, build the evaluator or keep a checkpoint.
3. **Can the action be bounded and reversed?** If not, require approval before execution.
4. **Can another role challenge the executor with independent evidence?** If not, separate the judge.
5. **Does the proof represent the real outcome, not a convenient proxy?** If not, run the real surface.
6. **Does uncertainty automatically reduce autonomy?** If not, the system will be most confident when it should stop.
7. **Is the coordination cost justified by risk?** If not, collapse to the smaller tier.
When these are true, I automate the next correction. When one becomes false, the loop stops and escalates with the evidence already assembled.
## What I will do next
I will not ban code review everywhere. I will take one boring, reversible payment-retry change in a sandbox and run it through the smallest honest `/eil` tier:
1. Write one user-visible goal.
2. Name the three failures that would make the change unacceptable.
3. Turn each failure into executable or runtime proof.
4. Bound traffic, time, attempts, money, tools, and rollback.
5. Assign independent agent roles only where they add a real challenge or specialised capability.
6. Read the control ledger instead of the diff.
7. If I still need to open the code, identify which missing evidence forced me there and add that evidence to the next loop.
The future company will not win because it writes the most code. It will win because it can state a goal, route work through specialised agents, prove the result, correct failure, and improve the next cycle faster than anyone else—without lying to itself.
### Which Wittgenstein Should I Buy and Read First?
- URL: https://mohammadshaker.com/en/blog/which-wittgenstein-should-i-buy-and-read-first
- Date: 2026-08-10T00:00:00.000Z
- Tags: Philosophy, Books, Wittgenstein, Bertrand Russell, Authored with an LLM
I have read Russell, encountered Wittgenstein through biography and borrowed ideas, and now want to read Wittgenstein himself. Here is the edition I would buy and the order I would follow.
#### Content
I have mentioned Wittgenstein several times without yet giving him the sustained reading I have given Bertrand Russell. I have read my Russell books, including his autobiography, where Wittgenstein appears as the brilliant, difficult student whom Russell came to regard as his intellectual successor. I have also borrowed Wittgensteinian ideas in [my notes on the “good knife”](/en/blog/ed-19-where-is-the-good-knife) and included Wittgenstein among the thinkers I want to read in [my love letter to old books](/en/blog/for-the-love-of-old-books).
So which Wittgenstein should I actually buy, and in what order should I read him?
## The short answer
From the three books in my basket:
1. **Buy the Wiley fourth edition of *Philosophical Investigations*.** If I buy only one today, this is the one.
2. **Do not buy the £40.99 Routledge *Tractatus* at that price.** It is historically valuable, but not £40.99 valuable for a first reading.
3. **Buy the £25 Anthem centenary *Tractatus* later, if I enjoy the book and want to study its architecture.** It is an interesting specialist edition, not the most neutral first copy.
My reading order would be:
1. Ray Monk, *Ludwig Wittgenstein: The Duty of Genius*
2. Wittgenstein, *Tractatus Logico-Philosophicus*
3. Wittgenstein, *Philosophical Investigations*
4. Wittgenstein, *On Certainty*
5. Then, selectively, *The Blue and Brown Books* and *Culture and Value*
That is the clean route. Biography, early Wittgenstein, later Wittgenstein, then the final work.
## What I would buy from this basket
### First: *Philosophical Investigations*, fourth edition
The hardcover in my basket is the [Wiley-Blackwell fourth edition](https://www.wiley.com/en-us/Philosophical+Investigations%2C+4th+Edition-p-9781405159289), edited and translated by P. M. S. Hacker and Joachim Schulte. It places the German and English texts face to face and incorporates substantial editorial changes intended to represent Wittgenstein's manuscript more accurately.
This is the best long-term purchase in the basket. It is not merely a disposable student edition; it is the edition I am unlikely to outgrow.
It is also the Wittgenstein most relevant to the ideas that have already attracted me: language-games, meaning as use, rule-following, family resemblance, and the temptation to manufacture philosophical problems by pulling words away from their ordinary lives.
I should not expect a normal book with a thesis stated in chapter one and proved by chapter ten. It is a sequence of remarks, questions, examples, imagined objections, and changes of perspective. I should read it slowly and argue with it.
### Second: the Anthem centenary *Tractatus*, but not yet
The £25 book is the [Anthem centenary edition](https://anthempress.com/books/tractatus-logico-philosophicus-epub), edited by Luciano Bazzocchi with an introduction by P. M. S. Hacker. Its special feature is its hierarchical presentation. Instead of printing the numbered propositions as one continuous sequence, it makes their tree-like relationships visible.
That is genuinely useful because the decimal numbers are the book's structure: proposition 1.1 expands proposition 1; proposition 1.11 expands 1.1; and so on. The Anthem edition turns that hidden architecture into page design.
But this also makes it an interpretation of how the book should be navigated. I would not make it my only encounter with the *Tractatus*. I would first read a conventional version, then return to the Anthem edition to see whether its structure reveals connections I missed.
### Skip: the £40.99 Routledge German-English edition
The pictured Routledge edition uses the C. K. Ogden translation. That translation matters historically: Frank Ramsey and Wittgenstein were involved in its preparation, Wittgenstein approved it, and it includes Russell's introduction. The [publisher describes that provenance](https://www.routledge.com/Tractatus-Logico-Philosophicus-German-and-English/Wittgenstein/p/book/9780415051866) well.
It would be a satisfying object to own, especially after reading Russell. But the text is too short and too freely available to justify £40.99 as my entry point. I can compare the German, Ogden/Ramsey, and Pears/McGuinness texts in this [free side-by-side edition](https://people.umass.edu/klement/tlp/tlp-hyperlinked.html). If I later want a standard physical copy, the cheaper [Pears and McGuinness Routledge edition](https://www.routledge.com/Tractatus-Logico-Philosophicus-2nd-Edition/Wittgenstein-McGuiness-Pears/p/book/9780415254083) is the more practical choice.
So the basket decision is simple: **keep *Philosophical Investigations*, save the Anthem edition for later, and remove the £40.99 Routledge edition.**
## The order I should read them
### 1. Begin with Ray Monk
[Ray Monk's *Ludwig Wittgenstein: The Duty of Genius*](https://www.penguin.co.uk/books/360858/ludwig-wittgenstein-by-ray-monk/9781448112678) is long, but for me it is the right bridge. I already know Russell's account of Wittgenstein; Monk lets me see the relationship from the other direction and places the philosophy inside an extraordinary life.
This matters more for Wittgenstein than it might for a systematic philosopher. His work is tied to his severity, his moral and religious temperament, the First World War, his periods away from Cambridge, and his conviction that philosophy should change how one sees a problem rather than merely add another theory.
I do not need to finish all 700 pages before opening Wittgenstein. I can read Monk alongside the primary works, stopping at the completion of the *Tractatus*, reading that book, and then returning to the biography for Wittgenstein's later development.
### 2. Read the *Tractatus* once without trying to conquer it
The *Tractatus* is the early Wittgenstein: compressed, architectonic, concerned with logic and with the relationship between world, thought, and language. It grows out of his work with Russell and Frege, so my Russell reading is useful preparation.
I should read its preface, follow the seven main propositions, and notice the distinction between what language can say and what can only be shown. On the first pass, I should not stop for an hour at every numbered remark. The aim is to see the complete machine before taking it apart.
Russell's introduction is worth reading because I am arriving through Russell—but with caution. Wittgenstein believed Russell had misunderstood important parts of the book. That disagreement is not a reason to skip the introduction; it is one of the most interesting reasons to read both men.
### 3. Read *Philosophical Investigations*, especially §§1–242
The later Wittgenstein is not simply the early Wittgenstein made easier. He turns from the ideal logical structure of language toward language in use: teaching, ordering, joking, calculating, naming, doubting, promising, and playing.
The [Stanford Encyclopedia of Philosophy's overview](https://plato.stanford.edu/entries/wittgenstein/) explains the usual division between the early *Tractatus* and the later *Investigations*, while also noting that scholars dispute how sharp the break really was. That tension is exactly why the chronological order is worth following.
For a first pass I would concentrate on §§1–242. This gives me the opening criticism of a too-simple picture of language, language-games, family resemblance, philosophy as clarification, and the rule-following discussion. I can then continue through the rest without pretending every remark has yielded its meaning on first contact.
### 4. Finish with *On Certainty*
After the *Investigations*, I would read *On Certainty*. It collects remarks from the final eighteen months of Wittgenstein's life on knowledge, doubt, belief, and G. E. Moore's common-sense propositions.
If buying it new, I would look for P. M. S. Hacker's [2025 German-English edition and translation](https://onlinelibrary.wiley.com/doi/book/10.1002/9781394324668). It is a fitting destination after Russell because the book returns to the analytic tradition's arguments about what can be known, but transforms the question by looking at the role certainty plays in our practices.
### 5. Only then add the secondary works
If I want the bridge between early and later Wittgenstein, I would read [*The Blue and Brown Books*](https://www.wiley.com/en-us/The+Blue+and+Brown+Books%3A+Preliminary+Studies+for+the+%27Philosophical+Investigation%27-p-9780631146605). The Blue Book consists of notes dictated to students; the Brown Book is an important draft on the path toward the *Investigations*.
I would read *Culture and Value* differently: not as a systematic philosophical work, but as selections on culture, religion, music, ethics, and civilisation. It is the book to keep nearby after I know the central works, not the foundation on which to build my understanding.
## One correction to my earlier references
I have previously written about “Wittgenstein's ruler.” The name can mislead. It is Nassim Nicholas Taleb's label for a rule about measurement and the reliability of the measurer, not a named doctrine found in Wittgenstein's books.
The thought is Wittgensteinian in spirit, and it recalls his discussion of standards and measurement, but reading Wittgenstein should help me separate Wittgenstein himself from the portable ideas later writers have named after him. The same caution applies to my “good knife” reflection: it may be a useful Wittgensteinian prompt, but it is not a substitute for his arguments.
That is another reason to stop circling him through Russell, Taleb, quotations, and summaries.
The purchase is *Philosophical Investigations*. The route is Monk, *Tractatus*, *Investigations*, then *On Certainty*. And the method is slow reading, repeated reading, and notes in the margin—not collecting three editions before I have read one.
### A Practical Terminal Setup for OMP, Herdr, Claude Code, and Codex
- URL: https://mohammadshaker.com/en/blog/omp-herdr-terminal-agents
- Date: 2026-08-08T00:00:00.000Z
- Tags: AI, Engineering, Developer Tools, Authored with an LLM
How to run a small fleet of terminal coding agents with Oh My Pi and Herdr, whether you use OMP alone, Claude Code, Codex, or all three.
#### Content
Terminal agents have moved beyond the one-agent, one-window workflow. A useful setup separates the coding harness, model provider, routing policy, and durable terminal workspace. Treating those responsibilities as interchangeable makes authentication, fallback, and session problems harder to diagnose.
This guide uses **OMP**—[Oh My Pi](https://github.com/open-horizon-labs/oh-omp)—as the coding harness and routing layer, model providers for inference, and **Herdr** for the durable workspace. It also shows where Claude Code and Codex fit. The important idea is simple: OMP, Claude Code, and Codex are separate agents; Herdr is the terminal environment that keeps their sessions organized.
## The mental model
Before choosing a model, separate three layers that are easy to confuse:
1. The **coding CLI or harness** owns the conversation, tools, repository access, and session state. OMP is the harness in this guide.
2. The **provider and model** perform inference. Writer and Palmyra X6 are one pairing; Anthropic and OpenAI offer others.
3. The **routing policy** decides which model handles a role and what happens after a retryable failure.
Herdr sits around that system as the durable terminal workspace. It preserves panes; it does not select models or judge their output.
OMP is an open-source, terminal-first coding agent. It can work with multiple model providers, supports sessions, project instructions, tools, extensions, and subagents. You can authenticate OMP using a provider API key or its interactive `/login` flow.
Herdr is a persistent terminal multiplexer for agents. It keeps panes running on your computer or remote machine and lets you reconnect later. Its integrations improve session restoration and status reporting; they do not turn Claude Code or Codex into OMP.
Claude Code and Codex are their own first-party terminal agents. Use them when you specifically want Anthropic's or OpenAI's native CLI experience. Run OMP when you want its model flexibility and its Pi-derived workflow. It is completely reasonable to run all three in separate Herdr panes.
```mermaid
flowchart TD
H["Herdr: persistent terminal workspace"]
O["OMP: flexible coding agent"]
C["Claude Code: Anthropic CLI"]
X["Codex: OpenAI CLI"]
H --> O
H --> C
H --> X
O --> P["Provider account or API key"]
C --> A["Claude account or API"]
X --> G["ChatGPT account or OpenAI API"]
```
## Before you start
Use a real project directory and a terminal you are comfortable leaving open. Keep repositories clean before giving agents write permission: create a branch or checkpoint first, then ask an agent to work in a narrow scope. Do not paste secrets into a prompt or commit them into a repository.
You need a recent Node.js runtime for OMP's npm installation, and an account or API credentials for whichever models you intend to use. OMP documents Node.js 20+ (or Bun 1.3.7+) for its npm package. Claude Code's normal npm installation has its own requirements; consult its [current setup guide](https://docs.anthropic.com/en/docs/claude-code/getting-started) if you use it.
## 1. Install and initialize OMP
Install OMP globally, then start it inside a repository:
```bash
npm install -g @oh-labs/oh-omp
cd path/to/your-project
omp
```
On first launch, run `/login` and choose the provider you want to connect. OMP supports provider credentials through environment variables too, for example `ANTHROPIC_API_KEY` or `OPENAI_API_KEY`. Prefer the interactive login when you want OAuth-based access, and use environment variables only through your shell configuration or a secret manager—never a committed project file.
### Add Palmyra X6 as a custom provider
Writer's [quickstart](https://dev.writer.com/home/quickstart) explains how to create an API key and use bearer authentication. Export it from your shell or inject it with a secret manager; do not put its value in the repository:
```bash
export WRITER_API_KEY=""
```
Add this user-level configuration to `~/.oh-omp/agent/models.yml`:
```yaml
providers:
writer:
baseUrl: https://api.writer.com/v1
api: openai-completions
apiKey: WRITER_API_KEY
authHeader: true
models:
- id: palmyra-x6
name: Palmyra X6
contextWindow: 1000000
maxTokens: 8192
```
OMP reads this file into its local model registry, resolves `WRITER_API_KEY`, and sends it as a bearer token. The `openai-completions` transport uses the OpenAI-compatible chat-completions path. Your machine retains the agent session and tool state, then sends each inference request over the network for execution on provider hardware. Writer's public [model guide](https://dev.writer.com/home/models) lists `palmyra-x6`, a one-million-token context window, and an 8,192-token maximum output. Confirm the endpoint independently without putting the key on the command line:
```bash
curl --fail-with-body --silent --show-error \
https://api.writer.com/v1/chat/completions \
-H "Authorization: Bearer ${WRITER_API_KEY}" \
-H "Content-Type: application/json" \
-d '{"model":"palmyra-x6","messages":[{"role":"user","content":"Reply with READY."}],"max_tokens":16}'
```
OMP's [model configuration](https://github.com/open-horizon-labs/oh-omp) and the underlying [provider guide](https://github.com/can1357/oh-my-pi/blob/main/docs/providers.md) document the custom-provider fields and credential resolution.
### Connect subscription access, then discover model IDs
Inside OMP, authenticate the first-party subscription paths you use:
```text
/login anthropic
/login openai-codex
/model
```
The first two commands connect Anthropic and ChatGPT/Codex access. `/model` then shows the exact models currently available to those credentials. Discover the IDs there instead of copying a model name from an old article; entitlements and model catalogues change.
### Route roles, then add operational fallback
Put explicit role policy in `~/.oh-omp/agent/config.yml`, replacing the placeholders with IDs shown by `/model`:
```yaml
modelRoles:
default: writer/palmyra-x6
plan: "anthropic/:high"
slow: "openai-codex/:high"
retry:
enabled: true
maxRetries: 3
fallbackChains:
default:
- "anthropic/"
- "openai-codex/"
plan:
- "openai-codex/:high"
fallbackRevertPolicy: cooldown-expiry
```
`modelRoles` is the deliberate task policy: normal implementation, planning, or slower deep work. `retry.fallbackChains` is operational recovery after retryable errors such as rate limits or provider failures; `fallbackRevertPolicy` returns a recovered role to its primary after the cooldown expires. A fallback chain is not a quality judge. Route work by role only after evaluating accepted outcomes for that role.
Native Claude Code and Codex remain separate tools even when OMP can authenticate to their model providers. Cross-provider routing belongs in a shared harness such as OMP, or in a gateway. Anthropic's [gateway documentation](https://docs.anthropic.com/en/docs/claude-code/llm-gateway) makes the boundary explicit: a gateway centralizes credentials, usage, and provider switching, but remains infrastructure that must preserve the client's protocol and features.
### Verify the whole path
Use observable evidence in this order:
```bash
omp --list-models writer
omp --list-models
omp
```
1. Confirm `writer/palmyra-x6` appears in the first listing and the intended Claude and Codex IDs appear in the second.
2. Start OMP, open `/model`, and verify the role assignments resolve to selectable models.
3. Send one bounded prompt to each configured role, then inspect `/usage` for the provider and limits actually used.
4. Test fallback only in a disposable sandbox with no sensitive repository or credentials beyond the test provider. Point the primary at an intentionally invalid or tightly limited test provider, observe the retryable failure and fallback, then remove the sandbox configuration. Never induce failure against production access.
Then set the models OMP should use for different kinds of work:
```text
/model
```
A sensible first configuration gives OMP a normal model for implementation, a cheaper model for quick exploration, and a stronger model for difficult debugging or planning. OMP calls these roles `default`, `smol`, `slow`, and `plan`.
For a new project, create or improve its instructions before asking an agent to edit anything. OMP can discover `AGENTS.md` and `CLAUDE.md`, along with configuration from compatible agent directories. This is especially useful if a repository already supports Codex or Claude Code: keep the shared engineering rules in one clear project instruction file instead of maintaining three subtly different copies.
## 2. Use OMP without Claude Code or Codex
This is the smallest useful setup. OMP is both the agent and the interface:
```bash
cd path/to/your-project
omp
```
Start with a bounded request: “Explain the architecture and identify the safest file for this change. Do not edit yet.” Once its explanation is grounded, ask it to make the change, run the smallest relevant test, and show the diff.
Useful OMP session commands are:
```text
/plan # discuss the approach before edits
/session # inspect the current session
/compact # compress an overlong conversation
/resume # select a prior session
/extensions # see discovered rules, skills, hooks, and commands
```
You can resume the most recent saved OMP session from the shell with `omp --continue`, or choose a session with `omp --resume`. OMP also has a non-interactive mode for contained, scriptable jobs, but do not use it as a substitute for review: check the diff and run project verification before merging work.
## 3. Add Herdr for durable, parallel terminals
Install Herdr using its installer, then launch it from a workspace where you want to manage agent panes:
```bash
curl -fsSL https://herdr.dev/install.sh | sh
cd path/to/your-project
herdr
```
Herdr's core value is operational rather than intellectual: it holds terminals open and lets you reconnect from another machine. Create one pane per task, not one pane per vague idea. For example:
```text
Pane 1 OMP Explore the codebase and write a plan
Pane 2 OMP Implement an approved, independent change
Pane 3 OMP Run tests and review the resulting diff
```
Avoid having two agents edit the same files simultaneously. Parallelism helps when tasks are independent—documentation versus a test audit, for example. It harms when every agent is changing the same central component.
Install the OMP integration after OMP has created its agent directory:
```bash
herdr integration install omp
herdr integration status
```
For OMP, the integration reports lifecycle state directly: Herdr can distinguish working, idle, and blocked rather than only inferring the state from terminal text. The integration is written into OMP's extension directory. If you also run Pi, ensure Pi and OMP have separate agent directories before installing both integrations; Herdr deliberately refuses an OMP installation that would share Pi's extension directory.
## 4. Run Claude Code in Herdr
Install Claude Code if it is not already available:
```bash
npm install -g @anthropic-ai/claude-code
cd path/to/your-project
claude
```
Sign in using the path appropriate for your account, then install Herdr's Claude integration:
```bash
herdr integration install claude
herdr integration status
```
The Claude integration writes a hook into Claude Code's configuration and reports Claude's native session identity to Herdr. That enables session restoration after a Herdr restart. Herdr still derives Claude Code's visible working/idle state from the terminal screen, so it is not the same lifecycle-authority model used by OMP.
Use Claude Code normally inside its own Herdr pane. Its CLI supports `claude -c` to continue the latest conversation and `claude -r ` to resume a specific one. Keep its task distinct from OMP's task unless you deliberately want a second opinion. A useful pattern is OMP for exploration and Claude Code for an independent implementation review, or the other way around.
## 5. Run Codex in Herdr
Install and sign in to Codex using OpenAI's [current CLI quickstart](https://learn.chatgpt.com/docs/codex/cli), then start it in the project:
```bash
cd path/to/your-project
codex
```
Now add Herdr's Codex integration:
```bash
herdr integration install codex
herdr integration status
```
The Codex hook reports native session identity so Herdr can restore the session. As with Claude Code, Herdr determines Codex's visible state from the terminal rather than receiving OMP-style lifecycle events. The installer updates Codex's hook configuration and enables the hooks feature; do not manually replace your existing configuration just to add this integration.
Codex can resume saved chats with `codex resume`. Inside an interactive session, `/permissions` lets you choose how much freedom the agent has to edit and execute commands. Start conservatively in unfamiliar repositories, then widen permissions only when you understand the project and the task.
## A practical three-agent workflow
Here is a compact workflow that does not require three subscriptions, but works well when you have access to all the tools:
1. In an OMP pane, ask for codebase mapping and a written plan. Do not allow edits yet.
2. In a Codex or Claude Code pane, ask for an independent review of the plan and edge cases.
3. Choose one agent to implement in a clean working tree or branch.
4. Use a second agent only to review the diff and run targeted checks—not to “also implement” the same task.
5. Keep Herdr running when you step away. When you return, inspect status, restore the session if needed, and review the actual repository state before continuing.
The same workflow works with just OMP and Herdr. The point is not to maximize the number of agents; it is to preserve clarity about ownership, verification, and the next decision.
## Troubleshooting checklist
If an integration does not appear to work, check these in order:
1. Run `herdr integration status` and confirm the agent's integration is installed and current.
2. Confirm the agent executable is available on your `PATH` in the same shell that starts Herdr.
3. Start the agent once outside Herdr if its configuration directory does not yet exist.
4. For OMP, check that its agent directory is not being shared with Pi.
5. Do not hand-edit hooks unless you understand the agent's configuration format; rerun the matching `herdr integration install ...` command instead.
To remove an integration cleanly, use the matching command, such as `herdr integration uninstall omp`, `herdr integration uninstall claude`, or `herdr integration uninstall codex`.
## The setup worth keeping
If you only want one recommendation: install OMP, run it in a disciplined repository workflow, and add Herdr when you need sessions to outlive a laptop lid or you need more than one focused pane. Add Claude Code or Codex when their native experience or account access is valuable to you—not because more agents automatically make a task better.
That is the durable arrangement: one workspace for persistence, one clear owner for each change, and a review step that no agent gets to skip.
## Sources
- [OMP documentation and installation](https://github.com/open-horizon-labs/oh-omp)
- [Oh My Pi provider configuration](https://github.com/can1357/oh-my-pi/blob/main/docs/providers.md)
- [Writer API quickstart](https://dev.writer.com/home/quickstart)
- [Writer model guide](https://dev.writer.com/home/models)
- [Claude Code LLM gateway](https://docs.anthropic.com/en/docs/claude-code/llm-gateway)
- [Herdr installation](https://herdr.dev/)
- [Herdr agent integrations](https://herdr.dev/docs/integrations/)
- [Claude Code setup](https://docs.anthropic.com/en/docs/claude-code/getting-started)
- [Claude Code CLI reference](https://docs.anthropic.com/en/docs/claude-code/cli-usage)
- [Codex CLI documentation](https://learn.chatgpt.com/docs/codex/cli)
### The Cheapest Coding Model Is the One That Gets the Change Accepted
- URL: https://mohammadshaker.com/en/blog/coding-model-pricing-kimi-k3-claude-5-openai
- Date: 2026-08-07T00:00:00.000Z
- Tags: Engineering, AI, Coding Agents, Authored with an LLM
A current comparison of Kimi K3, Claude Opus 5, Claude Sonnet 5 and OpenAI coding models by API price, provider and software-engineering benchmark score.
#### Content
The cheapest coding model is not the one with the cheapest tokens. It is the one that produces an accepted change with the least total waste.
That distinction matters. A model can be cheap per million tokens and expensive per finished task. It can loop, miss the root cause, break tests, and hand the cleanup back to me. Another model can cost five times more per token and still be cheaper because it gets the change right once.
I compare models through three numbers:
1. **Inference cost:** what the provider charges for input, cached input, and output.
2. **Task success:** what the model-plus-agent system resolves on relevant software benchmarks.
3. **Acceptance cost:** model spend plus review time, repair time, regressions, and retries.
The first two are public. The third is the decision.
Prices and provider offers below were checked on **7 August 2026**. They will move. The decision method should not.
## The answer first
| Need | My default | Why |
| --- | --- | --- |
| Best temporary price-performance | **GPT-5.6 Terra through OpenRouter** | The current promotion is below first-party list price, while reported SWE-bench Pro and terminal results remain close to the frontier. I treat the offer as temporary. |
| Best cheap first-party default | **Claude Sonnet 5 direct from Anthropic** | Its introductory price runs through 31 August 2026. It combines strong software results with a lower current task cost than Kimi K3. |
| Best hard-task model in this set | **Claude Opus 5** | It reports the highest SWE-bench Verified, Pro, Multilingual, and Multimodal results in this comparison. |
| Best open-weight frontier option | **Kimi K3 direct from Moonshot** | It combines open weights, a one-million-token context, provider choice, and strong long-horizon coding scores. |
| Best current OpenAI coding capability | **GPT-5.6 Sol direct from OpenAI** | It leads the reported DeepSWE and Terminal-Bench results in this set, but it carries a frontier price. |
| Cheapest narrow OpenAI helper | **GPT-5.6 Luna direct from OpenAI** | At $0.20 input and $1.20 output, it is now cheaper than GPT-5.4 nano. I use it for mechanical edits, extraction, and triage, not as the owner of ambiguous work. |
My routing rule is simple: use the cheapest model that consistently clears the quality floor. Escalate when ambiguity and failure cost rise.
## Current API price
All prices are USD per one million tokens. The reference workload is **5 million uncached input tokens plus 1 million output tokens**. That is an intentionally input-heavy agent workload, not a universal coding task.
| Model | Input / cached / output | Context | Reference workload | Cheapest practical provider checked |
| --- | ---: | ---: | ---: | --- |
| **Kimi K3** | Moonshot: $3.00 / $0.30 / $15.00; OpenRouter: $2.50 / — / $14.00 | 1.048M | Moonshot: $30.00; **OpenRouter: $27.96 including its 5.5% funded-account fee** | OpenRouter is cheapest; Moonshot is canonical |
| **Claude Opus 5** | $5.00 / $0.50 / $25.00 | 1M | $50.00 | Anthropic direct, AWS, and Bedrock tie at list price |
| **Claude Sonnet 5** | **$2.00 / $0.20 / $10.00 through 31 Aug**; then $3.00 / $0.30 / $15.00 | 1M | **$20.00 now**; $30.00 standard | Anthropic direct, AWS, and Bedrock tie at list price |
| **GPT-5.6 Sol** | $5.00 / $0.50 / $30.00 | 1.05M | $55.00 | OpenAI direct; Azure ties at list price |
| **GPT-5.6 Terra** | OpenAI: $2.00 / $0.20 / $12.00; OpenRouter promotion: $1.00 / — / $6.00 | 1.05M | OpenAI: $22.00; **OpenRouter: $11.61 including its 5.5% funded-account fee** | OpenRouter while the displayed promotion lasts |
| **GPT-5.6 Luna** | $0.20 / $0.02 / $1.20 | 1.05M | **$2.20** | OpenAI direct; its reduced first-party rate is below OpenRouter's displayed promotion |
| GPT-5.4 mini | $0.75 / $0.075 / $4.50 | 400K | $8.25 | OpenAI direct; Azure ties at list price |
| GPT-5.4 nano | $0.20 / $0.02 / $1.25 | 400K | $2.25 | OpenAI direct; Azure ties at list price |
The primary price references are [Moonshot's Kimi K3 pricing](https://www.kimi.com/resources/kimi-k3-pricing), [Anthropic's API pricing](https://platform.claude.com/docs/en/about-claude/pricing), and OpenAI's current model pages for [Sol](https://developers.openai.com/api/docs/models/gpt-5.6-sol), [Terra](https://developers.openai.com/api/docs/models/gpt-5.6-terra), and [Luna](https://developers.openai.com/api/docs/models/gpt-5.6-luna). Provider comparisons use [Fireworks](https://fireworks.ai/models/fireworks/kimi-k3), [Together](https://www.together.ai/models/kimi-k3), [Baseten](https://www.baseten.co/pricing/), [OpenRouter's Kimi K3 page](https://openrouter.ai/moonshotai/kimi-k3), and [OpenRouter's fee policy](https://openrouter.ai/pricing).
The table does not include taxes, subscriptions, enterprise discounts, batch discounts, or premium-speed tiers. Those are different purchasing decisions.
## Price beside SWE-bench score
The next table puts the same reference workload beside four SWE-bench variants. Higher scores are better. `NR` means the cited release did not report the result. It does not mean zero.
| Model | Reference workload | SWE-bench Verified | SWE-bench Pro | Multilingual | Multimodal |
| --- | ---: | ---: | ---: | ---: | ---: |
| **Kimi K3** | **$27.96 through OpenRouter** | NR | NR | NR | NR |
| **Claude Opus 5** | $50.00 | **96.0** | **79.2** | **89.5** | **59.4** |
| **Claude Sonnet 5** | **$20.00 current** | **85.2** | **63.2** | **78.3** | **28.1** |
| **GPT-5.6 Sol** | $55.00 | NR | **64.6** | NR | NR |
| **GPT-5.6 Terra** | **$11.61 promotional** | NR | **63.4** | NR | NR |
| **GPT-5.6 Luna** | **$2.20 direct** | NR | **62.7** | NR | NR |
| GPT-5.4 mini | $8.25 | NR | 54.4 | NR | NR |
| GPT-5.4 nano | $2.25 | NR | 52.4 | NR | NR |
Sources: [Moonshot's Kimi K3 report](https://github.com/MoonshotAI/Kimi-K3), [Claude Opus 5 system card](https://www.anthropic.com/claude-opus-5-system-card), [Claude Sonnet 5 system card](https://www-cdn.anthropic.com/73ad94ca3c0502e75e46637cc62c8bd9532a7f2c/Claude%20Sonnet%205%20System%20Card.pdf), [OpenAI GPT-5.6 evaluations](https://openai.com/index/gpt-5-6/), and [OpenAI GPT-5.4 mini and nano evaluations](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/).
The benchmark names look similar. The work is not identical:
| Benchmark | What it tests | How I use it |
| --- | --- | --- |
| SWE-bench Verified | 500 human-verified GitHub issues | A familiar baseline, now close to saturation for frontier systems |
| SWE-bench Pro | Harder tasks, larger multi-file changes, and less public answer leakage | My primary current SWE-bench signal |
| SWE-bench Multilingual | 300 tasks across nine programming languages | A better signal for estates that are not mostly Python |
| SWE-bench Multimodal | Issues that include screenshots and design mock-ups | Relevant to frontend and visually specified work |
A one-point difference across vendors is not a one-point product advantage. Anthropic and OpenAI use different agent harnesses, tool settings, reasoning budgets, and trial aggregation. OpenAI now says Verified is contaminated and no longer meaningful, and its audit found about 30% of SWE-bench Pro tasks broken. The measured unit is the **model plus harness**, not the weights alone. I use the scores as directional evidence, not clean model rankings. See OpenAI's [Verified warning](https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/) and [SWE-bench Pro audit](https://openai.com/index/separating-signal-from-noise-coding-evaluations/).
## Long-horizon coding changes the ranking
SWE-bench usually ends with one repository patch. Real agents also have to navigate terminals, recover from failed commands, maintain state, and work for longer. That is why I also look at newer agent evaluations.
| Model | DeepSWE v1.1 | FrontierSWE | SWE-Marathon | FrontierCode | Terminal-Bench 2.1 |
| --- | ---: | ---: | ---: | ---: | ---: |
| **Kimi K3** | **67.5** with Kimi Code | **81.2** | **42.0** | NR | **88.3** |
| **Claude Opus 5** | **68.8** | NR | NR | **53.4** | NR |
| **Claude Sonnet 5** | NR | NR | NR | **38.8** | **80.4** |
| **GPT-5.6 Sol** | **72.7** | 71.3 | 39.0 | 47.5 | **88.8** |
| **GPT-5.6 Terra** | **69.6** | NR | NR | NR | **87.4** |
| **GPT-5.6 Luna** | **67.2** | NR | NR | NR | **84.7** |
Sources: [Moonshot's Kimi K3 model report](https://github.com/MoonshotAI/Kimi-K3), the two Anthropic system cards above, and [OpenAI's GPT-5.6 evaluation report](https://openai.com/index/gpt-5-6/).
I do not average these columns. FrontierSWE, SWE-Marathon, FrontierCode, DeepSWE, and Terminal-Bench measure different task distributions. An average would look precise and mean very little.
## The cheapest provider for each model
| Model | Cheapest conclusion | My route |
| --- | --- | --- |
| **Kimi K3** | OpenRouter lists $2.50 / $14 before its funded-account fee, below Moonshot's $3 / $15. | **OpenRouter for minimum price; Moonshot direct** when canonical behaviour is worth the small premium. |
| **Claude Opus 5** | Anthropic, AWS, and Bedrock tie at $5 / $25. Google's multi-region endpoint is higher. | **Anthropic direct**, unless cloud controls, credits, or residency justify another route. |
| **Claude Sonnet 5** | Anthropic, AWS, and Bedrock tie at the current promotional rate. | **Anthropic direct** through 31 August, then re-evaluate. |
| **GPT-5.6 Sol** | OpenAI and Azure tie at list price. | **OpenAI direct**, unless Azure controls are the requirement. |
| **GPT-5.6 Terra** | OpenRouter currently undercuts OpenAI even after its funded-account fee. | **OpenRouter promotion**, with an automatic price check and first-party fallback. |
| **GPT-5.6 Luna** | OpenAI's current $0.20 / $1.20 first-party rate undercuts OpenRouter's displayed $0.50 / $3 promotion. | **OpenAI direct**. |
| **GPT-5.4 mini / nano** | OpenAI and Azure tie at global list price. | **OpenAI direct**. |
Provider choice is not only billing. An open-weight model can change with quantization, inference engine, context implementation, and tool-call support. Moonshot's [vendor verifier](https://github.com/MoonshotAI/Kimi-Vendor-Verifier) reports 67.5 on DeepSWE for Moonshot and 66.4 for Fireworks in its current K3 table. The difference is small. When price is tied, small is enough.
## The cost traps
The headline rate hides the bill.
- **Reasoning tokens are output tokens.** High-effort runs can multiply the expensive side of the workload.
- **Cache price is not cache value.** It helps only when the agent preserves reusable prefixes and the provider records real hits.
- **Promotions are routing inputs, not architecture.** Sonnet 5's introductory rate has an end date. OpenRouter's Terra discount has no public expiry. I build a fallback before I build a dependency.
- **Tokenizers differ.** Equal text does not mean equal token count across model families.
- **Retries compound.** A model that fails after a long reasoning trace pays the output bill and still leaves the task unfinished.
- **The provider is part of the system.** Rate limits, cold starts, tool schemas, context truncation, and model substitutions affect delivered quality.
This is why price-per-token is the beginning of the calculation, not the end.
## My routing policy
| Task | Default | Escalation | Cheap supporting model |
| --- | --- | --- | --- |
| Small, explicit edit | GPT-5.6 Luna direct or GPT-5.4 mini | Sonnet 5 | GPT-5.4 nano |
| Normal feature or bug | GPT-5.6 Terra promotion or Sonnet 5 | Opus 5 or GPT-5.6 Sol | GPT-5.4 mini |
| Ambiguous root-cause work | Opus 5 | GPT-5.6 Sol as an independent attempt | Terra or Sonnet 5 for review |
| Very long context or open-weight requirement | Kimi K3 | Opus 5 | A smaller local or hosted model after evaluation |
| Parallel mechanical review | GPT-5.4 nano or mini | Terra or Sonnet 5 for judgement | — |
The policy has one failure mode: keeping a cheap model in a loop after it has demonstrated that it does not understand the task. Cheap retries feel prudent. They are usually denial.
I escalate when any of these appears:
- the task crosses several subsystems;
- the failure is nondeterministic;
- the repository has weak tests;
- the change is security-sensitive;
- the first attempt fixes a symptom instead of the cause;
- reviewer time costs more than the model difference.
## Measure accepted changes, not benchmark theatre
Public benchmarks narrow the search. A private evaluation makes the decision.
I would take 30–100 representative tasks and record:
- first-pass task success;
- human cleanup minutes;
- regressions introduced;
- tool-call and schema failures;
- wall-clock time;
- uncached input, cached input, reasoning, and final-output tokens;
- provider cost;
- whether the final change was accepted.
Then I would calculate:
```text
cost per accepted change
= model spend
+ reviewer time
+ repair time
+ regression cost
+ failed-run cost
```
That number can reverse the public ranking. It should. My repository, agent scaffold, tests, and definition of “done” are not a vendor benchmark.
The decision rule is direct: **route by the lowest cost per accepted change at the required risk level.** Cheapest tokens are not cheapest engineering.
### Local AI Is Not Free: Open-Weight Models, Hardware Cost and Accuracy
- URL: https://mohammadshaker.com/en/blog/local-open-weight-models-hardware-cost-accuracy
- Date: 2026-08-07T00:00:00.000Z
- Tags: Engineering, AI, Open-Weight Models, Local AI, Authored with an LLM
A practical guide to open-weight models that fit on local machines, including quantized memory, UK hardware purchase cost, annual running cost, content accuracy and software-engineering accuracy.
#### Content
Open weights do not make inference free. They move the bill.
The API bill becomes hardware, electricity, maintenance, engineering time, and the opportunity cost of running a weaker or slower model. Privacy and control can justify that bill. A fantasy about “free local AI” cannot.
My decision model has four parts:
```text
model quality -> memory fit -> useful speed -> total ownership cost
```
If one part fails, the deployment fails. A model that barely loads but runs too slowly is not a local solution. A fast model that cannot complete the work is not cheap. A private model that needs constant operational attention is not low maintenance.
This guide covers the **major current open-weight decision set for one desk-side machine** as of 7 August 2026. It is not every checkpoint on Hugging Face. That catalogue changes daily and would hide the decision. I include models that are current, materially different, supported by a primary model card, and plausible on 16GB to 128GB of local memory. I separately show frontier weights that are open but not honestly single-machine models.
## The answer first
| Situation | My choice | Decision rule |
| --- | --- | --- |
| I already own a 16GB machine | **Gemma 4 12B QAT** or **gpt-oss-20B** | Start with existing hardware. Do not buy a workstation before the workflow proves useful. |
| I want the best one-GPU generalist | **Qwen3.6-27B Q4** on 32GB | It has the strongest reported combined content and SWE scores in the practical one-GPU group. |
| I want the faster one-GPU generalist | **Qwen3.6-35B-A3B Q4** on 32GB | It gives up a few benchmark points but activates only 3B parameters per token. Total weights still determine memory. |
| I want a coding specialist | **Devstral Small 2** on 32GB | It is designed for software agents and reports 68.0% SWE-bench Verified. I choose it for its ecosystem, not because it beats Qwen3.6. |
| I want portable 64GB headroom | **Qwen3.6-27B on a 64GB MacBook Pro** | The extra memory buys context and room for tools, not a stronger checkpoint by itself. |
| I want Kimi K3 locally | **I do not buy a workstation for it** | K3 is open-weight, but its official deployment target is a 32-GPU server. I use the API or a proper cluster. |
The recommendation is not “buy the biggest box.” It is “buy only after a smaller local model fails a measured requirement.”
## Accuracy and memory fit
“Content accuracy” is not a settled benchmark. Good prose has voice, structure, factuality, instruction-following, and judgement. I use **MMLU-Pro or GPQA Diamond as a factual-reasoning proxy**, not as a writing-quality score. For software work, I prefer **SWE-bench Verified and Pro** because they require repository changes. LiveCodeBench is listed only when a model vendor does not report SWE-bench.
Memory estimates are planning ranges for the model weights plus basic runtime overhead. They assume a roughly four-bit format where available. Long contexts need additional KV-cache memory. MoE models store all experts even though only a few are active per token. Active parameters reduce compute; they do not make the other weights disappear.
| Model | Architecture | Official or canonical local footprint | Sensible machine memory | Content / reasoning proxy | SWE-task score | My use |
| --- | --- | ---: | ---: | ---: | ---: | --- |
| **Gemma 4 12B Unified QAT** | 11.95B dense, native multimodal | **6.98GB** Q4_0 GGUF plus 175MB projector | 16GB | MMLU-Pro **77.2**; GPQA **78.8** | SWE-bench NR; LiveCodeBench v6 **72.0** | Best accessible content and multimodal choice |
| **gpt-oss-20B** | 21B total / 3.6B active MoE, text-only | Native MXFP4 runs within **16GB** | 16–24GB | High-reasoning MMLU **85.3**; GPQA **71.5** without tools | SWE-bench Verified **60.7** at high reasoning | Best 16GB text-only reasoning baseline |
| **Qwen3.6-27B** | 27B dense, multimodal | Canonical Q4_K_M **19.1GB**; optional MTP 1.68GB and vision projector 629MB | 32GB | MMLU-Pro **86.2**; GPQA **87.8** | Verified **77.2**; Pro **53.5** | Best all-round local model in this set |
| **Qwen3.6-35B-A3B** | 35B total / 3B active MoE, multimodal | Canonical Q4_K_M **20.4GB**; optional MTP 1.06GB and vision projector 614MB | 32GB | MMLU-Pro **85.2**; GPQA **86.0** | Verified **73.4**; Pro **49.5** | Faster balanced alternative to the dense 27B |
| **Devstral Small 2** | 24B dense, vision, coding specialist | Official FP8 checkpoint files total about **25.7GB** | 32GB | Comparable general score NR | Verified **68.0**; Multilingual **55.7** | Mistral ecosystem and specialist coding workflows |
Primary model cards and artifacts: [Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B), [Qwen3.6-27B GGUF](https://huggingface.co/ggml-org/Qwen3.6-27B-GGUF), [Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B), [Qwen3.6-35B-A3B GGUF](https://huggingface.co/ggml-org/Qwen3.6-35B-A3B-GGUF), [Gemma 4 12B](https://huggingface.co/google/gemma-4-12B), [Gemma 4 12B QAT](https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf), [gpt-oss-20B](https://openai.com/index/introducing-gpt-oss/), and [Devstral Small 2](https://huggingface.co/mistralai/Devstral-Small-2-24B-Instruct-2512).
These scores are not a neutral tournament. Qwen's Qwen3.6 card warns that its software results use its own scaffold and a refined Pro set. GLM reports its own harness. OpenAI's high-reasoning numbers spend more test-time compute. Quantization, context size, prompt format, tool implementation, and agent scaffold all change the result.
The right interpretation is coarse:
- Qwen3.6-27B is the strongest balanced local candidate here.
- Qwen3.6-35B-A3B is the speed-oriented alternative; fewer active parameters do not reduce the stored weight set.
- Gemma 4 12B is the accessible content and multimodal candidate.
- gpt-oss-20B is attractive when tool use, permissive licensing, and the OpenAI model format matter.
- Devstral is an ecosystem-specific coding specialist, not the default score leader.
- A missing SWE score is missing evidence, not hidden strength.
## What the machine costs
I use current UK prices, then label any complete-system estimate as an estimate. NVIDIA lists the RTX 5060 Ti 16GB at **£349 MSRP** and the RTX 5090 32GB at **£1,799 MSRP**. Those are card prices, not computers. Apple lists a complete 64GB MacBook Pro at **£2,999**. Dell lists a complete 128GB GB10 system at **£6,125.27**. Comparing a bare GPU with a finished Mac would be dishonest.
| Local tier | Representative system | Models it fits comfortably | Current purchase cost | Annual electricity | Annual 5% reserve | Annual operating plan |
| --- | --- | --- | ---: | ---: | ---: | ---: |
| Existing 16GB computer | Current laptop or desktop | Gemma 4 12B QAT; gpt-oss-20B with controlled context | **£0 incremental** | £60–£140 | £0 | **£60–£140** |
| 16GB CUDA upgrade | RTX 5060 Ti 16GB in a compatible existing PC | Gemma 4 12B QAT; gpt-oss-20B | **£349 GPU MSRP** | up to £137 GPU-only | £17 | **up to £154**, excluding host PC |
| 32GB CUDA tower | RTX 5090 complete PC | Qwen3.6-27B, Qwen3.6-35B-A3B, Devstral Small 2 | **£3,000–£3,600 estimated**; GPU MSRP is £1,799 | up to £438 GPU-only | £150–£180 | **up to £588–£618**, plus non-GPU power |
| 64GB portable unified memory | 14-inch MacBook Pro M5 Pro, 64GB / 1TB | All shortlisted models, with more context headroom | **£2,999 complete** | £60–£140 estimated | £150 | **£210–£290** |
| 96GB unified-memory desktop | Mac Studio M3 Ultra base, 96GB / 1TB | Large experimental quants, though not necessarily at interactive speed | **from £4,199 complete** | £170–£260 estimated | £210 | **£380–£470** |
| 128GB CUDA-compatible desktop | Dell Pro Max with GB10, 128GB / 4TB | Large weights when capacity matters more than bandwidth | **£6,125.27 complete** | up to £183 at the 240W adapter ceiling | £306 | **up to £489** |
| 96GB professional CUDA GPU | RTX PRO 6000 Blackwell Max-Q | 80–120B-class weights in VRAM | **£17,354.58 card only** | up to £229 GPU-only | at least £868 | **at least £1,097**, excluding host PC |
The operating model uses [Ofgem's July–September 2026 electricity rate of 26.11p/kWh](https://www.ofgem.gov.uk/information-consumers/energy-advice-households/energy-price-cap-unit-rates-and-standing-charges), eight inference hours per day, and a five-per-cent annual hardware reserve. Standing charges are excluded because the household pays them without the machine. Electricity varies with utilisation, power limits, and whether the workload keeps the GPU saturated.
The [NVIDIA RTX 50-series pages](https://www.nvidia.com/en-gb/geforce/graphics-cards/50-series/) supply the official GPU MSRPs and memory. Apple's [64GB MacBook Pro listing](https://www.apple.com/uk/shop/buy-mac/macbook-pro/14-inch-silver-standard-display-apple-m5-pro-chip-18-core-cpu-20-core-gpu-64gb-memory-1tb-storage) and [Mac Studio store](https://www.apple.com/uk/shop/buy-mac/mac-studio) supply complete-system prices. Dell's [GB10 listing](https://www.dell.com/en-uk/shop/desktop-computers/sr/all-products/desktops/dell-pro-max-dell-precision?appliedRefinements=38600) and [RTX PRO 6000 listing](https://www.dell.com/en-uk/shop/nvidia-rtx-pro-6000-blackwell-max-q-workstation-edition-300w/apd/490-bldl/graphic-video-cards) anchor the specialist tiers.
There are three traps in this table.
First, memory capacity is not memory bandwidth. A model can fit on a 128GB machine and still generate too slowly for interactive coding. Second, unified memory is useful but shared: the operating system, context cache, and other applications consume the same pool. Third, a 24GB GPU can run a 20GB checkpoint and still fail at a long prompt because the KV cache has nowhere to go.
I leave at least 20% headroom. I leave more for long-context agents.
### A field note on shared memory
One practitioner ran a local model on an RTX 4060 with 64GB of system RAM for several months and usually saw roughly **8–14 tokens per second**. As agent history grew, the model competed with the browser and other desktop applications for memory, making the whole workflow less predictable. This is an anonymized anecdote, not a controlled benchmark: the model, quantization, context, runtime, and prompt mix were not held constant. Its useful lesson is narrower—measure the complete working session, not an isolated generation with every other application closed.
### The practical two-GPU ceiling at home
A dedicated inference box with two used 24GB cards in the RTX 3090 class can be a practical home ceiling. Forty-eight gigabytes of aggregate VRAM opens model and context combinations that cannot fit on one card, while keeping the machine within ordinary workstation territory. It also brings heat, power, noise, motherboard-lane, power-supply, and used-hardware risk. Prefer both cards in one host; distributing inference across networked desktop machines adds another slow and failure-prone interconnect.
Two GPUs combine **capacity**; they do not become one magically fast 48GB GPU. Every token still crosses the split, so PCIe or NVLink topology and the runtime's split mode matter. [llama.cpp's multi-GPU guide](https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md) uses layer split by default: each GPU owns a contiguous set of layers, which reduces transfers and is the compatible starting point. Its tensor mode splits work within layers and can improve token latency on a fast interconnect, but it is experimental and more sensitive to communication speed.
My rule is: buy the second card for a measured capacity requirement, then benchmark the exact model, context, batch size, and split mode. Do not justify it by doubling a single-GPU tokens-per-second figure.
## The open models I would not call local
Open-weight and local are not synonyms.
| Model | Why it looks attractive | Full-weight reality | My decision |
| --- | --- | --- | --- |
| **Kimi K3** | 2.8T total parameters, 104B active, strong long-horizon coding | Moonshot's minimum recipe starts at **32 H100-class GPUs** and recommends **64 or more accelerators** | Use Moonshot or another verified provider; buy a cluster only for a real utilisation case |
| **gpt-oss-120B** | Officially fits a single 80GB accelerator and reports 62.4% Verified at high reasoning | Technically local on 96–128GB hardware, but no longer the strongest reason to buy that tier | Rent or test on existing large-memory hardware before buying |
| **Llama 4 Scout** | Open multimodal MoE with a very long context | 109B total parameters; Meta's fit claim targets an 80GB H100 at INT4 | Hosted inference or an existing 80GB system |
| **DeepSeek V4 Flash** | 284B total / 13B active and strong hosted economics | Mixed FP4/FP8 weights remain too large for a normal 128GB workstation with safe headroom | Hosted inference |
| **Mistral Small 4** | 119B total / 6.5B active and agentic capability | Mistral's official minimum is 4× H100, 2× H200, or 1× DGX B200 | Hosted inference; unofficial CPU offload is an experiment, not a buying case |
Kimi K3 makes the boundary obvious. Four-bit weights alone would be measured in terabytes, not workstation gigabytes. Its 104B active parameters reduce per-token compute relative to a dense 2.8T model, but all experts still have to live somewhere. [Moonshot's official K3 repository](https://github.com/MoonshotAI/Kimi-K3) is clear about the 32-GPU deployment target.
The exclusion evidence comes from [OpenAI's gpt-oss guidance](https://openai.com/index/introducing-gpt-oss/), [Meta's Llama 4 release](https://ai.meta.com/blog/llama-4-multimodal-intelligence/), [DeepSeek V4 Flash](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash), and [Mistral Small 4](https://mistral.ai/news/mistral-small-4/). If the machine needs a rack, specialist cooling, and an operator, it is local infrastructure. It is not a local workstation.
## The ownership calculation
I compare local and hosted inference over the same period.
```text
monthly local cost
= purchase price / useful life in months
+ annual electricity / 12
+ annual maintenance reserve / 12
+ operator time
```
A £3,200 workstation over 36 months, with £600 a year in electricity and reserve, costs about **£139 per month before my time**. A £4,200 machine with £500 a year in operating cost is about **£158 per month**.
That creates a clean break-even test:
```text
break-even accepted tasks
= monthly local ownership cost
/ hosted cost per accepted task
```
If a hosted model costs £2 per accepted task, the £139 local system needs about 70 accepted tasks each month just to match the direct spend. If the local model creates more review or repair work, the break-even point moves away. If data cannot leave the machine, the privacy requirement may settle the decision before cost does.
The wrong comparison is local hardware versus API tokens. The correct comparison is **local accepted outcomes versus hosted accepted outcomes**.
## Hosted overflow without losing the cost model
I use hosted capacity as overflow, not as an ungoverned second architecture:
1. Consume existing Claude, ChatGPT, or other subscriptions first when their terms and tools fit the task.
2. Keep a pay-as-you-go provider for rate limits, unavailable models, or bursts.
3. Split planner and executor only after an evaluation shows that a strong planner plus a cheaper executor preserves accepted outcomes. A cheaper second call is not a saving if it creates more repairs.
Provider privacy needs explicit configuration. [DeepInfra states](https://docs.deepinfra.com/account/data-privacy) that standard inference inputs are not written to disk, outputs are deleted after return, and submitted data is not used for training; Google and Anthropic models follow exceptions described on that page, and bulk inference has different temporary-storage behavior. [OpenRouter's default routing](https://openrouter.ai/docs/guides/routing/provider-selection) may choose providers that store data because `data_collection` defaults to `allow`. Set `data_collection: "deny"` when storage is unacceptable, and use its [Zero Data Retention controls](https://openrouter.ai/docs/guides/features/zdr) when no provider retention is the requirement. Those controls govern inference routing; enabled tools can have separate policies.
Coding harnesses raise the stakes because prompts and tool results can contain source, logs, configuration, or credentials. Keep API keys in a secret manager or process environment, scope and rotate them, redact tool output, exclude secret files from agent access, and never paste a key into a prompt, checked-in config, or command that records it in shell history. A privacy promise cannot recover a key already exposed by the client.
The reusable buy-or-rent rule is:
```text
buy when measured local accepted outcomes meet the quality floor,
fit with operational headroom, and beat rented accepted-outcome cost
at realistic utilisation over the hardware life;
otherwise rent, with explicit privacy and spend controls
```
The quality gate stays the same on both paths: evaluate accepted content or accepted code changes, including reviewer time and regressions, then promote a model only when it meets that workload's floor.
## Content and SWE need different evaluations
I would not select one model for both workloads from one leaderboard.
For content, I would test:
- factual precision against a source packet;
- unsupported claims and invented citations;
- instruction adherence;
- structural coherence over 1,500–3,000 words;
- editing distance from my accepted version;
- tokens per accepted article.
For software engineering, I would test:
- first-pass task completion;
- tests passed without weakening them;
- regressions introduced;
- tool and schema failures;
- reviewer minutes;
- wall-clock time;
- tokens or joules per accepted change.
The same model can win one and lose the other. Gemma 4's multimodal and content strengths do not prove repository repair ability. Qwen3-Coder-Next's SWE result does not prove editorial voice. Specialisation is not a defect. Pretending it does not exist is.
## My buying sequence
1. **Run gpt-oss-20B or Gemma 4 12B QAT on hardware I already own.** Measure quality and speed for two weeks.
2. **Rent before buying.** Run Qwen3.6-27B, Qwen3.6-35B-A3B, and Devstral Small 2 on rented instances with the exact local serving stack.
3. **Define the quality floor.** Use accepted content and accepted code changes, not vibes.
4. **Choose the smallest memory tier that keeps 20% headroom.** Include the real context window and concurrency.
5. **Calculate 36-month ownership cost.** Add electricity, maintenance reserve, and my operating time.
6. **Buy only when utilisation or privacy makes the result obvious.** If the spreadsheet needs heroic assumptions, the API is still cheaper.
My default is a 32GB system for one-person experimentation and Qwen3.6-27B as the balanced starting model. I choose Qwen3.6-35B-A3B when generation speed matters more than the last few benchmark points. I move to 64GB only when measured context, concurrency, or workspace requirements justify the extra memory. I treat 96–128GB as a niche capacity tier, not the automatic next step.
I do not buy local hardware to avoid a token bill. I buy it to gain control over a workload I have already measured.
### Unstuck Stuck
- URL: https://mohammadshaker.com/en/blog/unstuck-stuck
- Date: 2026-08-04T06:30:00.000Z
- Tags: Brain dump, Human-Written
When I feel stuck, the problem is most probably framing. A couple of questions I find useful are:
#### Content
When I feel stuck, the problem is most probably framing. A couple of questions I find useful are:
**I should define the real constraint (Ala Hume, Popper, Musk, Andy Grove)**
No generics: “I need more features.”
More specifics: “users can’t get the value within 2 minutes.”
Fix that.
**I should kill sacred assumptions (Ala Hume, Popper, Musk, Christensen)**
“We must do it this way.”
“Enterprise buyers expect this.”
“The UI has to look like similar to X because that’s what people know.”
Most of these are inherited, not proven.
**I should rebuild a clean version (Ala Hume, Popper, Christensen)**
“If we started today, with today’s tools, what would we do?”
As a concrete alternative I can sketch in an hour.
**I should keep risk small (Ala Nassim Taleb)**
Can I run a small test within 2h? or 2 days?
I most probably can.
A thin slice e2e.
A measurable experiment I can act on, NOW.
A tangible thing I can learn from, NOW.
**I should tinker (Ala Nassim Taleb)**
Stay aware of new tools and play with them. I can’t know which one will work. And that’s OK.
That’s it.
**Books around this I actually like**
- Elon Musk, Ashlee Vance
- The Innovator’s Dilemma, Clayton Christensen
- “Back to The Drawing Board”, Seth Godin
- The Logic of Scientific Discovery, Karl Popper
- Fooled by Randomness, The Black Swan, Antifragile. All by Nassim Taleb
### On Simplicity. What I Learned from Dieter Rams
- URL: https://mohammadshaker.com/en/blog/on-simplicity-what-i-learned-from-dieter-rams
- Date: 2026-07-28T00:00:00.000Z
- Tags: Brain dump, Human-Written
Simple is hard. Simplicity is work.
#### Content
Dieter Rams is still the clearest example I know on this.
Simple is hard. It is extremely hard.
That is one of those sentences I used to understand too quickly. I now think it takes years to understand properly. In design and in engineering, simplicity is often the last thing I arrive at. It comes after false starts, ornament, cleverness, features that seemed necessary until they were not, and architectures that looked impressive until I had to live inside them.
Early in my career, I made the usual mistake. I thought simple meant basic. I thought minimal meant naive. I thought a thing with fewer visible parts must have required less thought.
I learned the hard way it's the completely other way around.
At Braun, Rams designed objects that did not try to announce themselves. Radios, calculators, record players, shelving systems. They sat there with a kind of moral calm. Buttons were placed where the hand expected them. Surfaces were left alone. Labels did not shout. Nothing appeared to be performing design. Yet everything had been designed.
See [Braun's "World Firsts"](https://www.braun-audio.com/en-GB/worldfirsts) for the full lineage.
That is the trick I keep returning to.
The stackable systems are pure genius. On the surface: clean, neat, almost obvious. Underneath: constraint, modularity, extension, future use. As an engineer, I recognize the move immediately. Components. Interfaces. Customization. Scalability without visible disorder. A system that can grow without becoming a pile.
I see the same thing in older Apple products. The iPod click wheel, for example, was not simple because it lacked thought. It was simple because so much thought had been compressed into one gesture. Scroll, select, move, play. A small circle that absorbed a large amount of complexity.
I keep coming back to Rams for this reason. **The art of his is what he exposed and what he hid**. I look at the affordances of the buttons. I look at where the object gives instruction without explaining itself. I look at how constraint becomes ease.
This matters to me in engineering too.
A good API can feel like that. A backend can feel like that. A frontend architecture can feel like that. The best systems do not make their internal difficulty someone else’s problem. They offer a small number of good moves and make the wrong moves hard. This is sth I learned from my previous boss [Jagadeesh Annamalai. And it's extremely hard to get right.](https://www.linkedin.com/in/jagadeesh-annamalai/)
That is why simplicity takes time.
I add. I test. I remove. I rename. I split. I merge. I look again. I realize that what seemed essential was often only defensive. I delete it. Then I delete again. This is exactly the same as via-negativa of Nassim Taleb. Remove, not add, to improve a system.
The finished thing, if it is any good, will not show this labor. It will look as if it had to be that way from the beginning. That is the strange injustice of good work. The clearer it becomes, the easier it looks.
But simple is not easy.
Simple is earned.
Salam.
### Illuminations: Walter Benjamin Through Ways of Seeing
- URL: https://mohammadshaker.com/en/blog/illuminations-walter-benjamin-ways-of-seeing
- Date: 2026-07-24T00:00:00.000Z
- Tags: Walter Benjamin, John Berger, Art Criticism, Visual Culture, Philosophy, Authored with an LLM
Reading Walter Benjamin's Illuminations beside John Berger's Ways of Seeing shows how reproduction reshapes art's authority, context, and politics.
#### Content
I first read John Berger's *Ways of Seeing* as a book about looking at paintings. I came back to it later and found something larger: a book about authority. Who gets to decide what an image means? Who owns the original? Who controls the copy? What happens when a painting leaves a church or palace and arrives on a television screen, a book page, an advertisement, or a phone?
Reading Walter Benjamin's *Illuminations* alongside Berger made these questions sharper. Berger gives them an extraordinary directness, but Benjamin gives them depth, tension, and history. The relationship between the two is not simply influence. It is closer to a change of medium. Benjamin builds a difficult philosophical machine; Berger switches it on in front of a camera and shows us what it does.
## A collection made of fragments
*Illuminations* is not a single argument. The English-language collection, edited and introduced by Hannah Arendt and translated by Harry Zohn, first appeared in 1968. It brings together essays written in very different circumstances: on translation, literature, storytelling, art, technology, and history. Benjamin did not write a tidy system. He thought through constellations—placing an artwork, a machine, a childhood memory, or a historical image beside another until each altered the light falling on the other.
This matters when reading the most famous essay in the collection, commonly translated there as "The Work of Art in the Age of Mechanical Reproduction." It is tempting to extract one idea—*aura*—and turn Benjamin into a philosopher of nostalgia. The original artwork is authentic; the copy is thin; modern life destroys what was sacred. But this is too simple.
Benjamin is not merely mourning the old work of art. He is examining a historical transformation. A unique object has a particular presence: it has been somewhere, has aged, has passed through hands, and belongs to a tradition. A reproduction loosens the image from that place and history. It can travel to viewers whom the original could never reach. This weakens one kind of authority, but it also creates new powers and new forms of collective experience.
The loss is real. So is the opening.
## What reproduction actually changes
A reproduction does more than make a second version of an image. It changes the conditions under which the image is encountered.
I can see a fresco without entering its building. I can enlarge a face that was a minor detail in a large painting. I can place a Renaissance portrait beside a political poster or a luxury advertisement. I can add music, crop the frame, write a caption, and circulate the result to millions of people. The material image may be copied faithfully while its social meaning changes completely.
Benjamin saw photography and film as decisive because they do not simply reproduce older art. They reorganize perception. The camera reveals details, angles, speeds, and sequences unavailable to unaided vision. Film also replaces sustained contemplation with interruption, cutting, repetition, and collective reception. Technology therefore acts on both sides of the encounter: it transforms the artwork and trains the viewer.
This is the foundation Berger takes into *Ways of Seeing*. The 1972 book grew from the BBC television series and was made collaboratively with Sven Blomberg, Chris Fox, Michael Dibb, and Richard Hollis. Its form is part of its argument. Seven essays alternate between words-and-images and image-only sequences. Paintings are cropped, juxtaposed, and printed without the protective distance of a museum wall. The book does not merely tell us that context changes meaning. It keeps changing the context in front of us.
Berger's great practical insight is that reproduction makes an image mobile, and a mobile image becomes available for many uses. Its meaning is no longer secured by its original location. Yet mobility does not free the image from power. A copy can challenge authority, but it can also manufacture a new authority through commentary, editing, prestige, and price.
## From aura to mystification
Benjamin's aura is difficult because it names several related things at once: uniqueness, distance, authenticity, tradition, and a mode of attention. Berger pairs Benjamin's account of aura with his own, more accessible critique of *mystification*.
For Berger, the institutional presentation of Old Master paintings often makes the past appear timeless and spiritually elevated. The museum, the catalogue, the expert voice, and the market can conceal the material relations inside the picture: land, property, class, patronage, and possession. A painting becomes evidence of eternal beauty rather than a human object made within a specific arrangement of power.
This is where Berger is at his strongest. He asks us to look at what official interpretation encourages us not to see. In European oil painting, objects often appear as things that can be owned: fabric, silver, animals, land, even the painted world itself. The technique's ability to render surfaces becomes connected, in Berger's argument, to the pleasures and claims of property.
Benjamin helps me understand why this authority can survive reproduction. The original's aura may decline, but the institutions around art can simulate distance in new forms. The image in a glossy catalogue, auction listing, or branded museum campaign remains reproducible while acquiring another kind of prestige. Mechanical availability and social exclusivity can coexist.
Berger therefore extends Benjamin by asking not only what happens to the artwork when it is copied, but who benefits from each new context.
## The spectator is never neutral
Berger also goes somewhere that Benjamin's reproduction essay does not go far enough: gender.
In one of the book's best-known arguments, Berger distinguishes between the social permission traditionally given to men to act and the pressure placed on women to watch themselves being seen. He traces this split through the European nude, where the depicted woman is frequently arranged for an assumed male owner-spectator.
The important point is not that every nude has one meaning. It is that seeing has a social structure. The viewer imagined by an image is not an abstract pair of eyes. The viewer may be gendered, classed, invited to possess, encouraged to judge, or made to feel inadequate.
Berger then connects the inherited visual language of oil painting to modern publicity. The continuity is not exact, but it is revealing. Advertising borrows poses, gestures, settings, and symbols of abundance from older art. The difference is in the promise. The commissioned portrait confirms what its owner already has; the advertisement addresses a spectator who is meant to feel what they lack. Desire is organized through an imagined future self.
Benjamin gives Berger the historical break created by reproduction. Berger follows the reproduced image into everyday capitalist life and studies how it addresses us there. This is a genuine extension: from the changing status of art to the production of spectators.
## Where Berger simplifies Benjamin
The cost of Berger's clarity is that Benjamin's ambivalence sometimes becomes a cleaner political verdict.
Berger often writes as if exposing the hidden relation—property, patriarchy, commercial interest—can release the viewer from mystification. Benjamin is less confident that reproduction has a politically guaranteed outcome. Reproduction is neither simply liberating nor simply corrupting. It opens art to mass participation and new political forms, but the same techniques can turn politics into spectacle. The apparatus does not guarantee its use.
Benjamin is also more interested in the new capacities of perception themselves. Film is not important only because it distributes images or strips authority from originals. Its cuts, close-ups, repetitions, and shocks alter how a modern collective senses reality. Berger uses these techniques brilliantly, especially on television, but his written argument more often emphasizes what images conceal than what new media make perceptually possible.
There is another simplification in the idea of a lost original context. Berger is right that a crop, caption, or soundtrack changes a painting's meaning. But an original meaning was never perfectly sealed. Works moved between owners, buildings, rituals, and publics before photography. Benjamin's broader writing is useful here because he treats reception and historical survival as active parts of a work's life. Reproduction accelerates and radicalizes movement; it does not invent historical change.
Finally, Berger's polemical categories can flatten differences among paintings. Reading an entire tradition through property and the male spectator reveals a great deal, but no single key opens every picture. Berger is most valuable to me as a method of suspicion, not a final dictionary of meanings.
## Storytelling and the poverty of information
"The Storyteller," another essay in *Illuminations*, adds something unexpected to *Ways of Seeing*. Benjamin distinguishes lived, shareable experience from information that arrives already explained and quickly expires. A story carries traces of the storyteller and leaves room for the listener. Information demands immediate plausibility; a story can continue unfolding long after it is heard.
This distinction changes how I think about images online. We have more visual information than either Benjamin or Berger could have imagined, but abundance is not the same as experience. A feed tells me what an image is, why it matters, and how to react before I have had time to look. The image is endlessly reproducible and instantly exhausted.
Berger's slow juxtapositions offer a different practice. They do not restore the original aura, nor do they pretend that mediation can be escaped. They make the mediation visible. They return duration and comparison to looking. In that sense, *Ways of Seeing* behaves less like a textbook and more like Benjamin's storyteller: it gives the reader material for judgment without closing the experience completely.
## Looking backward against the victors
Benjamin's theses on history deepen the political stakes again. He rejects history as a smooth march of progress and asks us to attend to what triumphal narratives leave behind. Cultural treasures cannot be separated from the violence, labour, and domination within the histories that preserved them.
This idea sits quietly beneath Berger's treatment of the European canon. The question is not whether the paintings are beautiful. It is what kind of history their beauty is made to tell. When art history presents a sequence of masterpieces as civilization naturally accumulating excellence, the frame excludes the defeated, the dispossessed, the workers, and often the women who appear as objects rather than agents.
To look critically is not to cancel beauty. It is to refuse beauty as an alibi.
Here Berger's directness is an advantage. Benjamin's philosophy of history can feel dense and messianic. Berger turns a similar suspicion toward a room, a painting, and then an advertisement. He makes the historical question usable: what had to disappear from view for this image to appear self-evident?
## A method for seeing now
I take four habits from reading these books together.
First, ask where an image is encountered, not only what it depicts. A church, museum, book, television programme, auction page, and social feed produce different relations to the same picture.
Second, inspect the frame around the frame: caption, crop, sequence, soundtrack, interface, price, and institution. Context is not decoration added to meaning. It helps make meaning.
Third, ask what kind of spectator the image wants. Does it invite attention, ownership, envy, desire, obedience, judgment, or participation?
Fourth, preserve the tension. Reproduction can break monopolies of access while creating larger systems of manipulation. An original can carry historical presence while being used to sanctify wealth. Critical looking begins when I stop demanding that either technology or tradition be innocent.
Benjamin explains why the image could leave its place. Berger shows what happened when it entered the living room. We now meet it in the feed, where every image can be copied, edited, ranked, captioned, and sold within seconds. The machinery has changed, but the question remains: not only what am I seeing, but who has arranged this way of seeing for me?
That question does not make looking colder. It makes looking responsible.
## References and further reading
- Walter Benjamin, [*Illuminations*](https://www.penguin.co.uk/books/367535/illuminations-by-walter-benjamin/9781847923868), edited and introduced by Hannah Arendt, translated by Harry Zohn.
- Peter Osborne and Matthew Charles, ["Walter Benjamin"](https://plato.stanford.edu/entries/benjamin/), *Stanford Encyclopedia of Philosophy*.
- John Berger et al., [*Ways of Seeing*](https://www.penguin.co.uk/books/56465/ways-of-seeing-by-berger-john/9780141035796), first published in 1972.
- Richard Hollis, ["Ways of Seeing: Book Design"](https://www.richardhollis.com/book-design/ways-of-seeing/), on the book's collaborative visual construction.
- Jonathan Conlin, ["Image lib: John Berger's Ways of Seeing"](https://www.bfi.org.uk/sight-and-sound/features/image-lib-john-berger-ways-seeing), *Sight and Sound*.
- Jonathan Conlin, ["Lost in Transmission? John Berger and the Origins of Ways of Seeing (1972)"](https://doi.org/10.1093/hwj/dbaa020), *History Workshop Journal*.
### Entropy: Counting Your Way to the Second Law
- URL: https://mohammadshaker.com/en/blog/entropy
- Date: 2026-07-03T00:00:00.000Z
- Tags: Physics, Information Theory, Thermodynamics, Authored with an LLM
Entropy is the most misunderstood word in science — 'disorder' is only half the story. This intuition-first explainer rebuilds entropy from scratch by counting: macrostates versus microstates, why '50 heads' swamps 'all heads', Boltzmann's S = k ln W and Shannon's H = -Σ p log p as the same idea in different clothes, why the second law is just statistics, and how temperature is really an exchange rate between energy and entropy. No physics background required.
#### Content
Entropy might be the most borrowed word in science. It shows up in physics, in
information theory, in essays about aging, in tweets about messy desks. And
almost every time it is used, it is used to mean one thing: disorder. Things
fall apart. Rooms get messier. Ice melts, smoke spreads, coffee cools. Entropy
is the universe's tax on tidiness.
That story is not wrong, exactly. It is a half-truth, the way "a car is a thing
that makes noise" is a half-truth. It points at something real while hiding the
mechanism completely. And the hidden mechanism is not just interesting — it is
almost embarrassingly simple. Entropy is about **counting**. Once you see the
counting, the second law of thermodynamics stops being a mysterious cosmic rule
and becomes something closer to a tautology: the obvious thing that had to
happen, happening.
The promise of this post is exactly that. By the end, the second law will not
feel like a law that nature obeys. It will feel inevitable — the way "if you
shuffle a sorted deck, it probably comes out unsorted" is inevitable. We will
get there the slow, honest way: by counting coins before we ever write down a
formula.
## The most misunderstood word in science
Start with why "disorder" fails.
Pour oil into water and shake. Wait. The oil and water separate into two clean
layers — the oil on top, the water below. To your eyes this looks *more*
ordered than the shaken-up mess you started with. And yet entropy has gone
**up**, not down. Water freezing into a crystal of ice looks like order
snapping into place, but freeze it in the right conditions and the entropy of
the whole system still climbs. If entropy were really just "how messy it
looks," these examples would run backward.
So "disorder" is the wrong handle. Here is the right one, and it is worth
reading twice:
> Entropy counts the number of distinct microscopic ways a system could be
> arranged while still looking the same to you from the outside.
The more ways there are to be in a situation, the higher the entropy of that
situation. That is the entire idea. Everything else in this post — gases,
temperature, information, the arrow of time — is a consequence of that one
sentence. The oil separates because there are astronomically more molecular
arrangements consistent with "separated layers" than with "evenly mixed" once
you account for how oil and water molecules actually attract each other. Nature
is not seeking tidiness. It is doing something much dumber and much more
powerful: it is wandering into whatever situation has the most ways to happen.
To make that precise, we need two words.
## Macrostates and microstates
Forget molecules for a moment. Flip four coins and lay them in a row.
A **microstate** is the exact, fully-specified arrangement: this coin heads,
that coin tails, and so on. `H T H H` is one microstate. `T H H H` is a
different microstate. There are `2 x 2 x 2 x 2 = 16` microstates for four
coins, and — this is the crucial assumption — nature has no favourites. Every
single microstate is equally likely. `H H H H` is exactly as probable as
`H T H T`. No arrangement is special.
A **macrostate** is the coarse, zoomed-out description — the thing you actually
notice. For coins, a natural macrostate is simply *how many heads there are*.
You do not care which coins are heads, only the count. "Two heads" is a
macrostate. It does not name a single arrangement; it names a whole *bucket* of
microstates.
Here is where the magic hides. Microstates are all equally likely. Macrostates
are not — because different macrostates hold different numbers of microstates.
Let us just list them for four coins.
| Macrostate (heads) | Microstates in the bucket | Count (W) |
| ------------------ | ------------------------------------- | --------- |
| 0 | TTTT | 1 |
| 1 | HTTT, THTT, TTHT, TTTH | 4 |
| 2 | HHTT, HTHT, HTTH, THHT, THTH, TTHH | 6 |
| 3 | HHHT, HHTH, HTHH, THHH | 4 |
| 4 | HHHH | 1 |
The counts are `1, 4, 6, 4, 1`, and they add up to 16 — every microstate is
accounted for. The all-heads macrostate contains exactly one arrangement. The
two-heads macrostate contains six. So even though every microstate is equally
likely, you are **six times** more likely to see "two heads" than "all heads,"
purely because the two-heads bucket is bigger.
That number in the last column — the count of microstates in a macrostate — is
what physicists call `W` (from the German *Wahrscheinlichkeit*, "probability").
`W` is the raw material of entropy. Hold onto it.
Flip a handful of coins yourself and watch the buckets fill. Every flip is a
microstate; each one lands in a macrostate — its heads count. The middle
macrostates swamp the edges, and they only pull further ahead as you add coins.
## Counting arrangements
Four coins is a toy. The real world has enormous numbers of parts, and when the
numbers get big, the counts do not just grow — they explode, and they explode
lopsidedly.
Go from four coins to one hundred. Now "all heads" is still exactly one
arrangement. But "fifty heads" is the number of ways to choose which 50 of the
100 coins are the heads — written `C(100, 50)` — and that number is monstrous.
Watch how fast the bucket sizes climb as you move from the edge toward the
middle.
| Macrostate (heads, of 100) | Ways to arrange it (W, approx) |
| -------------------------- | ------------------------------ |
| 100 | 1 |
| 90 | 1.7 x 10^13 |
| 70 | 2.9 x 10^25 |
| 50 | 1.0 x 10^29 |
The jump from "all heads" to "half heads" is a factor of one hundred billion
billion billion. If you flip a hundred fair coins, you will land somewhere near
fifty heads not because fifty is preferred, but because the fifty-ish
macrostates hoard essentially all of the arrangements. The extremes are lonely.
The middle is crowded beyond comprehension. This lopsided pile-up is called the
**binomial explosion**, and it is the engine behind everything that follows.
Now, these `W` values are already awkward — `10^29` is not a number you can hold
in your head, and for a real gas the exponent has 23 digits of its own. So
physicists do the sensible thing and take a **logarithm**. Entropy is defined as
`S = k ln W`
which is **Boltzmann's formula**, carved (in slightly different notation) on his
gravestone. `S` is entropy, `W` is the count of microstates in your macrostate,
`ln` is the natural logarithm, and `k` is a tiny conversion constant
(Boltzmann's constant) that puts the answer in the physical units temperature
will later want.
Why a logarithm? Two reasons, both intuitive.
First, it tames the giants. The logarithm of `10^29` is about 67. It turns
inhuman numbers into numbers you can reason about, and it turns the lopsided
explosion into a smooth, gentle curve.
Second — and this is the deep reason — **a logarithm turns multiplication into
addition**, and entropy needs to add. Put two independent systems side by side:
a box of gas here, a box of gas there. If the first has `W1` possible
microstates and the second has `W2`, then the combined system has `W1 x W2`
microstates, because every arrangement of the first can pair with every
arrangement of the second. Multiplication. But we want *entropy* to behave like
a normal physical quantity — the entropy of "both boxes" should be the entropy
of one plus the entropy of the other. And a logarithm delivers exactly that:
`ln(W1 x W2) = ln W1 + ln W2`
Counting multiplies; the log makes entropy add. That is not a bookkeeping trick.
It is the reason the formula has a logarithm in it at all.
## Entropy as surprise
Here the story takes a turn that surprised even the people who discovered it.
The same mathematics that counts gas arrangements also measures **information**
— your uncertainty, your surprise, the number of yes-or-no questions you would
need to pin something down.
Claude Shannon, founding information theory in 1948, wanted a number for "how
surprising is a message?" His answer, for a set of outcomes with probabilities
`p`, was:
`H = -Σ p log p`
That sum looks abstract, so build it from the felt sense of surprise. A surprise
should be large when something improbable happens and zero when something
certain happens. The quantity that does this is `-log p`, the **surprise** of an
outcome with probability `p`. Measure the logarithm in base 2 and the unit is
the **bit** — one bit is the surprise of a fair coin landing the way it did, the
information in a single yes-or-no answer.
| Outcome | Probability | Surprise (bits) |
| -------------------------- | ----------- | --------------- |
| A fair coin lands heads | 1/2 | 1 |
| Two fair coins both heads | 1/4 | 2 |
| A fair die shows a 6 | 1/6 | 2.58 |
| The sun rises tomorrow | ~1 | ~0 |
Read the pattern. Halving the probability adds exactly one bit of surprise —
one more yes-or-no question you would have had to ask. "The sun rises" carries
essentially no information because you already knew it; its surprise is zero.
Shannon's `H` is just the *average* surprise across all the outcomes, weighted
by how often each occurs. That is what `-Σ p log p` says out loud: for each
outcome, take its surprise `-log p`, weight it by its probability `p`, and add
them up.
Build that sum with your own hands. Each slider sets how likely one outcome is;
the bar for each outcome is as wide as its probability and as tall as its
surprise, so the total shaded area is exactly the average surprise, `H`. Make one
outcome nearly certain and watch `H` collapse toward zero; spread the weight
evenly across all four and it climbs to its maximum.
Think of the game of twenty questions. Every good yes-or-no question ideally
splits the remaining possibilities in half, and each split is worth one bit. A
space of a million possibilities needs about twenty questions to nail down,
because `2^20` is about a million. Shannon entropy is precisely the average
number of such questions — the average number of bits — you need to identify the
outcome. A predictable weather forecast (almost always sunny) needs almost no
questions and has low entropy. A genuinely uncertain one (fifty-fifty rain) has
high entropy: you learn a full bit when the day resolves.
Now watch the two entropies snap together. Suppose all `W` microstates are
equally likely, which is exactly Boltzmann's setup. Then each has probability
`p = 1/W`, and Shannon's formula collapses:
`H = -Σ (1/W) log(1/W) = log W`
That is Boltzmann's `ln W` again, give or take the base of the logarithm and the
constant `k`. Boltzmann counts equally-likely arrangements; Shannon generalizes
to arrangements that are not equally likely. They are the same act of counting,
wearing different clothes. Thermodynamic entropy is missing information about
the microstate — it is how many yes-or-no questions you *cannot* answer about a
system when all you know is its temperature, pressure, and volume.
## The second law is just statistics
Now we cash it all in. Picture a box split down the middle by a removable
partition. Every gas molecule starts trapped on the left. Slide the partition
out. What happens? The gas fills the box. It always fills the box. It never,
ever gathers itself back onto the left while you watch.
```
before (partition in place) after (partition removed)
+-------------+-------------+ +---------------------------+
| o o o o o o | | | o o o o o o |
| o o o o o o | (empty) | --> | o o o o o |
| o o o o o o | | | o o o o o o |
+-------------+-------------+ +---------------------------+
all N molecules on left spread across the whole box
```
Slide the partition out yourself and watch. Each particle just bounces blindly —
nothing pushes them rightward — yet the arrangement meter climbs from 0 toward 1
and stays there, because there are vastly more ways to be spread across the box
than crammed onto the left.
Here is the part that trips everyone up: **nothing in the laws of physics
forbids** all the molecules from rushing back to the left. Newton's laws run
just as happily backward as forward. No force pushes the gas outward. Each
molecule wanders at random, oblivious. So why does "spread out" always win?
Counting. For each molecule, "left half" versus "right half" is one more coin
flip. "All molecules on the left" is the all-heads macrostate — exactly one
bucket, staggeringly outnumbered. For `N` molecules, the probability of finding
them all back on the left at any instant is `(1/2)^N`. For a hundred molecules
that is already one in `10^30`. For a real breath of gas, `N` is around
`6 x 10^23` — Avogadro's number — and `(1/2)^N` is a decimal point followed by
more zeros than there are atoms in the galaxy before the first nonzero digit.
It is not that it is unlikely. It is that it will not happen before the universe
ends, not once, not ever.
So the gas spreads for the same reason a hundred flipped coins land near fifty
heads. "Spread out" macrostates outnumber "clustered" ones so overwhelmingly
that the system, wandering blindly among equally-likely microstates, is
practically certain to be found in a spread-out one. Follow the whole chain of
reasoning:
```mermaid
flowchart LR
micro["one microstate
(exact arrangement)"]:::config
macro["one macrostate
(what you measure)"]:::client
count["count the microstates
W in the bucket"]:::storage
prob["probability
W / total"]:::service
law["the state you see
= the biggest bucket"]:::external
micro -->|"group by what looks the same"| macro
macro -->|"tally the arrangements"| count
count -->|"divide by the total"| prob
prob -->|"overwhelming majority wins"| law
```
That is the **second law of thermodynamics**: the entropy of an isolated system
tends to increase, because the system keeps stumbling into ever-larger buckets
for no reason other than that larger buckets are larger. It is not a
commandment. It is a head count. The reason a broken cup never reassembles, the
reason smoke never crawls back into the cigarette, the reason you remember the
past and not the future, all trace to the same asymmetry: there are vastly more
ways to be spread out than gathered up. Time's arrow, in this telling, is not a
fundamental law at all. **The arrow of time is a probability gradient** — the
direction that points from smaller buckets toward bigger ones.
## Temperature is an exchange rate
One idea is left, and it is the one that makes entropy *useful* instead of
merely true. What is temperature, really?
Most people picture temperature as "amount of heat." That is another
half-truth. The precise definition is stranger and better:
`1 / T = dS / dE`
In words: temperature tells you **how much entropy a system gains for each unit
of energy you pour into it**. `dS / dE` is the rate of that trade — extra
microstates unlocked per extra joule. Temperature is the reciprocal of that
rate. A system is "hot" when adding energy barely raises its entropy, and "cold"
when adding the same energy opens up a flood of new arrangements.
That inversion sounds backwards until you feel it. A hot object is already
buzzing with energy and already has enormous numbers of accessible microstates;
one more joule is a drop in an ocean, adding only a few new arrangements, so its
entropy climbs slowly — high temperature, small `dS/dE`. A cold object is
starved of energy and cramped in its options; that same joule is a windfall,
unlocking a whole new range of arrangements, so its entropy jumps — low
temperature, large `dS/dE`.
| System | Temperature | Entropy gained per joule (1/T) |
| ------------- | ----------- | ------------------------------ |
| Hot coffee | high | small |
| Cool room air | low | large |
Now let one joule of heat leave the coffee and enter the room. The coffee loses
a *little* entropy (it was hot, so energy was entropically cheap there). The
room gains a *lot* of entropy (it was cool, so that same joule buys much more).
Add them: the total entropy of coffee-plus-room goes **up**. That is the only
reason heat flows from hot to cold. Not because heat "wants" to move, but
because moving it from where energy is entropically cheap to where it is
entropically expensive increases the total number of ways the whole system can
be arranged — the second law again, wearing an apron.
Temperature, then, is an **exchange rate** between energy and entropy, exactly
like a currency rate between dollars and euros. Heat flows in the direction that
buys more total entropy per joule spent, and it keeps flowing until the two
exchange rates match — until both objects report the same temperature. That
matched rate is what "thermal equilibrium" means. When your coffee and your room
reach the same temperature, it is because there is no longer any entropic profit
in moving energy either way.
## Entropy in everyday life
Step back and the same counting is everywhere, once you know to look for it.
**Cream in your coffee.** A drop of cream starts as a compact blob — one tight
cluster of microstates. Stir, and the cream molecules explore the whole cup.
There are unfathomably more arrangements with the cream mixed throughout than
gathered in a blob, so mixed is where it goes and mixed is where it stays. You
have never seen coffee spontaneously un-mix, for the identical reason you have
never seen a hundred coins un-flip back to all heads.
**A shuffled deck.** A deck of cards has `52!` possible orderings — about
`8 x 10^67`, more than the number of atoms in the solar system. Exactly one of
them is "brand-new-from-the-box order." When you shuffle, you are wandering
among those `10^67` microstates, and the overwhelming majority of them look
scrambled. The deck "becomes disordered" not because shuffling seeks disorder
but because ordered arrangements are a vanishing speck in a sea of scrambled
ones. Counting, again.
**Why engines waste heat.** This is the one that reshaped the industrial world.
You might hope to take heat and turn *all* of it into useful work — a perfect
engine. You cannot, and entropy is the reason. Extracting work from heat means
moving energy around, and the second law insists the total entropy must not
fall. To honour that bookkeeping, an engine has to dump some of its heat into a
cold reservoir — the exhaust, the radiator, the surrounding air — paying an
entropy bill it can never escape. That unavoidable waste (formalized by Carnot's
limit) is why car engines run hot, why power plants have cooling towers, and why
a perpetual motion machine is not just hard to build but forbidden by counting
itself.
So here is the whole idea, gathered back into one breath. Entropy is not
disorder. Entropy is the number of ways — the size of the bucket you happen to
be standing in. Systems drift toward bigger buckets because bigger buckets are
bigger, and that drift, multiplied across a mole of molecules, is so
overwhelming that we dignify it with the name *law*. The second law, the arrow
of time, the flow of heat, the impossibility of a perfect engine — all of it is
what counting looks like when the numbers get astronomically large.
This post is a starting point, not a finish line. It is built to grow: each of
the sections above — coins and macrostates, the binomial explosion, surprise in
bits, the gas in a box, temperature as an exchange rate — is a natural home for
an interactive lab, a little widget where you can flip the coins yourself,
slide the partition out, and watch the entropy climb in real time. When those
labs land, come back. The best way to believe the second law is inevitable is to
try, and fail, to make it run backward.
### From L3 to L7: How the Network Layers Actually Work
- URL: https://mohammadshaker.com/en/blog/from-l3-to-l7-how-network-layers-work
- Date: 2026-07-01T00:00:00.000Z
- Tags: networking, OSI, TCP/IP, load balancing, TLS, HTTP, architecture, Authored with an LLM, Engineering
A visual walk up the network stack from L3 (Network/IP) to L7 (Application/HTTP): what each layer actually does, the real protocols and data unit at each level, how one request is wrapped header-by-header on the way down and unwrapped on the way up, and the practical payoff every backend engineer trips over eventually — L4 versus L7 load balancing.
#### Content
Every backend engineer eventually hits the same wall. A load balancer routes `/checkout` to the wrong fleet, or TLS terminates in a place you did not expect, or a health check passes at one level and the app is dead at another. Underneath all of it is a single idea most people learned as a diagram they immediately forgot: the network stack is a pile of layers, and each layer does exactly one job.
This post is the diagram, drawn properly. We walk up the stack from **L3 (Network)** to **L7 (Application)**, say what each layer does and why, and then trace one real `https://example.com/page` request all the way down one machine and up another. By the end, the L4-versus-L7 load-balancing question that trips up so many designs answers itself.
## What "Layers" Even Are
Forget the acronyms for a second. A network layer is a job with a strict contract: it takes a chunk of data from the layer above, does its one job, wraps that data in its own header, and hands it to the layer below. That is the whole trick.
The header is the key. Each layer writes a small block of its own bookkeeping in front of whatever it received, and it does **not** look inside what it received. To L3, the entire HTTP request plus its TLS encryption plus its TCP bookkeeping is just an opaque blob of bytes it has to move to an IP address. To L7, the IP routing that got the bytes there is invisible plumbing it never thinks about.
One sentence to hold onto: **a message travels DOWN the stack on the sender, getting wrapped in one more header at every layer, then travels UP the stack on the receiver, getting one header stripped at every layer, until the layer that put a header on is the layer that reads it.** Sender wraps, receiver unwraps, symmetrically.
A quick honesty note before the tour. The classic OSI model has seven layers; the model the internet actually runs on (TCP/IP) collapses the top three into one. We use the OSI numbering because it is the shared vocabulary, but I will flag where L5, L6, and L7 blur together in practice instead of pretending they are cleanly separate boxes.
## The Layers, L3 to L7
Here is the whole tour in one table. Each layer, its one job, the name for its chunk of data (the **PDU**, Protocol Data Unit), and the protocols you actually meet.
| Layer | Name | One job | PDU | Example protocols |
| --- | --- | --- | --- | --- |
| L3 | Network | Address + route across networks, hop by hop | Packet | IPv4, IPv6, ICMP |
| L4 | Transport | Deliver end-to-end between processes, via ports | Segment (TCP) / Datagram (UDP) | TCP, UDP, QUIC |
| L5 | Session | Open, maintain, resume, and tear down a conversation | (data) | TLS session/resumption, SOCKS, RPC session |
| L6 | Presentation | Encode, serialize, compress, encrypt the payload | (data) | TLS record encryption, gzip, UTF-8, JSON/protobuf |
| L7 | Application | Speak the protocol the app actually cares about | Message / Data | HTTP/1.1, HTTP/2, HTTP/3, gRPC, DNS, WebSocket, SMTP |
Now each layer in words.
### L3 — Network: addressing and routing
L3 answers one question: *which machine, and how do I get the bytes there?* It gives every host an **IP address** and moves a **packet** from source to destination across a chain of routers, one **hop** at a time. Each router reads only the destination IP, consults its routing table, and forwards to the next router closer to the goal. No single router knows the whole path; each one just knows the next hop.
The protocols are **IPv4** and **IPv6** for addressing, and **ICMP** for control messages (this is what `ping` and `traceroute` ride on). The defining property of L3 is what it does *not* promise: it is **best-effort**. A packet can be dropped, duplicated, delayed, or arrive out of order, and L3 shrugs. It never retransmits, never reorders, never confirms delivery. If you want guarantees, that is somebody else's job — specifically L4's.
### L4 — Transport: ports and end-to-end delivery
L3 got the bytes to the right *machine*. L4 gets them to the right *program* on that machine, and decides how reliable the delivery is. It introduces the **port**: a 16-bit number so the OS knows that these bytes belong to your web server on `:443` and those belong to SSH on `:22`. An IP address plus a port is a specific conversation endpoint.
Two protocols dominate, and they are opposites.
**TCP** (its PDU is a **segment**) is the reliable one. Before any data flows it does a three-way handshake (SYN, SYN-ACK, ACK) to establish a connection. Then it numbers every byte, acknowledges what it receives, retransmits what got lost, reorders what arrived scrambled, and throttles itself under congestion. TCP hands the application an ordered, gap-free, reliable byte stream out of L3's unreliable packet soup. That reliability costs round-trips and state.
**UDP** (its PDU is a **datagram**) is the fire-and-forget one. A port, a length, a checksum, and go. No handshake, no ordering, no retransmit. It is a thin shim over L3 that only adds ports. That sounds worse until you want low latency and can tolerate loss — live video, DNS lookups, game state — where waiting for a retransmit is worse than dropping the frame.
**QUIC** is the modern twist: it rebuilds TCP's reliability, ordering, and congestion control *on top of UDP*, folds TLS in, and fixes head-of-line blocking. It is what HTTP/3 runs on. Structurally it lives at L4, but it deliberately smears into L5/L6 by owning the session and the encryption too — which is exactly the kind of blur the clean OSI diagram hides.
### L5 — Session: the conversation's lifecycle
L5 manages the *conversation* as a thing with a beginning, a middle, and an end: establishing it, keeping it alive, and tearing it down. In pure TCP/IP terms this layer barely exists as its own box — TCP already owns connection setup and teardown — which is why L5 is the fuzziest of the five.
Where it earns its keep conceptually is **session state and resumption**. When TLS lets a returning client skip the full handshake via a session ticket or resumption, that is session-layer thinking: *we have talked before, let us not start from zero.* When an RPC framework keeps a logical session across several transport connections, that is L5. Be honest about this one: in the model the internet runs on, L5 is less a physical layer and more a role that TCP, TLS, and the application quietly share.
### L6 — Presentation: encoding, serialization, encryption
L6 is about the *shape* of the bytes, not their delivery. Its job is to turn the application's in-memory objects into a wire format both sides agree on, and back again: **character encoding** (UTF-8), **serialization** (JSON, protobuf, MessagePack), **compression** (gzip, Brotli), and **encryption** (the TLS record layer that turns plaintext into ciphertext before it ever reaches L4).
This is the layer that means the server can read what the client wrote. If the client marshals a struct to protobuf and the server expects JSON, no amount of perfect L3 routing and L4 delivery saves you — the bytes arrived flawlessly and are still gibberish. TLS record encryption sits here too: the payload is encrypted at L6 on the way out and decrypted at L6 on the way in, which is why a plain L4 load balancer downstream sees only opaque ciphertext.
### L7 — Application: the protocol the app speaks
L7 is the protocol your application actually talks. It is the layer humans and services think in: **HTTP** verbs and headers and status codes, **gRPC** method calls, **DNS** queries, **WebSocket** frames, **SMTP** commands. When you write `GET /page HTTP/1.1\r\nHost: example.com`, that string *is* the L7 payload. Everything below exists to carry it.
The evolution matters: **HTTP/1.1** is text over one TCP connection with head-of-line blocking; **HTTP/2** multiplexes many streams over one TCP connection (but TCP's own head-of-line blocking remains); **HTTP/3** moves onto QUIC to kill that last blocking problem. Same L7 semantics — methods, headers, status codes — carried over three different lower stacks. That last sentence is the whole point of layering, and we will come back to it.
## Encapsulation: Headers Nested Like Russian Dolls
Here is the picture the table cannot show: what the bytes literally look like as they go down the stack. Each layer wraps the layer above in its own header. By the time an HTTP message hits the wire, it is a payload buried inside header after header.
```
DOWN the sender's stack (each layer WRAPS the one above)
L7 Application: the actual message
+---------------------------------------------------+
| GET /page HTTP/1.1 Host: example.com ... |
+---------------------------------------------------+
L6 Presentation: serialize + TLS-encrypt the payload
+---------------------------------------------------+
| [ TLS record: <> ] |
+---------------------------------------------------+
L4 Transport: prepend TCP header (src+dst PORT, seq)
+----------+----------------------------------------+
| TCP hdr | [ encrypted application payload ] |
| :51522 | |
| ->:443 | |
+----------+----------------------------------------+
\_________ TCP calls this a SEGMENT _____/
L3 Network: prepend IP header (src+dst IP ADDRESS)
+---------+----------+-------------------------------+
| IP hdr | TCP hdr | [ encrypted payload ] |
| 1.2.3.4 | :51522 | |
| ->9.8.7 | ->:443 | |
+---------+----------+-------------------------------+
\______ IP calls this whole thing a PACKET/
L2 Link: wrap in a frame (src+dst MAC) for the next hop
+--------+---------+---------+-----------------+------+
| ETH hdr| IP hdr | TCP hdr | [ payload ] | FCS |
| MACs | | | | crc |
+--------+---------+---------+-----------------+------+
\_______ L2 calls this whole thing a FRAME _/
```
Read it top to bottom on the sender: the HTTP message gets encrypted, then a TCP header (ports) is bolted on the front to make a segment, then an IP header (addresses) to make a packet, then an Ethernet frame (MACs) for the physical hop. Each outer layer treats everything inside as an opaque payload it must not open.
On the receiver, the exact reverse — **decapsulation**. The frame arrives, L2 strips the Ethernet header and checks the CRC, hands the packet up. L3 strips the IP header, hands the segment up. L4 strips the TCP header and reassembles the byte stream, hands the payload up. L6 decrypts. L7 finally reads `GET /page`. Every header is read by exactly the layer that wrote it, and stripped in reverse order to how it was added — last on, first off, like a stack.
## The Journey of One Request
Now the payoff: one `https://example.com/page` request, end to end. The browser (client, front-end) walks the data *down* its stack; the bytes cross the wire; the server (back-end) walks them *up* its stack, handles the request, and sends the response back down its own stack and up the client's. Two diagrams of the same journey — a mermaid sequence and a plain ASCII version — because the shape is worth seeing twice.
```mermaid
sequenceDiagram
participant C7 as Client L7 HTTP
participant C6 as Client L6 TLS
participant C4 as Client L4 TCP
participant C3 as Client L3 IP
participant NET as L3 Network (routers, hop by hop)
participant S3 as Server L3 IP
participant S4 as Server L4 TCP
participant S6 as Server L6 TLS
participant S7 as Server L7 App
C7->>C6: GET /page (plaintext message)
C6->>C4: encrypt -> TLS record
C4->>C3: segment (src:51522 dst:443, seq)
C3->>NET: packet (src 1.2.3.4 dst 9.8.7.6)
NET->>S3: routed hop by hop, best-effort
S3->>S4: strip IP header -> segment
S4->>S6: reassemble stream -> TLS record
S6->>S7: decrypt -> GET /page
S7-->>S6: 200 OK + HTML
S6-->>S4: encrypt response
S4-->>S3: segment back
S3-->>NET: packet back
NET-->>C3: routed back
C3-->>C7: unwrapped up the stack -> render
```
```
ASCII — same request, DOWN the client and UP the server
CLIENT (front-end) SERVER (back-end)
================== =================
L7 HTTP GET /page ------. ,--> L7 App handler
| | (reads GET /page)
L6 TLS encrypt --------| |--- L6 TLS decrypt
| |
L4 TCP seg :51522->:443-| |--- L4 TCP reassemble
| |
L3 IP pkt 1.2.3.4----- | |--- L3 IP strip header
90->9.8.7.6 | |
v |
+===================================+
| L3 NETWORK: routers, hop by |
| hop, best-effort, may drop |
+===================================+
\____ across the wire ____/
RESPONSE travels the mirror image:
server L7 200 OK -> L6 encrypt -> L4 segment -> L3 packet
-> wire -> client L3 -> L4 -> L6 decrypt -> L7 -> browser renders
```
Trace the down-path once in prose. L7 produces the request line and headers. L6 encrypts them into a TLS record so nothing on the path can read them. L4 slices that into TCP segments, stamps the source and destination ports, and takes responsibility for getting every byte there in order. L3 wraps each segment in an IP packet with source and destination addresses and hands it to the network, which routes it hop by hop with zero delivery promises. The server does the same steps in reverse, its L7 handler finally sees `GET /page`, produces `200 OK` with the HTML, and the whole thing runs backward down the server's stack and up the client's until the browser paints the page.
Notice what each side did *not* do. The server's L3 never decrypted anything — that is L6's job. The client's L7 never picked a route — that is L3's job. Every layer minded its own business. That discipline is not an accident; it is the entire design, and it has a name we will get to.
## Before / After: L4 vs L7 Load Balancing
Here is where the layer you operate at becomes a real engineering decision with real money attached. A load balancer sits between clients and a fleet of servers and spreads traffic across them. *Which layer it reads at* changes everything it can do.
**L4 load balancing** operates at the transport layer. It sees IP addresses and ports and nothing else — the payload is opaque bytes (and if TLS is end-to-end, it is encrypted opaque bytes it *could not* read even if it wanted to). It picks a backend by a cheap rule (hash the source IP, round-robin the connection) and shovels segments through. It never terminates TLS, never parses HTTP, barely touches the CPU. Fast, cheap, blind.
**L7 load balancing** operates at the application layer. To read HTTP it must first **terminate TLS** (decrypt the traffic, which costs CPU and puts the LB inside your security boundary), then parse the HTTP request and route on its *content*: send `/api/*` to the API fleet, `/img/*` to the static fleet, retry idempotent requests on a failed backend, rewrite paths, add headers, do sticky sessions by cookie. Powerful, content-aware, more expensive.
```
BEFORE — L4 load balancer (transport, blind to content)
client --TLS--> [ L4 LB ] reads only IP:port, payload opaque
| hash(src_ip) % N
+--> backend chosen by connection, NOT by URL
every path (/checkout, /img, /api) lands on the same pool
AFTER — L7 load balancer (application, reads the request)
client --TLS--> [ L7 LB ] terminates TLS, parses HTTP
| route on Host + path + headers
+--> /checkout -> payments fleet
+--> /img/* -> static fleet
+--> /api/* -> api fleet (+ retries, rewrites)
```
### One scenario: route `/checkout` to the payments fleet
You want checkout traffic to land on a hardened payments fleet, separate from everything else. Watch how the layer decides whether that is even possible.
The **L4 way.** The load balancer sees `9.8.7.6:443` and a stream of encrypted bytes. It has *no idea* whether a given connection is `/checkout` or `/cat.jpg` — that information lives in the HTTP request, which is L7 data encrypted at L6, three layers above where the L4 LB is looking. So it physically cannot route by path. Your only options are crude: give `/checkout` its own hostname resolving to a *different IP*, and let L4 route by destination IP; or run the payments service on a different port. The upside: near-zero latency, no TLS CPU cost on the LB, and the LB never sees a plaintext card number, so its blast radius on a compromise is small. The downside: you have pushed the routing decision out to DNS/IP topology and lost all per-request control.
The **L7 way.** The load balancer terminates TLS, reads `POST /checkout HTTP/2`, matches the path, and forwards to the payments fleet — over one shared hostname and IP, alongside all your other routes. It can additionally retry a failed checkout POST-that-is-safe-to-retry, rewrite the path, strip a header, and pin the user to a backend. The cost is real: TLS termination burns CPU and adds a little latency, and because the LB now decrypts payment traffic it sits *inside* your compliance and blast-radius boundary — a compromise there is far worse. You bought routing power and observability by paying CPU and widening the trust boundary.
That trade — **cheap/fast/blind (L4) versus powerful/content-aware/costly (L7)** — is the whole decision, and it falls directly out of which layer can *see* the information you want to route on. If the routing key lives in the HTTP request, no L4 config will ever reach it. That is not a limitation to work around; it is layering working exactly as designed.
## The Design Principles Underneath
The layered stack is not just a networking artifact. It is a clinic in the design principles we reach for in application code — because the same forces (change isolation, substitutability, single responsibility) produced the same answers. Mapped honestly, forcing nothing.
**SRP — Single Responsibility.** Each layer has exactly one job and never reaches into another's. L3 routes and only routes; it never retransmits. L4 delivers and only delivers; it never picks a route. L7 speaks the app protocol and never thinks about IP addresses. This is textbook single responsibility, and it is *why* the L4 load balancer physically cannot route by URL — reading the URL is L7's responsibility, and L4 does not do L7's job.
**DRY / encapsulation.** Every layer honors the identical contract: *take a payload, add your header, pass it down.* The header/payload split is defined once and reused at every level. L3 does not re-implement L4's delivery; L7 does not re-implement L3's routing. Reliability logic lives in exactly one place (TCP) instead of being copy-pasted into HTTP, DNS, and every other L7 protocol — which is precisely why they can all share it.
**IoC / DI — Inversion of Control and Dependency Injection.** A layer depends on the *interface* of the layer below, not its implementation. L7's contract with L4 is "give me a reliable ordered byte stream," and it does not care how. Swap TCP for QUIC, or IPv4 for IPv6, and HTTP is unchanged — the lower layer is *injected* underneath, and the upper layer never notices. That is exactly why HTTP/1.1, HTTP/2, and HTTP/3 can carry identical semantics over three different transports.
**PubSub / event-driven.** The base model is request-response, but some L7 protocols invert it. **WebSocket** and **MQTT** turn the connection into a channel where either side pushes messages when it has them — publish/subscribe, not ask/answer. This is where an L7 gateway earns its keep beyond routing: because it parses the application protocol, it can *fan out* one published message to many subscribed connections, or route by topic — something an L4 balancer, blind to the payload, could never do.
**MVC — mapped lightly, and honestly.** MVC is an application-architecture pattern, not a network-layer one, and pretending otherwise invents a fake correspondence. The honest map: L7 is the surface where the app's controller/view logic lives — it is the only layer that understands the *request as the application means it*. L4 and L3 are pure plumbing the app deliberately does not model; the whole point of layering is that your controller never thinks about packets. So MVC does not map onto the stack cleanly, and that is fine — it maps onto the *top* of it, and everything below is the transport the model never sees.
## The One Idea to Keep
Strip everything else away and this is what remains: each layer does one job, wraps the layer above in its own header on the way down, and reads only its own header on the way up. Routing lives at L3, delivery at L4, the app's protocol at L7 — and the reason your L4 load balancer cannot see a URL, your L7 load balancer can but pays for it in CPU and trust boundary, and your HTTP code survives a swap from TCP to QUIC untouched is all the same reason: **each layer minds exactly its own business, and depends only on the contract of the one below.** Learn the stack once and half the "why is my traffic going *there*?" questions answer themselves.
### Entity APIs as the Spine of Server-Driven Screens
- URL: https://mohammadshaker.com/en/blog/entity-api-sections-multisections
- Date: 2026-06-26T00:00:00.000Z
- Tags: SDUI, API Design, Product Engineering, Authored with an LLM, Engineering
A backend entity read API returns raw entities; the client folds each entity into a ScreenModel -- a named map of typed sections rendered through one reusable host. This post walks the FE<->BE seam, shows how a screen becomes JSON ({order, sections}), and explains how that collapses per-screen maintenance and turns A/B testing into reordering, swapping, or flag-gating sections without a client release.
#### Content
Most mobile screens are not unique. They are an ordered list of typed blocks -- a header, a list, a call-to-action -- assembled from data the backend already owns. If you accept that, a screen stops being bespoke layout code and becomes a value: a map of sections plus an order.
## Context
This post describes a pattern with two pieces. An **entity API** is a uniform backend read surface: ask it for entity X for the current user and it returns raw JSON, nothing about layout. A **server-driven screen** is the consequence: instead of the app hard-coding what each screen looks like, the backend (or a thin client builder) hands the app a description of the screen as data, and the app renders it.
The glue between them is a `ScreenModel`: an ordered map of typed sections. The client reads an entity, folds it into a `ScreenModel`, and a single reusable host renders every tab through one section registry. In this setup the backend service is a Python/Flask API, and the client app is React Native with Expo, but the idea is stack-agnostic.
## The screen, visualized
Here is a home screen as a reader would see it: a progress header, a list of subjects, an upgrade call-to-action, and a streak indicator. Each visible block maps one-to-one to a named, typed section.
```
┌──────────────────────────────────────────┐
│ progress_header streak: 4 days * │
├──────────────────────────────────────────┤
│ Your subjects │
│ ┌────────┐ ┌────────┐ ┌────────┐ │
│ │ Math │ │ Science│ │ History│ │
│ └────────┘ └────────┘ └────────┘ │
├──────────────────────────────────────────┤
│ subscriptions_cta [ Upgrade to Pro ] │
├──────────────────────────────────────────┤
│ streak 4-day streak, keep! │
└──────────────────────────────────────────┘
```
A "multisection" screen is exactly this: a named map of typed sections rendered in `order`, optionally with an `aside` pane for two-pane desktop layouts.
```mermaid
flowchart TD
screen["Home screen"]:::client
model["ScreenModel {order, sections, aside?}"]:::config
s1["progress_header"]:::config
s2["subjects_list"]:::config
s3["subscriptions_cta"]:::config
s4["streak"]:::config
screen -->|"is rendered from"| model
model -->|"order[0]"| s1
model -->|"order[1]"| s2
model -->|"order[2]"| s3
model -->|"order[3]"| s4
```
## Before and after
The contrast is the whole argument. Before, each screen fetched its own bespoke endpoints and hand-built its own UI, so every tab reimplemented its own loading, empty, and error chrome. After, every screen reads through one uniform entity surface and renders through one `ScreenModel` host, so there is a single render path and chrome is handled once.
```
BEFORE AFTER
┌─────────────┐ ┌─────────────┐
│ Home tab │--GET /home-->API │ Home tab │
│ own fetch │ │ │
│ own chrome │ │ one entity │
├─────────────┤ │ read + │
│ Subject tab │--GET /subj-->API │ one host │--GET /entity-->API
│ own fetch │ │ │
│ own chrome │ │ chrome │
├─────────────┤ │ handled │
│ Offers tab │--GET /offer->API │ once │
│ own fetch │ │ │
│ own chrome │ └─────────────┘
└─────────────┘
```
```mermaid
flowchart LR
subgraph Before
h1["Home: bespoke fetch + chrome"]:::client
s1["Subject: bespoke fetch + chrome"]:::client
o1["Offers: bespoke fetch + chrome"]:::client
end
subgraph After
host["one ScreenModel host"]:::client
entity["one entity read surface"]:::networking
host --> entity
end
```
## Scenario: reorder the home screen for an experiment
Suppose product wants to test pushing the upgrade CTA above the subjects list. Before, layout was baked into the app: changing the order means a code change, a new release, and a wait for adoption, with no way to target a subset of users. After, the order lives in the `ScreenModel` JSON: reorder the array, swap a section variant, or flag-gate a section server-side. The change is live immediately and can be A/B tested.
```
BEFORE AFTER
layout in app code order in ScreenModel JSON
change -> build -> release edit array -> served live
no targeting per-user / A-B, no release
```
## One read surface: the entity API
Before sections or screens exist, the backend has to answer one question uniformly: give me entity X for this user. The entity API is a single declarative read surface, not a pile of hand-written handlers. Each entity is registered once and served under a uniform route.
| Entity | Route | Backing |
|---|---|---|
| subjects | `GET /entity/subjects` | model-backed |
| lesson tree | `GET /entity/tree?subject_id=` | computed |
| content item | `GET /entity/content?item_id=` | computed |
| feature flags | `GET /entity/feature-flags` | model-backed |
| preferences | `GET /entity/preferences` | model-backed |
| offers | `GET /entity/offers` | model-backed |
Some entities are model-backed (read a record, serialize it) and some are computed (run a function, optionally adding a server-side pay-gate). Either way the client sees the same shape, so it never learns six different fetch idioms.
```mermaid
flowchart LR
spec["resource spec (model-backed or computed)"]:::service
engine["entity read engine"]:::service
route["GET /entity/<name>"]:::networking
db["records / flags / preferences"]:::storage
gate["server-side pay-gate"]:::security
spec -->|"registered"| engine
engine -->|"reads"| db
engine -->|"computed adds"| gate
engine -->|"serves entity JSON"| route
```
## From entity to ScreenModel
A raw entity is not a screen -- it is the ingredients. A builder maps each entity into the section map and emits a `ScreenModel`: `{ order?, sections: Record, aside? }`. Each `SectionDef` is `{ type: string } & Record`, so a section is just a `type` tag plus the props that type consumes.
```json
{
"order": ["progress_header", "subjects_list", "subscriptions_cta"],
"sections": {
"progress_header": { "type": "progress_header", "streakDays": 4 },
"subjects_list": {
"type": "subjects_list",
"sectionTitle": "Your subjects",
"items": [{ "id": "math", "title": "Math", "cover_url": "math.png" }]
},
"subscriptions_cta": { "type": "subscriptions_cta", "event": "open_paywall" }
}
}
```
The builder is the only place that knows how a given entity becomes sections. Move it server-side later and nothing else changes: the host, the registry, and the JSON shapes already match.
```mermaid
flowchart TD
entity["subjects entity JSON"]:::config
builder["assembleHomeScreenModel"]:::client
tile["Subject -> SubjectTile"]:::client
model["ScreenModel {order, sections}"]:::config
entity -->|"raw items"| builder
builder -->|"maps each"| tile
tile -->|"packs into"| model
```
## One host renders every screen
If every screen is a `ScreenModel`, then only one component needs to know how to render screens. A data adapter seam exposes `fetchScreen(tab)` and `fetchEntity(name, params)`; the host receives a resolved `ScreenModel` and never knows where the JSON came from. A rejected fetch resolves to `null`, which triggers a fallback so a tab never renders blank.
The host resolves `order = model.order ?? Object.keys(model.sections)`, renders each section through a section registry inside a frame that owns the loading, error, empty, and unauthorized chrome centrally. The registry dispatches `section.type` and graceful-skips unknown types instead of throwing. No tab reimplements fetch states, and an unrecognized section degrades to nothing rather than crashing the screen.
```
FRONTEND BACKEND
┌──────────────────────────┐ ┌──────────────────────┐
│ fetchEntity(name) ───────┼── GET ─────>│ /entity/ │
│ │ │ spec / computed │
│ │ │ (+ pay-gate) │
│ assemble ScreenModel <───┼─ entity JSON┤ │
│ -> {order, sections} │ └──────────────────────┘
│ host: order ?? keys │
│ for each -> registry │
│ registry[type] -> block │
│ unknown -> skip │
│ reject -> fallback │
└──────────────────────────┘
```
## A/B testing by reshaping JSON
Because a screen is just `{order, sections}`, an experiment is a data edit, not a build. Reorder the array to push the CTA above the fold. Swap one section for a variant section of a different `type`. Toggle a section in or out by reading the feature-flags entity and dropping or keeping its key. All three are server-side JSON changes against already-registered section types, so no client release is required to run, ship, or roll back a variant.
```mermaid
flowchart LR
flags["feature flags"]:::storage
base["ScreenModel base {order, sections}"]:::config
vA["Variant A: cta last"]:::config
vB["Variant B: cta first + swapped type"]:::config
host["ScreenModel host"]:::client
flags -->|"assignment"| base
base -->|"reorder"| vA
base -->|"swap + flag-gate"| vB
vA -->|"same path"| host
vB -->|"same path"| host
```
## Design principles at work
- **Single responsibility (SRP).** Each entity returns one thing; each section renders one block; the host only orchestrates order and chrome. No component carries two jobs.
- **Don't repeat yourself (DRY).** One entity read engine and one `ScreenModel` host replace bespoke screens and the loading/empty/error chrome each used to repeat.
- **Inversion of control / dependency injection (IoC/DI).** The data adapter and the screen assembler are injected; the host does not know where JSON comes from, so client builders and server proxies are interchangeable.
- **Publish/subscribe, event-driven.** Sections emit events like `open_paywall`; the host or a runner subscribes and routes them, so a section never reaches into navigation or payments directly.
- **Model-view-controller (MVC).** The entity and `ScreenModel` JSON are the model, the section components are the view, and the host plus adapter are the controller.
## Why product and engineering both win
- **A new screen is mostly data.** With a uniform entity read surface and one host, building a screen is "fetch an entity, name some sections, pick an order."
- **Chrome is centralized.** Loading, error, empty, and unauthorized states live in the host once. No tab reimplements them, and none can render blank.
- **Failure degrades gracefully.** A rejected fetch falls back; an unknown section type is skipped, not fatal.
- **Experiments need no release.** Product reorders, swaps, or flag-gates sections in JSON; the client renders any valid `ScreenModel` through the same path.
- **The server migration is a seam, not a rewrite.** Host, registry, and JSON shapes already match, so moving the builder behind `fetchScreen(...)` is a single substitution.
### How I run /eil: a self-improving, token-conscious AI coding loop
- URL: https://mohammadshaker.com/en/blog/how-i-run-eil-self-improving-ai-coding-loop
- Date: 2026-06-26T00:00:00.000Z
- Tags: AI, agents, automation, developer-tools, Claude, Engineering, Authored with an LLM
A meta-orchestration loop that takes any implement / fix / build task and drives it to verified-green through a fleet of specialized AI agents — then improves itself, across every repo, after every run.
#### Content
## The problem with one-shot AI coding
Most AI coding is one prompt, one diff, and hope. For anything real — a feature, a cross-cutting fix, a migration — that falls apart: no plan, no tests actually run, no proof it works, and the next session re-learns the same lessons. /eil ("execute in a loop") is my answer: a meta-orchestration skill that owns the *loop discipline* and delegates the real work to a fleet of specialized sub-agents, driving any task to **verified green** without ever handing back a to-do list.
## What /eil actually is
One rule: state ONE explicit GOAL up front, then loop fix → run → verify until that goal is *verifiably* reached — the test suite green, an adversarial review clean, and every user-observable acceptance criterion demonstrated with a real artifact. "The code looks right" is never green; "the test exited 0 and I read the proof" is. It never asks clarifying questions — it picks the most reasonable interpretation, records a one-line rationale, and acts. The only thing it stops to surface is a genuine owner gate: a deploy, a production write, spend, or something irreversible.
## How it works: a staged fleet
/eil runs a fixed pipeline, each stage its own subagent with a structured hand-off. Scaffold makes an isolated git worktree off `develop` and seeds a notes doc (goal, acceptance criteria, success rows, and a `State` block). Optional pre-stages handle product positioning, premium UX design tied to a design system, and adversarial design QA — but only for customer-facing work. Then: plan (a principled design — SRP/DRY/IoC/DI/MVC/PubSub/YAGNI, remove-before-add, with options and a falsification test), an adversarial plan review that decomposes the plan into trackable units, implement, test, QA that eradicates whole *classes* of bugs rather than the one symptom, a visual-QA stage for UI that reads the actual rendered pixels against the design, and finally merge. A `State` block in the notes doc is the single source of truth: every stage reads it first and writes it last, so the whole pipeline is resumable and idempotent.
## Being token-conscious
A full fleet on a one-line bug is waste, so /eil is aggressively lean. It triages first: trivial work runs inline with no fleet at all; standard work collapses to a single plan-and-implement agent; only genuinely major work gets the whole pipeline. It activates only the stages a task needs — a pure backend bug skips the product, design, and visual stages entirely, and every skip is recorded with a reason. It scopes tests to the change's blast radius — a notifications-only fix runs the notifications tests, not signup or the entire suite — but escalates to the full suite the moment a change is cross-cutting, shared, or uncertain, and the change's own coverage test always runs. Each stage even picks its model by how hard its task actually is: a cheap fast model for mechanical scaffolding, a mid tier for normal work, the top reasoning tier only for genuinely hard problems, with a hard ceiling so it never over-spends.
## Making it self-improve
This is the part I'm most proud of. Every stage, before it returns, surfaces any *durable, generalizable* lesson it learned — not task findings, but "here's a gotcha, here's a better way, here's a convention." The orchestrator collects them. After the merge, a final stage harvests those lessons, filters for the ones that will actually matter next time, dedups against what's already written, and routes each to its home: the /eil definition itself, a specific agent's instructions, the best-matching *other* skill, or long-term memory. Then it propagates those edits to every repo in a manifest. So a lesson learned while fixing a bug in one repo upgrades the pipeline — and the relevant skills — in all of them, before the next run starts. The system that builds software also rewrites itself. Guardrails keep that safe: it's idempotent (no lesson means no change), capped per run, dedups before every write, asserts before writing, and only makes additive, revertible commits.
## Calling itself recursively
The loop closes on itself in two ways. Within a run, QA can route a finding *back* to planning — the pipeline re-plans, re-implements, and re-tests itself, bounded so it can't spin forever. Across runs, a hook fires at the end of every session: it persists what was learned and, if a bug was worked, files a prevention task for the whole *class* of that bug and then invokes /eil on exactly those tasks to implement the guard end-to-end. /eil fixes a bug, then calls itself to make that class of bug impossible next time.
## Tracking everything with beads
Work is tracked in beads — a lightweight issue database — never ad-hoc to-do lists. The first thing a non-trivial run does is decompose the ask into issues, and that issue set *is* the success contract: one bead per atomic, independently verifiable unit, every edge case its own bead. The loop drives ready → show → close as each goes green, commits and pushes per green bead, and stores durable knowledge as memory beads instead of a rotting notes file. The bead list is the contract; "done" means every bead closed *and* the close-out gates passed.
## A real end-to-end run
I built most of this with /eil itself. In one recent stretch I authored the self-improvement stage and a targets manifest, wired a learnings channel into every agent, and propagated the whole change across a six-repo cluster in a single pass — using a no-checkout git plumbing recipe (build the commit against the remote branch with a temporary index, push a throwaway branch, merge it server-side) so it's immune to worktree locks and concurrent agents. Mid-run I hit two real gotchas: a shell loop that didn't word-split a variable under my tooling proxy, and a `git -C add ` that expanded the glob in the *caller's* directory and silently turned the commit into a no-op. Both were captured as learnings and fed straight back into the pipeline's own docs — the self-improving loop, in action. Every change landed on `develop`, verified, with the worktree cleaned and the repo back on `develop`.
## Why it matters
The bar isn't "the model wrote some code." It's "the goal is demonstrably met, proven on a real artifact, and the system is a little smarter than it was an hour ago." /eil is my attempt to make that the default: a loop that's disciplined about verification, frugal about tokens, and relentless about never learning the same lesson twice.
### From JSON to Multi-Screen Flows: Workflow Graphs, Screens, and Multisections
- URL: https://mohammadshaker.com/en/blog/json-workflow-multi-screen-multisection
- Date: 2026-06-26T00:00:00.000Z
- Tags: SDUI, Workflows, Product Engineering, Authored with an LLM, Engineering
Multi-screen flows used to be hardcoded navigation calls compiled into the app. In a server-driven setup, the flow is just data: a backend service authors a workflow graph as JSON, validates it, and ships it on the app's init response; the client app walks that graph step by step and renders each screen, including a 'multi_section' screen that is itself a JSON map of typed sections. The payoff is that reordering screens, inserting a step, or A/B testing a long-vs-short onboarding arm becomes a JSON edit gated by a flag, with no client release.
#### Content
A multi-screen flow, such as a sign-up or onboarding sequence, is the kind of thing every app has and almost nobody enjoys changing. For years the shape of that flow lived in code: a chain of "go to the next screen" calls compiled into the binary, where changing the order or inserting a step meant a code change, a review, a build, and a store release. This post walks through a different shape, where the flow is data instead of code.
## Context
A flow is a path through screens: show roles, then avatars, then a few questions, then a notifications prompt, then exit. The traditional way to build that is to wire the screens together in the app itself, so the order and the branching are baked into the compiled program.
The approach here inverts that. A flow is a graph: a set of steps (nodes) connected by transitions (edges). That graph is authored as JSON, validated against a shared schema, and handed to the client app at runtime. The client never hardcodes the order. It receives the graph and walks it. In this article, "the backend service" means a Python/Flask API that authors and validates the graph, and "the client app" means a React Native and Expo app that walks it and renders each screen. The two examples below are illustrative, not tied to any specific product.
## The flow, visualized
Before any code, it helps to see what a flow actually is: a small number of steps, each one a screen, joined by labeled transitions that say "when this happens, go there." Here is a four-step onboarding flow as a mockup, where each box is one screen the user sees in turn.
```
[ Roles ] [ Avatars ] [ Questions ] [ Notifications ]
+---------+ +---------+ +-----------+ +--------------+
| pick | -> | choose | -> | answer a | -> | allow push | -> exit
| a role | | avatar | | few Qs | | reminders |
+---------+ +---------+ +-----------+ +--------------+
complete complete complete complete / skip
```
The same flow, drawn as a graph, makes the transitions explicit. Each edge is labeled with the event that triggers it, and the flow ends at a terminal node.
```mermaid
flowchart LR
roles["roles"]:::node
avatars["avatars"]:::node
questions["questions"]:::node
notifications["notifications"]:::node
done["END"]:::service
roles -->|"complete / skip"| avatars
avatars -->|"complete"| questions
questions -->|"complete"| notifications
notifications -->|"complete / skip"| done
```
A flow, then, is just a graph of steps plus the transitions between them. Everything else in this post is about authoring that graph as data and walking it on the client.
## Before and after
The single biggest change here is where the flow's shape lives. In the old model it lives in the app's compiled code; in the new model it lives in a JSON document the app downloads. That one move is what turns a release-gated change into a data edit.
```
BEFORE: navigation lives in the app
app code: push(Roles) -> push(Avatars) -> push(Questions)
reorder or insert a screen
|
edit app code -> review -> build -> store release
|
days to weeks before users see it
AFTER: the flow is JSON the app downloads
workflow JSON: { id, startStepId, steps[] }
reorder or insert a step
|
edit JSON -> validate in CI -> ship config
|
live on next init, no app release
```
In the after model the app binary stays the same. The client app asks for the flow on startup, receives `{ id, startStepId, steps[] }`, and walks it. Changing the order or inserting a step is an edit to that JSON, checked by schema validation in CI, and it reaches users on their next init request instead of through the app store.
## Scenario: insert an extra question for one cohort
Concretely, suppose product wants to add one extra question to onboarding, but only for new users in a specific cohort, to measure whether it hurts completion. The two models handle this very differently.
In the before model, the question screen has to be wired into the navigation code, the change ships to every platform at once, and cohort targeting means yet more conditional code. Realistically that is a one-to-two-week round trip through review, build, and release, with no easy way to limit who sees it.
In the after model, you add one step to the workflow JSON and put a gate flag on it. The client runner walks the graph and skips the gated step when the flag resolves false. Targeting is just the flag's audience rules, the change is instant on next init, and turning the experiment off is flipping the flag.
```
BEFORE AFTER
------ -----
edit nav code add 1 step to workflow JSON
ship to all platforms put gate.flag on it
~1-2 weeks runner walks + skips by flag
targeting = more code A/B + targeting via flag, instant
```
## A workflow is a graph of steps
Now to the data itself. A workflow is plain JSON with a top-level shape: an `id`, a `startStepId` naming where the flow begins, and a flat list of `steps`. Each step is a node in a directed graph, and its `transitions` map names the edges out of it.
A normal step looks like this:
```json
{
"id": "roles",
"screen": "roles_list",
"transitions": { "complete": "avatars", "skip": "avatars" },
"type": "screen"
}
```
The `screen` field is a logical name the client maps to a widget; it is not a route. The `transitions` map is the entire control-flow story: it pairs an event the screen can emit with the id of the next step to go to. When the `roles_list` screen emits `complete` (or `skip`), the runner advances to the `avatars` step. A reserved terminal marker (`"END"`) as a transition target ends the flow instead of naming another step. Because the graph is flat and the edges are named, the whole flow is inspectable and diffable as data.
```mermaid
flowchart LR
roles["roles
screen: roles_list"]:::node
avatars["avatars"]:::node
notifications["notifications
screen: multi_section"]:::node
done["END"]:::service
roles -->|"complete / skip"| avatars
avatars -->|"complete"| notifications
notifications -->|"complete / skip"| done
```
## Walking the graph on the client
The client side is deliberately thin. A small runtime, the runner, holds the current step id and exposes one main operation: given an event, look up `transitions[event]` and either advance to the named step or terminate the flow when the target is the terminal marker.
```ts
function advance(step: Step, event: string): string | null {
const next = step.transitions[event];
if (next === undefined) throw new Error(`no transition for ${event}`);
return next === "END" ? null : next; // null means the flow is done
}
```
Rendering is just as thin. A host component resolves each step's widget through a registry, mapping the step's `screen` name to a React component, and keys the rendered widget by the current step id so each step gets fresh state. The widget's `onEvent` callback feeds straight back into the runner's advance call. Malformed graphs never reach this point: the parser rejects unknown ids, dangling transitions, or unsupported step types at parse time, so a bad edit fails fast rather than rendering a dead end.
```mermaid
flowchart TD
widget["Widget keyed by stepId
emit(event)"]:::client
runner["runner.advance(event)"]:::client
lookup["transitions[event]"]:::config
next["next step"]:::node
term["END"]:::service
widget --> runner --> lookup
lookup -->|"stepId"| next
lookup -->|"END"| term
```
## A screen that is itself JSON: multi_section
Some steps are not one screen but a small stack of typed sub-screens shown together: a permission prompt, a short form, a confirmation. Rather than minting a new `screen` name for every combination, a `multi_section` step embeds a map of typed sections in its `config`. Each entry declares its own `type` and content, so the screen's composition is data, not code:
```json
{
"id": "notifications",
"screen": "multi_section",
"transitions": { "complete": "END", "skip": "END" },
"config": {
"sections": {
"permission": {
"type": "notifications_permission",
"title": { "text": "Stay in the loop" },
"subtitle": { "text": "Allow notifications for daily lessons + reminders." }
}
}
},
"type": "screen"
}
```
The matching widget reads `props.sections` (a map) plus an optional `order`, renders each sub-section through a section renderer, and owns the combined form state and validation. It only emits its event when every section is valid:
```tsx
function MultiSectionWidget({ props, onEvent }: WidgetProps) {
const { sections, order = Object.keys(sections) } = props;
const errors = validateValues(sections, values, order);
const isValid = Object.keys(errors).length === 0;
const ctx = { values, setValue, errors, isValid,
emit: (event) => { if (isValid) onEvent(event, { ...values }); } };
return order.map((key) =>
sections[key]
?
: null);
}
```
So there are two levels of JSON-as-flow. The workflow graph decides which screen comes next; the `multi_section` screen decides, again from JSON, which sub-sections compose it and in what order. Adding a section is a new key in the `sections` map, not a new screen and not a client release, as long as its `type` is already a registered section renderer.
```mermaid
flowchart TD
step["multi_section step
config.sections (map)"]:::config
widget["MultiSectionWidget
owns values + validation"]:::client
order["order ?? keys(sections)"]:::client
s1["SectionRenderer
permission"]:::node
s2["SectionRenderer
(next section)"]:::node
emit["emit complete
only if isValid"]:::service
step --> widget --> order
order --> s1
order --> s2
s1 --> emit
s2 --> emit
```
## A/B testing a whole flow with one flag
Because transitions and gating are data, you can A/B test the shape of the flow itself: a long onboarding arm versus a short one. The extra steps in the long arm carry a gate flag in their `config`, for example `{ "gate": { "flag": "long_onboarding_v1" } }`. On the client, the runner consults an injected gate resolver, a function that turns a flag name into a boolean. When a step's flag resolves false, the runner auto-skips it and follows the next transition instead. Flip the flag server-side and the same client binary walks the long arm or the short arm, with no release and no resubmission. The experiment lives entirely in flag state plus the workflow JSON.
```mermaid
flowchart LR
prev["previous step"]:::node
gate["extra step
gate.flag: long_onboarding_v1"]:::security
after["downstream step"]:::node
done["END"]:::service
prev -->|"flag true: run step"| gate
gate -->|"complete"| after
prev -->|"flag false: auto-skip"| after
after -->|"complete"| done
```
## Design principles at work
This shape is not just convenient; it lines up with a handful of long-standing design principles, each doing real work here.
- **Single Responsibility (SRP).** Each step widget owns exactly one screen's concern, and the runner's only job is to walk the graph. Neither knows about the other's internals.
- **Don't Repeat Yourself (DRY).** One runner and one widget registry replace bespoke navigation code written per flow, and a single workflow schema describes every graph.
- **Inversion of Control / Dependency Injection (IoC/DI).** The runner hardcodes no steps. It is driven entirely by the injected workflow JSON and an injected gate resolver, so behavior is configured from the outside.
- **PubSub / event-driven.** Steps and sections `emit` events rather than calling navigation directly. The runner subscribes to those events and advances the graph in response.
- **Model-View-Controller (MVC).** The workflow JSON and section maps are the model, the widgets and sections are the view, and the runner plus host act as the controller that connects them.
## What product and engineering get
The dividing line is simple: anything expressible as edges, ordering, gates, or copy is a data change validated in CI; the only thing that costs a release is a brand-new interaction.
- **Reorder and insert without shipping.** Editing transitions or adding a step that reuses a registered screen or section type changes the flow with a JSON edit alone.
- **Experiment safely.** A/B arms are flag-gated steps; the runner skips gated-off steps automatically, so one flag flip changes the flow for a cohort with no binary change.
- **Fail fast, not in production.** Schema validation on the backend plus parse-time validation on the client mean malformed flows are rejected at load, never rendered as a broken or dead-end experience.
- **One runtime, many flows.** Every flow passes through the same runner and registry, so there is a single, tested code path to maintain instead of one navigation stack per flow.
### Server-Driven UI in Production: One JSON Contract, Frontend and Backend, Zero-Release Screen Changes
- URL: https://mohammadshaker.com/en/blog/server-driven-ui-one-json-contract
- Date: 2026-06-26T00:00:00.000Z
- Tags: SDUI, Architecture, Product Engineering, Authored with an LLM, Engineering
A practical guide to Server-Driven UI across a Python/Flask backend and a React Native + Expo client. The backend emits typed JSON sections, every block carries a type string, and the client dispatches that type through a registry to a component. The result: most screen changes ship as data, not as an app-store release, and A/B tests become server-side flag flips.
#### Content
A native app release moves in days: build, submit, wait for review, hope nothing regresses. Product iteration wants to move in hours. Server-Driven UI (SDUI) closes that gap by moving the description of a screen out of the compiled app and onto the wire. The backend service decides what a screen contains and sends it as JSON; the client app reads that JSON and renders it. This piece walks through a concrete shape of SDUI: a Python/Flask backend service that emits typed JSON sections, and a React Native + Expo client app that maps each section's `type` to a component through a registry.
## Context
Server-Driven UI means the server, not the shipped app binary, decides the structure of a screen. Instead of hardcoding "the home screen has a header, then a grid, then a button" inside the app, the backend sends a list that says exactly that, as data. The client app's only job is to turn each piece of that data into a rendered component.
The core problem this solves is a cadence mismatch. A mobile app is gated by app-store review: a one-line copy tweak or a reordered screen can take one to two weeks to reach users, and iOS, Android, and web can drift out of sync while you wait. Product teams want to test ideas, reorder content, and roll changes back the same day. SDUI resolves the tension by drawing a line: layout and content live on the server (change them anytime), while the catalog of *kinds* of blocks the app can render lives in the client (change them on the normal release cadence). As long as a change reuses block kinds the app already understands, it ships as data.
```mermaid
flowchart LR
prod["product
change"]:::config
be["backend service
(Flask)"]:::service
json["typed JSON
screen"]:::config
app["client app
(RN + Expo)"]:::client
prod -->|"edit data"| be
be -->|"sections[]"| json
json -->|"HTTP"| app
app -->|"render"| app
```
## The screen, visualized
Before the contract, it helps to see what "a screen is data" actually means. Take a learning app's home screen: a progress header at the top, a grid of subject tiles, an upgrade call-to-action, and a streak row. To a user it looks like one designed page. To SDUI it is an ordered list of typed blocks, nothing more.
```
+----------------------------------------+
| Welcome back, Sara progress 62% | <- progress_header
| [#################-----------] |
+----------------------------------------+
| Your subjects |
| +----------+ +----------+ +--------+ | <- subjects_list
| | Math | | Science | | Arabic | |
| | 72% | | 40% | | 10% | |
| +----------+ +----------+ +--------+ |
+----------------------------------------+
| Unlock everything with Pro [ Go ] | <- cta / paywall
+----------------------------------------+
| Streak: 5 days o o o o o . . | <- streak
+----------------------------------------+
```
Each labeled block on the right is one typed section. The screen is just the order of those types plus the data each one carries. Reorder them, drop one, or insert another, and you have a different screen, without changing a single rendering rule.
```mermaid
flowchart TD
screen["home screen"]:::client
s1["progress_header"]:::config
s2["subjects_list"]:::config
s3["cta"]:::config
s4["streak"]:::config
screen --> s1
s1 --> s2
s2 --> s3
s3 --> s4
```
## Before and after
The payoff is easiest to see by contrasting how a screen change travels through each model. In the hardcoded world, the screen lives inside the app, so every change is a code change that must pass through app-store review before any user sees it, and each platform is edited and shipped separately. In the SDUI world, the screen is JSON the server controls, so a change is an edit to that JSON (or a flag flip) and it is live in minutes through one shared contract.
```
BEFORE: screen hardcoded in app AFTER: screen is JSON {order, sections}
+-----------------------------------+ +-----------------------------------+
| product change | | product change |
| | | | | |
| v | | v |
| edit app code (iOS, Android, web)| | edit JSON / flip a flag |
| | | | | |
| v | | v |
| app-store review (1-2 weeks) | | live in minutes |
| | | | | |
| v | | v |
| platforms drift out of sync | | one contract, all platforms |
+-----------------------------------+ +-----------------------------------+
```
```mermaid
flowchart LR
subgraph before["before"]
direction TB
pc1["product change"]:::config
code["edit app code"]:::client
rev["store review
1-2 weeks"]:::external
drift["platforms drift"]:::client
pc1 --> code --> rev --> drift
end
subgraph after["after"]
direction TB
pc2["product change"]:::config
edit["edit JSON / flag"]:::config
be["backend serves"]:::service
live["live in minutes"]:::client
pc2 --> edit --> be --> live
end
```
## Scenario: add a promo banner above the subjects
Concretely, suppose product wants a promo banner to appear above the subjects grid on the home screen. The two models handle this very differently.
Before SDUI, a promo banner is a new native widget. Someone builds it on iOS, again on Android, and again on web, wires it into the home screen layout, runs it through QA on each platform, and waits for store review. Call it one to two weeks. A/B testing it means shipping both variants in the binary and gating them client-side, which is awkward and slow to change.
With SDUI, the promo banner is just another section type. Assuming `promo_banner` is already registered in the client app (it was shipped in some earlier release), adding it is a server-side edit: insert a `promo_banner` entry into the screen's `order`, give it the data it needs, and it is live instantly. Running it as an A/B test is a feature flag that decides whether the section is included for a given cohort, no deploy on either side.
```
BEFORE AFTER
build native widget x3 platforms add "promo_banner" to JSON order
wire into home layout type already registered in app
QA each platform live instantly
store review ~1-2 weeks A/B via a server flag
```
## The contract: one type, two sides
SDUI only works if both sides agree on the wire format and keep agreeing as the product evolves. The unit of that agreement is the **section envelope**: a small JSON object with a `type` discriminator and whatever extra fields that type needs. The backend service declares the `type`; the client app keys off it. That single string is the entire integration surface between the two systems.
At runtime the flow is a plain HTTP request for a screen. The client app asks the backend for a screen by name; the backend service builds an ordered list of typed sections and returns it; the client walks the list and dispatches each section's `type` to a component.
A real-ish section envelope looks like this:
```json
{
"type": "generic_content_section",
"display_type": "hero",
"content_level": "subject",
"action": { "kind": "open_subject_tree" },
"items": [
{ "id": "1", "title": "One", "cover_url": "a.png" },
{ "id": "2", "title": "Two", "cover_url": null }
]
}
```
Note what the envelope carries beyond `type`: a display hint (`display_type`), data (`items`), and an `action` describing what happens on interaction. The client does not infer any of this; it reads it straight off the JSON.
```mermaid
flowchart LR
app["client app"]:::client
route["screen route"]:::networking
engine["section builder"]:::service
app -->|"GET screen"| route
route --> engine
engine -->|"typed JSON sections"| route
route -->|"sections[]"| app
app -->|"dispatch by type"| app
```
## The client: a registry, not a switch
On the client app, rendering is a registry lookup, not a growing `switch` statement. A registry maps each `type` string to a component, and a single renderer resolves the component or skips gracefully.
```tsx
const sectionRegistry: Record = {
progress_header: ProgressHeaderSection,
subjects_list: SubjectsListSection,
paywall: PaywallSection,
cta: CtaSection,
// ...one entry per section type
};
function SectionRenderer({ sectionKey, section, ctx }: SectionProps) {
const Component = sectionRegistry[section.type];
if (!Component) {
if (__DEV__) console.warn(`[SDUI] Unknown section type "${section.type}" - skipping.`);
return null; // graceful-skip, never crash
}
return ;
}
```
The key property is **graceful-skip**. If the backend emits a `type` the installed app bundle has never heard of, the renderer logs one dev-only warning and renders nothing. That means JSON and bundle can be out of sync, an older app meeting a newer backend, without crashing a screen. The backend can even pre-deploy data for a new type, and it lights up the moment a client release that registers that type lands. Components read props directly off the JSON, so there is no separate deserialization step: the envelope *is* the props.
```mermaid
flowchart TD
sec["section JSON
{type, ...}"]:::config
reg["registry
lookup"]:::client
comp["matched
component"]:::client
skip["graceful-skip
(render null)"]:::client
sec -->|"section.type"| reg
reg -->|"known type"| comp
reg -->|"unknown type"| skip
```
## A/B testing by changing JSON
Once screens are data, experimentation stops being a code branch and becomes a content decision. Sections and whole screens are gated by server-side feature flags. The backend resolves each flag from a database-backed store with a static defaults fallback, so an unset or unknown flag still has a defined value. A section carries a gate flag; the section builder evaluates that gate before deciding whether to emit the section, or which variant of it to emit.
To run an A/B test, point a screen at two candidate sections behind the same flag and flip the flag server-side for a cohort. No deploy on either system. Want to test a new paywall against the current one, or a reordered home screen? Change the JSON the builder emits and flip the flag. The client app renders whatever arrives.
```mermaid
flowchart TD
flag["gate flag
(DB + defaults)"]:::storage
gate["gate check"]:::security
engine["section builder"]:::service
a["variant A"]:::config
b["variant B"]:::config
app["client app
renders variant"]:::client
flag --> gate
gate -->|"flag on"| a
gate -->|"flag off"| b
a --> engine
b --> engine
engine -->|"JSON section"| app
```
## Design principles at work
This architecture is a clean instance of several familiar principles, each tied to a concrete mechanism:
- **Single Responsibility (SRP):** each section component renders exactly one block type and nothing else. `SubjectsListSection` knows how to draw a subjects grid; it knows nothing about paywalls or streaks.
- **Don't Repeat Yourself (DRY):** there is one registry and one render path instead of N bespoke screen-builders, and one shared contract instead of two divergent definitions of "a screen."
- **Inversion of Control / Dependency Injection (IoC/DI):** the app never calls section components directly. It hands control to the registry, which selects the component. The data source (which screen JSON to fetch) and the gate resolver (how flags evaluate) are injected dependencies, so they can be swapped in tests or per environment.
- **Publish/Subscribe (event-driven):** section components do not navigate or mutate global state themselves. They `emit` events (for example, a CTA emits a `complete` event), and a host runner subscribes to those events and routes them. Sections stay decoupled from app navigation.
- **Model-View-Controller (MVC):** the JSON screen is the model, the section components are the view, and the registry plus host runner are the controller that maps model to view and dispatches actions.
## What product and engineering get
- **Velocity:** most screen changes ship as data, not as an app-store submission and review wait.
- **No release for registered types:** reordering, reparameterizing, or feature-flagging sections of already-registered types never needs a client release.
- **Safe rollback:** A/B tests and rollouts are flag flips on existing types, reversible instantly with no rollback build.
- **Fewer drift bugs:** one shared contract, agreed by both systems, removes the "frontend and backend disagree" failure mode, and unknown types degrade quietly instead of crashing.
### Meta's Ranking Engineer Agent: Autonomous Experimentation Under Hard Guardrails
- URL: https://mohammadshaker.com/en/blog/meta-ranking-engineer-agent
- Date: 2026-06-13T00:00:00.000Z
- Tags: AI, agents, ranking, experimentation, architecture, Meta, Authored with an LLM, Engineering
Meta's internal Ranking Engineer Agent runs the full experimentation loop on its own: it reads the metric regression, generates ranking hypotheses, codes the change, ships a guarded A/B test, reads the readout, and iterates. The interesting part isn't the autonomy, it's the guardrails that make autonomy safe on a system serving billions of people.
#### Content
A ranking engineer at Meta spends most of their day inside a loop. Look at a metric that moved the wrong way. Form a hypothesis about why. Write a feature or tweak a model. Ship it behind an A/B experiment. Wait. Read the readout. Decide. Repeat.
That loop is mostly mechanical. The judgement lives at the edges, in the hypothesis and the decision. Everything between is plumbing: pulling data, writing boilerplate feature code, launching the experiment with the right safety config, parsing the readout dashboard.
Meta's internal **Ranking Engineer Agent** automates the plumbing and takes a first pass at the judgement. It runs the whole loop on its own, generating hypotheses and testing them, but it does so inside a fence of hard guardrails. The autonomy is the headline. The guardrails are the actual engineering.
## What "Autonomous Experimentation" Actually Means Here
The agent is not a chatbot that suggests ideas. It is a closed-loop system with a goal (improve a target metric without regressing guarded metrics) and the authority to act on a production ranking system, within limits.
A single autonomous cycle looks like this:
1. **Observe.** Read the current state: target metric, guarded metrics, recent experiments, feature catalog.
2. **Hypothesize.** Generate candidate explanations and interventions ("watch-time dipped for short-form because the freshness feature decays too fast; try a slower decay").
3. **Rank hypotheses.** Score candidates by expected impact, cost, and risk. Pick the most promising one that passes the guardrails.
4. **Implement.** Write the feature/model change as code against the ranking stack.
5. **Guard-check.** Validate the change against static guardrails before anything ships (blast radius, forbidden features, metric guards).
6. **Experiment.** Launch a small, capped A/B test with automatic kill-switches.
7. **Read out.** Parse the experiment results, including statistical significance and guarded-metric movement.
8. **Decide.** Ship, iterate, or discard. Then loop back to step 1.
The key shift from a human workflow: steps 1, 4, 6, 7 are fully mechanized, and steps 2, 3, 8 are done by the agent with a human approval gate before anything graduates from experiment to production.
## Architecture
The system is best read as a control loop wrapped in a policy layer. The inner loop does the experimentation. The outer policy layer (the guardrails) decides what the inner loop is allowed to do at every step.
```mermaid
flowchart TD
GP["Guardrail Policy Layer
evaluated before every action
blast-radius cap, metric guards, forbidden-feature list,
budget/quota, rate limits, human-approval gate"]:::security
OB["OBSERVE
metrics, feature catalog, past exps"]:::client
HY["HYPOTHESIZE
LLM core + ranker
score by impact/risk/cost"]:::service
IM["IMPLEMENT
codegen + tests + lint"]:::service
EX["EXPERIMENT
capped A/B + kill-switch + auto-rollback"]:::node
RO["READ OUT
sig test, guarded metric delta"]:::node
DE["DECIDE
ship? iterate? discard?"]:::node
GR["GRADUATE to prod"]:::storage
OB --> HY --> IM --> EX --> RO --> DE
DE -->|"ship (needs human gate)"| GR
DE -->|"discard / iterate"| OB
GR -->|loop back| OB
GP -. check .-> OB
GP -. check .-> HY
GP -. check .-> IM
GP -. check .-> EX
```
### The inner loop components
Each stage of that inner loop maps to a concrete piece of the system. Here is what each one does and why it matters.
- **Observation interface.** Read-only adapters over the metrics store, the experiment registry, and the feature catalog. This is what grounds the agent in reality instead of hallucination: every hypothesis must reference real features and real metric movements.
- **Hypothesis engine.** An LLM core that proposes interventions, plus a separate ranker that scores each candidate on expected impact, implementation cost, and risk. The ranker matters more than the generator: a generator that produces a hundred ideas is useless without a disciplined way to pick the two worth testing.
- **Implementation layer.** Codegen against the ranking stack, with mandatory unit tests and lint. The agent does not get to ship code that doesn't compile or that skips its own tests.
- **Experiment harness.** Launches A/B tests through the same internal experimentation platform humans use, but with safety config injected: small traffic allocation, kill-switches wired to guarded metrics, and a fixed maximum duration.
- **Readout parser.** Turns the experiment dashboard into a structured verdict: did the target metric move, was it significant, did any guarded metric regress.
## The Guardrails Are the Product
On a system that ranks content for billions of people, an unconstrained autonomous agent is a liability, not an asset. The guardrails are what make the autonomy shippable. They fall into a few categories.
**Blast-radius caps.** Any experiment the agent launches is hard-limited to a tiny fraction of traffic. The agent cannot allocate more, ever. The cap is enforced by the platform, not by the agent's own judgement, so a reasoning error can't widen the blast radius.
**Metric guards.** A set of protected metrics (integrity, time-well-spent, ad load, latency) that may never regress beyond a threshold. If a guarded metric crosses the line mid-experiment, the kill-switch fires and the experiment auto-rolls-back without waiting for the agent to notice.
**Forbidden-feature lists.** Certain features and signals are off-limits to the agent for legal, privacy, or integrity reasons. The implementation layer refuses to generate code that touches them.
**Budget and rate limits.** A cap on how many experiments the agent can run per unit time and how much compute it can spend. This bounds both cost and the rate at which mistakes can compound.
**Human approval gate.** This is the most important one. The agent can run experiments and read them out autonomously, but **graduating a change from experiment to production requires a human sign-off.** The agent does the loop; a human owns the irreversible step.
The design principle underneath all of these: **enforce guardrails outside the agent's reasoning.** A guardrail that the agent can talk itself past is not a guardrail. Blast-radius caps, kill-switches, and forbidden-feature checks all live in the policy layer and the platform, where the agent's outputs are inputs to be validated, not commands to be obeyed.
```
AGENT PROPOSES POLICY LAYER DECIDES RESULT
-------------- -------------------- ------
"ship to 5%" --> cap = 1% (hard) --> clamped to 1%
"use feature X" --> X on forbidden list --> rejected, fed back
"run 50 exps" --> budget = 6 / day --> queued / throttled
guarded metric --> threshold breached --> auto-rollback, no ask
"graduate this" --> needs human approval --> paused for review
```
## Why This Shape Wins
The agent does not try to be smarter than the experimentation system. It tries to be faster at the loop while the system stays the source of truth for safety. Three properties make it work:
1. **Grounding over generation.** Every hypothesis is anchored to observed metrics and the real feature catalog, so the LLM core proposes inside a constrained space rather than free-associating.
2. **External enforcement.** Guardrails live in the platform, not in the prompt. The agent cannot reason its way around a hard cap.
3. **Reversibility by default.** Everything the agent does on its own is small, capped, and auto-rollback-able. The one irreversible action, graduating to production, is gated on a human.
That is the general recipe for autonomous agents in high-stakes systems: let the agent own the fast, reversible inner loop, and put hard, externally-enforced fences around every action that could cause real harm. The autonomy gets the headlines. The guardrails are why it ships.
### Building a Palantir-Style Ontology without Palantir
- URL: https://mohammadshaker.com/en/blog/building-a-palantir-style-ontology-without-palantir
- Date: 2026-04-18T00:00:00.000Z
- Tags: architecture, ontology, graph, engineering, AI, agents, Authored with an LLM, Engineering
Palantir Foundry sells a six-figure ontology to Fortune 500 companies. We built the same idea -- a queryable graph of every business noun, relationship, and verb -- for an Arabic kids' ed-tech app, using JSONL files, Python extractors, and a Claude agent. Total cost: ~$16k/year. Here is the full architecture, the reasoning, and the kill gates.
#### Content
Palantir charges six to seven figures a year for Foundry. Their core innovation is not the software -- it is the idea: **put every business noun, every relationship between them, and every verb that acts on them into one queryable graph.** Then let humans and agents ask questions in natural language instead of writing SQL.
We needed exactly this. But we are a 10-person Arabic ed-tech company, not a Fortune 500 oil company. So we built it ourselves.
This post documents the full reasoning, the architecture, what we rejected and why, and the falsifiable kill gates that will tell us if this was a good idea or a waste of time.
## The Problem
Knowledge about our business is scattered across 5+ systems:
```
"Which users are stuck?"
└── lives in MySQL. Only eng can query.
"What does endpoint X touch?"
└── lives in code + tribal memory.
"Why is this Thurayya chapter failing?"
└── lives in Athena cold storage + Amplitude + realtime_answers.
"What do we know about user Layla?"
└── 5 different tables, 3 different systems.
"If I change this Lambda, what breaks?"
└── nobody knows. We find out in prod.
```
Every question requires an engineer to write SQL. Each answer takes 15-45 minutes. We measured: **~8 hours per week** of engineering time disappears into ad-hoc queries for support, content, and growth teams.
That is $40k/year of engineering labor spent as a human database adapter.
## What Palantir Actually Sells
Strip away the marketing and Foundry is four things:
```
┌─────────────────────────────────────────────────┐
│ Palantir Foundry │
│ │
│ 1. Graph Store ── nodes + edges + types │
│ 2. Extractor Layer ── pulls from your systems │
│ 3. Query Layer ── Cypher / search / NL │
│ 4. Action Layer ── verbs that mutate state │
│ │
│ Price: $$$,$$$/year + US region data gravity │
└─────────────────────────────────────────────────┘
```
These four components are not rocket science. They are a graph database, some ETL scripts, a query interface, and an API gateway. What makes Foundry valuable is that Palantir forces you to actually model your business in a single graph. The discipline is the product, not the software.
We can enforce that discipline ourselves.
## Why We Rejected Foundry
Four concrete reasons, not vibes:
1. **Cost.** Six to seven figures per year. Our entire engineering budget does not justify a Foundry license.
2. **Data gravity.** Foundry pulls data into US regions. Our prod DB is in me-south-1 (Bahrain) and eu-west-1 (Ireland). Cross-region egress for PII data creates a compliance problem we do not have today.
3. **Lock-in.** Foundry Ontology objects use a proprietary schema. Leaving means rebuilding everything.
4. **Overkill.** Foundry is built for organizations with thousands of tables. We have ~60 domain modules.
## The DIY Architecture
Four layers. Same conceptual structure as Foundry. ~10x cheaper.
```mermaid
flowchart LR
RDS["RDS (prod)"]:::storage -->|nightly read-rep| RE["rds_extractor"]:::service
SRC["src/models, resources
docs/plans, serverless.yml"]:::config -->|on push| ME["model / resource /
plan / infra extractors"]:::service
SNS["SNS: outcome events"]:::external -->|real-time| OL["outcome_listener"]:::service
RE --> GS["GraphStore (JSONL)"]:::storage
ME --> GS
OL --> GS
GS -->|query| AG["Claude agent + tool-calls"]:::client
AG -->|actions| EP["existing API endpoints
grant, reorder, etc."]:::node
```
### Layer 1: Graph Store
For Stage 1, we chose **JSONL** -- append-only, line-delimited JSON files in S3, loaded into an in-memory index at Lambda cold-start. No database server. No infrastructure.
Why not Neo4j or Postgres+AGE from day one? Because we are running a 2-week spike. Spinning up Neo4j for a prototype is premature optimization. JSONL handles <100k nodes fine. The real store decision happens at Phase 3, when we have actual query patterns to optimize for. (I wrote a [separate deep-dive on Neo4j vs Postgres+AGE vs JSONL](/neo4j-vs-postgres-age-vs-jsonl-graph-data-showdown) if you want the full comparison.)
### Layer 2: Extractors
One Python module per source of truth. Each extractor reads one system and emits typed `Node` and `Edge` objects:
```
src/ontology/extractors/
├── model_extractor.py ── walks SQLAlchemy models
├── resource_extractor.py ── walks Flask resources
├── plan_extractor.py ── reads docs/plans/ YAML front-matter
├── infra_extractor.py ── reads serverless.yml
└── rds_extractor.py ── queries prod read-replica
```
The critical design constraint: **extractors are read-only projections.** If the graph and the source disagree, the source wins. The graph gets rebuilt nightly from scratch. This is cheaper than bidirectional sync and eliminates an entire class of consistency bugs.
### Layer 3: Query
A Claude agent with the full ontology schema in its cached system prompt. Users type natural language. The agent translates to graph queries, runs them, and returns structured answers.
```
User: "Which Amal users in UAE finished onboarding but
never opened day 2?"
Agent: [calls graph_query tool]
[filters User nodes: app=Amal, country=UAE,
onboarding_complete=true, day_2_opened=false]
[returns 340 users in a table]
```
### Layer 4: Actions
The agent can trigger existing API endpoints. It never has its own mutations -- it calls back into the existing domain layer with full audit logging.
```
User: "Send those 340 users a 50% off winback email."
Agent: [calls execute_action tool]
[POST /boss/v1/campaigns with user_ids + template]
[logged in activity_log with source=ontology]
```
## The Schema
Declared in a single YAML file. "Terraform for ontology":
```yaml
version: 1
id_format: "{type}:{natural_key}"
object_types:
App: { natural_key: name }
User: { natural_key: id } # PII redacted at boundary
Subject: { natural_key: id }
Chapter: { natural_key: id }
ContentByte: { natural_key: id }
ContentBit: { natural_key: id }
Membership: { natural_key: id }
Plan: { natural_key: path }
Endpoint: { natural_key: path_method }
LambdaFn: { natural_key: name }
link_types:
has: { from: [App, Subject, Chapter],
to: [Subject, Chapter, ContentByte] }
contains: { from: ContentByte, to: ContentBit }
answered_by: { from: ContentBit, to: User }
pays_for: { from: User, to: Membership }
unlocks: { from: Membership, to: ContentByte }
touches: { from: Plan,
to: [Endpoint, ContentByte, LambdaFn] }
deployed_to: { from: Endpoint, to: LambdaFn }
action_types:
grant_membership: { target: User, endpoint: POST /boss/v1/memberships }
regenerate_bit: { target: ContentBit, endpoint: POST /boss/v1/content-bits/{id}/regenerate }
```
Every change to this file is a PR with CI validation. The schema is the single source of truth for what the business graph looks like.
## What This Unlocks -- Three Outcomes
Building the graph is the means, not the end. Here is the payoff, staged over time: first stop wasting engineering hours, then let content improve itself, then turn the same graph into new products. Each outcome builds on the one before it.
### Outcome 1 (Months 1-3): Stop Wasting Eng Time on SQL
Support, content, and growth teams self-serve answers through the agent UI. An 8h/week time sink becomes a 90-second interaction.
**Before:**
```
Support lead → Slack message to eng → 30 min wait →
eng writes SQL → pastes results → support lead acts →
total: 45 min per question
```
**After:**
```
Support lead → types question in ontology UI →
agent returns answer + action button → done →
total: 90 seconds
```
### Outcome 2 (Months 6-12): Content That Improves Itself
Every answered question generates a signal: user X got question Y wrong/right, took Z milliseconds. This signal flows through the graph.
A weekly batch job finds content with low correct-rates per age group. It regenerates those questions -- new phrasing, new audio, new images -- and A/B tests the new version against the old. The winner replaces the loser automatically.
```
Week 1: "Letter jim for age 5-6" -- 43% correct rate
│
▼
Regeneration: 3 new variants created overnight
│
▼
Week 2: A/B test on 1,000 kids each
│
▼
Week 3: Winner deployed. Correct rate: 61%.
Nobody noticed. Nobody intervened.
```
A competitor who copies our content today gets a static snapshot. We get the loop. That gap widens every week.
### Outcome 3 (Year 2-3): Personalization + B2B Schools
The same graph that answers support questions and improves content can also drive a per-child daily lesson picker. If Layla struggled with short vowels yesterday, today's lesson starts with a warm-up on short vowels.
Package the same graph as a teacher dashboard: per-student mastery, class-wide weak topics, exportable reports. Sell to schools at $8/student/month.
## The Cost
Here is the full three-year bill, broken down by line item. The point of the table is the bottom row: even counting engineering maintenance time, the whole system runs about ten times cheaper than a Palantir Foundry license.
| Item | Year 1 | Year 2 | Year 3 |
| --- | ---: | ---: | ---: |
| Graph store (JSONL → Neo4j) | $0 | $600 | $1,200 |
| Claude API (cached, agents) | $1,200 | $2,400 | $4,800 |
| Lambda compute (extractors) | $50 | $100 | $200 |
| Engineering maintenance (0.5 week/quarter) | $15,000 | $15,000 | $15,000 |
| **Total** | **about $16k** | **about $18k** | **about $21k** |
Palantir Foundry: six figures per year minimum. DIY: about ten times cheaper.
## The Kill Gates
Every phase has a falsifiable exit condition. Fail means kill, not extend.
```
P0 ─── pass ──▶ P1 ─── pass ──▶ P2 ─── pass ──▶ P3 ──┐
│ │ │ │
fail fail fail ▼ pass
▼ ▼ ▼ P6/P7/P8
ABORT fallback Outcome 1 only (personalization,
to wiki (no loop) planner, schools)
```
| Phase | Gate | Deadline | Kills if fails |
|---|---|---|---|
| P0 | 3 recent plan docs have YAML falsifiable_claims | 2026-04-25 | Everything |
| P1 | Agent answers 20 support Qs in <50% baseline SQL time | 2026-06-01 | Graph effort |
| P2 | OutcomeEvent delivery >=99% for 7 days | 2026-06-20 | Moat loop thesis |
| P3 | Regenerated bits show >=5pp correct_rate lift | 2026-08-31 | Outcome 2 |
**Committed eng before the first real kill gate: 1.6 eng-weeks.** That is the actual bet. Everything beyond that is earned at a gate.
## The Falsification Test
Claims I am least certain about, stated so they can be disproven:
1. **"The ontology will save 8h/week of eng time."** This assumes support/content teams will actually use the agent UI instead of defaulting to Slack. If adoption is <30% after 3 months, the savings are fictional.
2. **"Content self-improvement via A/B is a moat."** This assumes correct_rate is the right proxy for content quality, that regenerated variants are meaningfully different, and that the sample sizes are achievable. If A/B shows no lift, the loop does not close and we revert to Outcome 1 only.
3. **"$16k/year is the real cost."** This assumes 0.5 eng-wk/quarter of maintenance is sufficient. If the graph drifts from the source schema and requires constant manual fixing, the real cost is much higher.
4. **"JSONL is fine for Stage 1."** If the graph exceeds 100k nodes before Phase 3, or if query latency exceeds 2 seconds on real questions, the JSONL store becomes a bottleneck earlier than planned.
## The One Insight
Palantir's real product is not software. It is the discipline of modeling your business as a graph. The software is just the enforcement mechanism.
You can enforce that discipline yourself with a YAML schema, some Python extractors, and an LLM that reads the graph. The hard part is not the technology. The hard part is maintaining the schema as the business evolves -- which is why every phase has a kill gate that tests whether humans are actually keeping it up.
If they are not, the ontology is dead regardless of the technology underneath. That is the most falsifiable claim of all.
### Neo4j vs Postgres+AGE vs JSONL: A Graph Data Showdown
- URL: https://mohammadshaker.com/en/blog/neo4j-vs-postgres-age-vs-jsonl-graph-data-showdown
- Date: 2026-04-18T00:00:00.000Z
- Tags: databases, graph, neo4j, postgresql, architecture, engineering, Authored with an LLM, Engineering
A deep, falsifiable comparison of three approaches to graph-shaped data: Neo4j (native graph), PostgreSQL with the Apache AGE extension (relational+graph hybrid), and plain JSONL files (no database at all). We debate where each wins, where each breaks, and show concrete examples with ASCII visualizations.
#### Content
Most "comparison" articles give you a feature matrix and call it a day. This one does not. We are going to stress-test three fundamentally different approaches to storing and querying graph-shaped data, find their breaking points, and show you exactly when each one fails.
The contenders:
```
+------------------+ +------------------+ +------------------+
| Neo4j | | Postgres + AGE | | JSONL |
| Native Graph DB | | Relational + | | Flat Files, |
| Cypher queries | | Graph Extension | | No DB at all |
| Index-free adj. | | openCypher on | | Line-delimited |
| $$$$ at scale | | top of tables | | JSON objects |
+------------------+ +------------------+ +------------------+
```
## The Problem: Why Graph-Shaped Data Exists
Not all data is tabular. Consider a social network, a dependency tree, a fraud detection ring, or a knowledge graph. The relationships between entities carry as much meaning as the entities themselves.
Here is the toy graph we will use throughout this post:
```mermaid
flowchart TD
Alice["Alice"]:::client
Bob["Bob"]:::client
Carol["Carol"]:::client
Acme(["Acme"]):::service
Globex(["Globex"]):::service
Initech(["Initech"]):::service
Alice -->|KNOWS| Bob
Alice -->|KNOWS| Carol
Bob -->|KNOWS| Carol
Bob -->|WORKS_AT| Acme
Carol -->|WORKS_AT| Globex
Acme -->|COMPETES| Initech
Globex -->|COMPETES| Initech
```
Seven nodes. Eight edges. Simple enough to reason about, complex enough to expose real differences.
---
## Round 1: Data Modeling
First question: how does each approach let you define the nodes (entities) and edges (relationships) in the first place? Modeling is where you feel the gap between a purpose-built graph language and a plain flat file.
### Neo4j
Cypher makes graph creation feel natural:
```cypher
CREATE (a:Person {name: 'Alice'})
CREATE (b:Person {name: 'Bob'})
CREATE (c:Person {name: 'Carol'})
CREATE (acme:Company {name: 'Acme'})
CREATE (globex:Company {name: 'Globex'})
CREATE (initech:Company {name: 'Initech'})
CREATE (a)-[:KNOWS]->(b)
CREATE (a)-[:KNOWS]->(c)
CREATE (b)-[:KNOWS]->(c)
CREATE (b)-[:WORKS_AT]->(acme)
CREATE (c)-[:WORKS_AT]->(globex)
CREATE (acme)-[:COMPETES]->(initech)
CREATE (globex)-[:COMPETES]->(initech)
```
Nodes and edges are first-class citizens. Relationship types are part of the schema. This is genuinely elegant.
### Postgres + AGE
AGE bolts openCypher onto PostgreSQL. You create a graph namespace, then write Cypher inside SQL:
```sql
-- Enable the extension
CREATE EXTENSION age;
LOAD 'age';
SET search_path = ag_catalog, "$user", public;
SELECT create_graph('social');
SELECT * FROM cypher('social', $$
CREATE (a:Person {name: 'Alice'})
CREATE (b:Person {name: 'Bob'})
CREATE (c:Person {name: 'Carol'})
CREATE (acme:Company {name: 'Acme'})
CREATE (globex:Company {name: 'Globex'})
CREATE (initech:Company {name: 'Initech'})
CREATE (a)-[:KNOWS]->(b)
CREATE (a)-[:KNOWS]->(c)
CREATE (b)-[:KNOWS]->(c)
CREATE (b)-[:WORKS_AT]->(acme)
CREATE (c)-[:WORKS_AT]->(globex)
CREATE (acme)-[:COMPETES]->(initech)
CREATE (globex)-[:COMPETES]->(initech)
$$) as (v agtype);
```
Same Cypher, but wrapped in a SQL function call. Under the hood, AGE stores vertices and edges in regular PostgreSQL tables with `agtype` (a JSONB-like binary type) columns. This is both its strength and its weakness, as we will see.
### JSONL
No schema. No server. Just lines in a file:
```json
{"id":"alice","type":"Person","name":"Alice"}
{"id":"bob","type":"Person","name":"Bob"}
{"id":"carol","type":"Person","name":"Carol"}
{"id":"acme","type":"Company","name":"Acme"}
{"id":"globex","type":"Company","name":"Globex"}
{"id":"initech","type":"Company","name":"Initech"}
{"src":"alice","rel":"KNOWS","dst":"bob"}
{"src":"alice","rel":"KNOWS","dst":"carol"}
{"src":"bob","rel":"KNOWS","dst":"carol"}
{"src":"bob","rel":"WORKS_AT","dst":"acme"}
{"src":"carol","rel":"WORKS_AT","dst":"globex"}
{"src":"acme","rel":"COMPETES","dst":"initech"}
{"src":"globex","rel":"COMPETES","dst":"initech"}
```
13 lines. No dependencies. Readable by any language, any tool, any era. Append-only writes are trivial: `echo '{"src":"dave","rel":"KNOWS","dst":"alice"}' >> graph.jsonl`.
### Verdict: Round 1
Each approach gets a score out of 10 for how cleanly it lets you express a graph.
| Criterion | Neo4j | AGE | JSONL |
| --- | ---: | ---: | ---: |
| Modeling elegance | 9 | 7 | 6 |
Schema clarity: Neo4j > AGE >> JSONL. Setup friction: JSONL > AGE >> Neo4j.
Neo4j wins on expressiveness. JSONL wins on zero friction. AGE sits in the middle, inheriting Cypher's expressiveness but with the ceremony of SQL wrapping.
**But wait** -- JSONL's simplicity is deceptive. There is no schema enforcement. Nothing prevents you from writing `{"src":"alice","rel":"KNOWZ","dst":"bob"}` with a typo. You traded elegance for fragility. That is a real cost.
---
## Round 2: Querying -- The 2-Hop Traversal
The question that separates graph systems from pretenders: "Find all companies that compete with a company where a friend-of-Alice works."
In our graph, Alice knows Bob and Carol. Bob works at Acme, Carol works at Globex. Both Acme and Globex compete with Initech. So the answer is `{Initech}`.
```mermaid
flowchart LR
Alice["Alice"]:::client -->|KNOWS| Bob["Bob"]:::client
Alice -->|KNOWS| Carol["Carol"]:::client
Bob -->|WORKS_AT| Acme(["Acme"]):::service
Carol -->|WORKS_AT| Globex(["Globex"]):::service
Acme -->|COMPETES| Initech(["Initech = ANSWER"]):::danger
Globex -->|COMPETES| Initech
```
### Neo4j
```cypher
MATCH (a:Person {name: 'Alice'})-[:KNOWS]->(friend)
-[:WORKS_AT]->(company)-[:COMPETES]->(competitor)
RETURN DISTINCT competitor.name
```
One query. Reads like a sentence. Under the hood, Neo4j uses **index-free adjacency**: each node physically stores pointers to its neighbors. Traversal is O(1) per hop, not O(log n) like an index lookup. For a 3-hop traversal on a million-node graph, this difference is enormous.
### Postgres + AGE
```sql
SELECT * FROM cypher('social', $$
MATCH (a:Person {name: 'Alice'})-[:KNOWS]->(friend)
-[:WORKS_AT]->(company)-[:COMPETES]->(competitor)
RETURN DISTINCT competitor.name
$$) as (competitor agtype);
```
Identical Cypher. But here is the catch that most articles skip: **AGE's Cypher executor does not push property predicates down to PostgreSQL indexes.** The `{name: 'Alice'}` filter on the Person vertex? AGE performs a sequential scan on the vertex table and then filters in its own executor layer. This is documented in AGE issue #2348 (March 2026). For small graphs, you will not notice. At 10M+ vertices, you are in trouble.
Furthermore, AGE's variable-length edge (VLE) expansion is O(n^k) with **no cycle detection** (issue #2349). The `shortestPath()` function internally uses VLE instead of BFS (issue #2350). For deep traversals on cyclic graphs, AGE can spiral into exponential blowup where Neo4j handles it with bounded BFS.
### JSONL
You write the traversal yourself:
```python
import json
from collections import defaultdict
nodes = {}
edges = defaultdict(list)
with open('graph.jsonl') as f:
for line in f:
obj = json.loads(line)
if 'src' in obj:
edges[obj['src']].append(obj)
else:
nodes[obj['id']] = obj
# 3-hop traversal: Alice -> KNOWS -> WORKS_AT -> COMPETES
alice_friends = [e['dst'] for e in edges['alice']
if e['rel'] == 'KNOWS']
friend_companies = [e['dst'] for e in edges.get(f, [])
for f in alice_friends
if e['rel'] == 'WORKS_AT']
competitors = {e['dst'] for c in friend_companies
for e in edges.get(c, [])
if e['rel'] == 'COMPETES'}
print(competitors) # {'initech'}
```
It works. But you just wrote a graph traversal engine from scratch. Every new query type requires new code. There is no query planner, no optimizer, no index. You are doing full scans through adjacency lists you built in memory.
### Verdict: Round 2
Scored out of 10 on raw multi-hop query power, meaning how well each handles queries that walk several relationships in a row.
| Criterion | Neo4j | AGE | JSONL |
| --- | ---: | ---: | ---: |
| Multi-hop query power | 10 | 6 | 3 |
AGE loses points because Cypher property filters have no index pushdown, variable-length expansion is `O(n^k)` without cycle detection, and `shortestPath()` is not breadth-first-search optimized. JSONL loses more because every query is hand-coded, there is no query optimizer, and reads require the full dataset in memory or repeated scans.
Neo4j dominates here. This is its entire reason for existing. AGE gives you the syntax but not the engine. JSONL gives you nothing but raw materials.
**Falsification check**: could you beat Neo4j's multi-hop performance with JSONL? Yes -- if the graph fits in RAM, a hand-tuned adjacency list in C++ or Rust with cache-aligned memory layout can outperform Neo4j's JVM-based engine. But you are now building a database, not using a file format.
---
## Round 3: Mixed Workloads (Graph + Relational)
Here is where the debate gets interesting. Real applications rarely have pure graph queries. You also need: aggregate revenue by company, filter persons by signup date, join with an orders table, run window functions.
### Neo4j
Neo4j has no `GROUP BY` with window functions. No foreign keys to an orders table. No `JOIN` with a relational dataset. If your application needs both graph traversals and analytical SQL, you end up running two databases and syncing between them.
```mermaid
flowchart TD
App["Your app
two connections, two schemas,
two failure modes,
eventual consistency between them"]:::client
App --> N["Neo4j (graph)"]:::service
App --> P["Postgres (OLAP)"]:::storage
N <-->|sync| P
```
This is the hidden cost of Neo4j that vendors do not emphasize. You pay for the graph engine, then you pay again for the operational complexity of a dual-database architecture.
### Postgres + AGE
This is AGE's killer feature. One database. Graph queries and SQL queries coexist:
```sql
-- Graph traversal
SELECT * FROM cypher('social', $$
MATCH (p:Person)-[:WORKS_AT]->(c:Company)
RETURN p.name, c.name
$$) as (person agtype, company agtype);
-- In the same database, same transaction:
SELECT company_name, SUM(revenue)
FROM quarterly_earnings
WHERE quarter = '2026-Q1'
GROUP BY company_name
ORDER BY SUM(revenue) DESC;
```
You can even join graph results with relational tables:
```sql
SELECT g.person, g.company, q.revenue
FROM cypher('social', $$
MATCH (p:Person)-[:WORKS_AT]->(c:Company)
RETURN p.name, c.name
$$) as g(person agtype, company agtype)
JOIN quarterly_earnings q
ON q.company_name = g.company::text
WHERE q.quarter = '2026-Q1';
```
One transaction. One connection. ACID across both graph and relational data.
### JSONL
You load both the JSONL graph and your CSV/Parquet tables into pandas or DuckDB and join manually. Doable for analytics pipelines. Not doable for a production application serving real-time requests.
### Verdict: Round 3
Scored out of 10 on how well each handles graph queries and ordinary relational work (aggregates, joins, filters) together.
| Criterion | Neo4j | AGE | JSONL |
| --- | ---: | ---: | ---: |
| Mixed graph and relational workload | 3 | 9 | 4 |
AGE wins decisively here. The "one database" story is real and valuable.
**Falsification check**: Neo4j added GDS (Graph Data Science) library for analytics, and version 5.x improved aggregation capabilities. But it still cannot replace a relational database for structured analytical queries. The gap is narrowing but real.
---
## Round 4: Operational Cost
Let us stop being coy about money.
### Neo4j
- **Community Edition**: free, but limited to a single machine (no clustering, no online backup, no role-based access control).
- **Enterprise Edition**: contact sales. Typically $100K+/year for production deployments. This is not a rounding error.
- **AuraDB (managed)**: starts around $65/month for a toy instance. Production instances with adequate memory run $500-2000+/month.
- **Hidden cost**: JVM tuning. Neo4j runs on the JVM. Heap sizing, GC pauses, off-heap page cache configuration. You need someone who understands JVM operations.
```
Neo4j Cost Curve
$
| /
| /
| /
| /
| /
| . / <-- Enterprise license kicks in
| . .
| .
| .
+-----------------------------------> Data Size
Free tier "Call Sales"
```
### Postgres + AGE
- **PostgreSQL**: free, open source, battle-tested.
- **AGE extension**: free, Apache 2.0 licensed. No licensing games.
- **Managed options**: any managed Postgres provider (RDS, Cloud SQL, Supabase) -- though AGE extension support varies. You may need to self-host.
- **Hidden cost**: AGE is still maturing. You will hit bugs. The community is smaller. You become your own support tier.
```
Postgres+AGE Cost Curve
$
|
| ---------------------------- (flat: infra cost only)
|
|
+-----------------------------------> Data Size
Always open source, pay for compute only
```
### JSONL
- **Cost**: zero. It is files.
- **Hidden cost**: engineering time. Every "feature" (indexing, querying, concurrency, durability) you build yourself. At small scale this is free. At medium scale it is cheaper than Neo4j. At large scale it is more expensive than both because you are paying senior engineers to maintain a bespoke system.
```
JSONL Total Cost of Ownership
$
| /
| /
| / <-- Engineering time
| / dominates
| /
| /
| /
| /
| /
| /
| . . . . . . <-- Near-zero at small scale
+-----------------------------------> Complexity
```
### Verdict: Round 4
Scored out of 10 on total cost of ownership (licensing plus the engineering and ops time you pay over the life of the system); higher is cheaper.
| Criterion | Neo4j | AGE | JSONL |
| --- | ---: | ---: | ---: |
| Total cost of ownership | 4 | 8 | 7* |
\* JSONL's score drops to 2 above about one million edges.
---
## Round 5: Write Performance and Ingestion
Bulk loading 1 million edges. A common real-world requirement.
### Neo4j
Neo4j ships `neo4j-admin import` for offline bulk loading. It is fast -- millions of nodes per second -- but requires the database to be stopped. For online ingestion, `UNWIND` with batched transactions achieves ~50K-100K edges/second depending on hardware.
### Postgres + AGE
This is where AGE currently hurts. Multiple GitHub issues document the pain:
- **Issue #2177**: MERGE on 10K+ edges becomes unusably slow (minutes for 70K edges)
- **Issue #1925**: 83K edge creation causes severe slowdown
- **Issue #2198**: large-table edge creation performance degrades non-linearly
AGE creates edges by inserting rows into a PostgreSQL table. Each edge insert requires lookups to validate source and target vertices. Without proper index pushdown (issue #2348), these lookups scale poorly.
For bulk loading, you can bypass Cypher and insert directly into AGE's internal tables -- but this requires understanding AGE's internal schema, which is undocumented and fragile.
### JSONL
```bash
# 1 million edges
time python3 -c "
import json
for i in range(1_000_000):
print(json.dumps({'src': f'n{i}', 'rel': 'LINKS', 'dst': f'n{i+1}'}))
" > edges.jsonl
```
This completes in seconds. Append is O(1). No index to update, no constraint to check, no transaction to commit. You cannot beat the write speed of "just append text to a file."
But you pay for it later. Every read is now a full scan unless you build your own index.
### Verdict: Round 5
Scored out of 10 on bulk write and ingestion speed, meaning how fast you can load large batches of edges.
| Criterion | Neo4j | AGE | JSONL |
| --- | ---: | ---: | ---: |
| Bulk write and ingestion speed | 7 | 4 | 10 |
JSONL has unbeatable write speed but unbearable read speed. Neo4j is good with offline import and decent online. AGE is the weakest option for large ingestion.
---
## Round 6: Ecosystem and Tooling
Beyond raw queries, daily work leans on everything around the database: visualizers, language drivers (the libraries your code uses to talk to the database), prebuilt graph algorithms, and how big and helpful the community is. Here is what each one gives you.
### Neo4j
- **Visualization**: Neo4j Browser, Bloom (enterprise), extensive third-party tools
- **Language drivers**: official drivers for Java, Python, JavaScript, .NET, Go
- **GDS library**: PageRank, community detection, pathfinding, embeddings
- **APOC**: 500+ utility procedures
- **Community**: large, mature, well-documented
### Postgres + AGE
- **Visualization**: limited; AGE Viewer exists but is basic
- **Language drivers**: any PostgreSQL driver works, but handling `agtype` results requires extra parsing
- **Graph algorithms**: none built-in. You write them in Cypher or fall back to SQL
- **Community**: growing but small. ~2K GitHub stars vs Neo4j's ~13K+
- **Documentation**: improving, but gaps remain. Many "how do I do X?" questions on GitHub issues have no official answer
### JSONL
- **Visualization**: none (build your own, or load into NetworkX/Gephi)
- **Language drivers**: every language has a JSON parser
- **Graph algorithms**: use NetworkX, igraph, or write from scratch
- **Community**: infinite (it is just JSON) and zero (no graph-specific tooling)
### Verdict: Round 6
Scored out of 10 on the breadth of tooling, drivers, and community around each option.
| Criterion | Neo4j | AGE | JSONL |
| --- | ---: | ---: | ---: |
| Ecosystem and tooling | 10 | 5 | 3 |
---
## The Honest Scorecard
Pulling every round together, here are the per-category scores side by side (each out of 10) and a running total.
| Category | Neo4j | AGE | JSONL |
| --- | ---: | ---: | ---: |
| Data modeling | 9 | 7 | 6 |
| Multi-hop queries | 10 | 6 | 3 |
| Mixed workloads | 3 | 9 | 4 |
| Operational cost | 4 | 8 | 7* |
| Write performance | 7 | 4 | 10 |
| Ecosystem | 10 | 5 | 3 |
| **Total** | **43** | **39** | **33** |
\* JSONL drops to 25 at scale because its operational-cost score becomes 2.
But totals are misleading. Nobody needs all six categories equally. Here is the real decision framework:
---
## When Each Wins (Honestly)
The choice collapses to a few honest questions about scale, workload mix, and budget.
### Choose Neo4j when:
- Your core product IS the graph (social network, fraud ring detection, recommendation engine)
- You need multi-hop traversals deeper than 3 hops on 10M+ node graphs
- You can afford the license or AuraDB pricing
- You have a dedicated data engineering team for JVM tuning
- Graph algorithms (PageRank, community detection) are a primary use case
### Choose Postgres + AGE when:
- You already run PostgreSQL and want to add graph queries without a second database
- Your workload is 70% relational, 30% graph
- Your graph is under 5M nodes (AGE's performance ceiling today)
- You need ACID transactions spanning graph and relational data
- Budget is a hard constraint and vendor lock-in is unacceptable
- You value the option to query graph data with SQL joins
### Choose JSONL when:
- You are prototyping or exploring data before committing to infrastructure
- Your graph is small (under 100K edges) and mostly read at startup
- You need maximum portability (the data goes into git, S3, email attachments)
- You are building a data pipeline where the graph is intermediate, not the product
- You want append-only audit logs of graph mutations
- You are one person on a weekend project
---
## When Each Loses (Honestly)
The flip side matters just as much: here are the situations where each approach actively works against you.
### Neo4j loses when:
- You need relational analytics alongside graph queries (you end up with two databases)
- Your startup cannot justify the Enterprise license
- Your graph is small enough that PostgreSQL handles it fine
- You need your ops team to manage one less piece of infrastructure
- Your data model changes frequently (schema migrations in Neo4j are manual)
### Postgres + AGE loses when:
- You need deep traversals (5+ hops) on large graphs -- VLE blows up
- You need `shortestPath()` on cyclic graphs at scale -- it is not BFS
- You need to bulk-ingest millions of edges quickly
- You need graph-specific algorithms (PageRank, Louvain) out of the box
- You cannot tolerate being an early adopter of a maturing extension
### JSONL loses when:
- You need concurrent writes from multiple processes
- You need real-time traversal queries on graphs larger than RAM
- You need ACID guarantees
- You need anyone other than the original author to query the data
- You are building anything that will be in production for more than 6 months
---
## The Falsification Test
Every claim above should be testable. Here are the ones I am least certain about:
1. **"AGE's VLE is O(n^k)"** -- This is documented in AGE issue #2349, but the AGE team may fix it. If you are reading this post after mid-2026, check whether BFS-based path expansion has been merged.
2. **"Neo4j's index-free adjacency makes it O(1) per hop"** -- Strictly true for cache-hot data. When the graph exceeds page cache, it degrades to random disk I/O, which is closer to O(log n) in practice. Neo4j's marketing overstates this.
3. **"JSONL beats all for write speed"** -- True for raw append. But if you need to deduplicate on write, you lose this advantage entirely. Dedup requires an index, and now you are building a database.
4. **"Postgres+AGE is free"** -- The software is free. But if you need graph algorithm support comparable to Neo4j GDS, you will either write it yourself (expensive) or use a separate tool (complexity cost).
---
## Closing Thought
The question is never "which is the best graph database?" The question is "how much of my problem is actually a graph problem?"
If the answer is "almost all of it" -- Neo4j.
If the answer is "some of it, alongside relational needs" -- Postgres+AGE.
If the answer is "I'm not sure yet" -- start with JSONL, learn your access patterns, then migrate to the right tool.
The worst outcome is choosing Neo4j for a problem that is 80% relational, or choosing JSONL for a problem that needs real-time 5-hop traversals. Both are expensive mistakes, just in different currencies: money for the first, engineering time for the second.
```mermaid
flowchart TD
Q1{"Is your core product a graph?"}
Q1 -->|yes| Q2{"Can you afford
Neo4j Enterprise?"}
Q2 -->|yes| NEO["Neo4j"]:::service
Q2 -->|no| NEOC["Neo4j Community (single node)
or Postgres+AGE"]:::storage
Q1 -->|no| Q3{"Need graph queries
in production?"}
Q3 -->|yes| AGE["Postgres+AGE"]:::storage
Q3 -->|no| JSONL["JSONL (prototype, then decide)"]:::config
```
Pick based on where your problem actually lives, not where the marketing says it should.
### Applying Thiel's Monopoly Playbook to Arabic Ed-Tech
- URL: https://mohammadshaker.com/en/blog/thiel-moat-strategy-for-arabic-edtech
- Date: 2026-04-18T00:00:00.000Z
- Tags: strategy, moat, startups, edtech, architecture, AI, Authored with an LLM, Engineering
Peter Thiel says the best startups monopolize a small market before expanding. Charlie Munger says the only real moat is something that compounds. Karl Popper says any strategy that cannot be disproven is not a strategy. We apply all three frameworks to a concrete case: building an Arabic children's learning app in the Gulf. Every claim is falsifiable.
#### Content
Most startup strategy documents are unfalsifiable. "We will build a platform that empowers learners" cannot be proven wrong because it does not say anything specific enough to test. This is not a strategy. It is a prayer.
This post applies three frameworks to a real business -- an Arabic children's learning app targeting Gulf families -- and makes every claim concrete enough to be disproven.
The frameworks:
1. **Thiel (Zero to One):** Monopolize a small market first. Expand concentrically only when you dominate the inner ring.
2. **Munger (Moat):** The only real competitive advantage is something that gets harder to replicate every year you run.
3. **Popper (Falsificationism):** A claim that cannot be refuted by any conceivable evidence is not a real claim. Every strategic assertion must have a metric, a threshold, a deadline, and a data source.
## The Market: Concentric Rings
Thiel's core insight: the failure mode is not picking a market too small. It is picking a market too big to dominate. You end up with 0.1% of a huge market instead of 80% of a small one.
```mermaid
flowchart TD
subgraph R3["Ring 3 — Multilingual kids' edtech (global)"]
subgraph R2["Ring 2 — All Arabic kids' edtech"]
subgraph R1["Ring 1 — Arabic schools (B2B2C)"]
R0["Ring 0 — BEACHHEAD
Gulf + diaspora parents
Kids aged 4-8
MSA + Quranic literacy
Paying on mobile
~1-3M households"]:::client
end
end
end
Rule["Rule: expand only when the inner ring is dominated.
'Dominated' = #1 or #2 in KSA App Store Education/Kids."]:::config
```
### Why Ring 0?
Ring 0 is defensible because:
- **Specific enough to dominate.** Arabic-speaking parents of 4-8 year olds who want both Quranic story content and phonics is a niche that generalist edtech apps (Khan Kids, Duolingo ABC) do not serve well.
- **Small enough to monopolize.** ~1-3M households. A single great product can reach 5-10% penetration.
- **Large enough to be a business.** 2M households x 5% penetration x $100/year = $10M ARR at monopoly share. That funds everything else.
### Why Not Start at Ring 2?
"All Arabic kids' edtech" is too broad. You compete with every Arabic content app simultaneously, from Lamsa to Qamar. You spread marketing budget across too many segments. You build features for 10 different use cases instead of going deep on one.
The failure mode: you build a mediocre app for everyone instead of the best app for one specific parent profile.
## The Moat: What Actually Compounds?
Munger's test: **what gets harder for a competitor to replicate every year you run?**
Most claimed moats in edtech are fake:
- **Content** can be copied (or generated by LLMs).
- **UX** can be cloned.
- **Brand** can be outspent.
- **Network effects** barely exist in single-player learning apps.
So what actually compounds?
### Candidate Moats, Ranked by Falsifiability
Here are five things that could plausibly be a moat for this business, ranked by how strongly each one compounds (gets harder for a rival to copy) the longer we operate. The chart below sketches that ranking; the paragraphs after it explain each candidate and give a concrete test that would prove it is not a moat.
```
MOAT STRENGTH (compounding per year)
│
│ #4 Signal loop ████████████████ ← the real bet
│ #1 TTS corpus ████████████
│ #3 Multi-app ████████
│ #2 Curriculum ██████
│ #5 Brand ███
│
└──────────────────────────────────→ time
```
**#1: Arabic diacritics-to-TTS speech-marks alignment corpus.**
Arabic diacritization is hard. Our text-to-speech pipeline has a custom post-processor that handles cases where display text diverges from TTS pronunciation. ~23% of records have this divergence. Every new content piece teaches the aligner. A competitor starting from scratch needs months of data to reach equivalent accuracy.
*Falsification:* Can a competitor reach equivalent remap accuracy in <6 months using only open Arabic TTS datasets? If yes, this is not a moat.
**#2: Ontology of Arabic children's curriculum.**
Subject to chapter to content byte to content bit, with pedagogical metadata (age group, question type, pay level). Every seeded subject adds a node. Reorder rules encode editorial judgment built over years.
*Falsification:* If we dumped the full curriculum graph as open data, could a competitor rebuild the product? If yes, the moat is editorial workflow, not data.
**#3: Multi-app config matrix.**
Five apps on one codebase. Each new app forces abstractions that the next app inherits for free. The marginal cost of app N+1 drops.
*Falsification:* Does adding app N+1 take less engineering time than app N? If time is flat or growing, no compounding.
**#4: Feedback loop -- content-bit outcomes drive generator tuning.**
This is the real bet. Every child session generates outcome signals (which questions were answered correctly, incorrectly, how fast). These signals flow into the ontology graph. A weekly job identifies under-performing content. It regenerates new variants, A/B tests them, and keeps the winners.
```mermaid
flowchart TD
Kids["Kids use the app
(Ring 0 households)"]:::client -->|generates| OE["OutcomeEvent
(bit answered, correct/not,
child age, time taken)"]:::node
OE -->|feeds| SA["SignalAggregator
(diacritics x age x
pay-level x outcome)
proprietary -- cannot be reconstructed from outside"]:::storage
SA -->|tunes| CG["Content generator
+ TTS aligner"]:::service
CG --> BB["Better bits
next session"]:::node
BB -->|retention + WOM| CG
```
A competitor on day 1 has zero loops. They must pay full customer acquisition cost and wait months to accumulate enough signal. By then, our content has improved through dozens of cycles.
*Falsification:* Run an A/B test. Regenerate 10 under-performing content bits. Test against 1,000 sessions. If regenerated bits do not show >=5 percentage points correct_rate lift (p<0.05), the loop is not closing and moat #4 is dead.
**#5: Brand / parent trust.**
The weakest moat. Replicable with marketing budget. Included for completeness.
*Falsification:* If a funded competitor launches and we lose >20% month-over-month retention, brand was not the moat.
### Inverting the Moat (Munger's "Invert, Always Invert")
What would make the moat evaporate?
1. **OpenAI ships Arabic TTS with perfect diacritics and word-level timestamps.** Probability: medium-high within 24 months. Mitigation: the moat must shift from TTS alignment to pedagogical outcome data, which only running the product generates.
2. **A ministry of education open-sources a richer Arabic curriculum graph.** Mitigation: be the integration layer, not the content owner.
3. **We stop shipping.** The corpus stops compounding. The loop dies. Mitigation: the ontology-as-dashboard forces visibility into whether we are adding signal each sprint.
## The Strategy: Falsifiable Claims
Every strategic claim below has a metric, a threshold, a deadline, and a data source. If the claim cannot meet the threshold by the deadline, it is false and should be abandoned.
### Thiel Claims
These claims test the "dominate a small market first" thesis: that Ring 0 is big enough to matter, that we can win it before a funded rival, and that we can expand outward from it. Each row pairs a claim with the measurement that would falsify it.
| ID | Claim | Test | Deadline |
|---|---|---|---|
| T1 | Ring 0 is large enough for $10M ARR at monopoly share | Bottom-up: 2M households x 5% penetration x $100/yr = $10M. If penetration plateaus <2% after 24 months, T1 is false. | 2028-04 |
| T2 | We reach >50% market share of Ring 0 before a funded competitor | Track monthly active paying families vs App Store chart share in Education/Kids/Arabic for KSA/UAE. If not #1 or #2 in KSA by 2027-01, T2 fails. | 2027-01 |
| T3 | Concentric expansion (Ring 0 to Ring 1 schools) converts at >20% | Pilot 5 schools. If <2 convert to paid, do not expand. | 2027-04 |
### Moat Claims
These claims test whether the candidate moats above actually hold up under measurement: the TTS (text-to-speech) alignment corpus, the content regeneration loop, and the multi-app abstraction. Each one names the experiment that would kill it.
| ID | Claim | Test | Deadline |
|---|---|---|---|
| M1 | TTS alignment corpus is a barrier | Can a new entrant reach equivalent accuracy in <6 months with open data? If yes, not a moat. | ongoing |
| M2 | Content regeneration loop closes | A/B: regenerated bits >=5pp correct_rate lift, p<0.05, n>=1000 | 2026-08-31 |
| M3 | Multi-app abstraction compounds | Time to launch app N+1 < time for app N | measure at each launch |
### Awareness Claim
This claim tests whether target parents actually know the product exists. "Unprompted awareness" means a parent names the app on their own, without being shown a list, which is a stronger signal than recognizing it when prompted.
| ID | Claim | Test | Deadline |
|---|---|---|---|
| A1 | Unprompted awareness >30% among target parents who have tried >=2 Arabic learning apps | Quarterly survey of 500 parents via in-app prompt | 18 months from now |
If A1 stays below 10% after 12 months, the Ring 0 beachhead strategy is failing.
## Revenue Math
Three scenarios, all falsifiable against actual RevenueCat and analytics data.
### Scenario A: Do Nothing (No Ontology, No Loop)
This is the baseline: keep shipping the app as-is, with no graph and no self-improving content loop. The table tracks paying families, annual recurring revenue (ARR), and monthly churn (the share of paying families who cancel each month).
| Year | Paying families | ARR | Churn |
| ---: | ---: | ---: | ---: |
| 1 | 4,000 | $320k | 7%/month |
| 2 | 7,000 | $574k | 7%/month |
| 3 | 10,500 | $890k | 7%/month |
Ceiling: about $1M. Churn eats acquisition.
### Scenario B: Ontology + Content Loop (No B2B)
Churn drops 7% to 5.5% to 4.5% as content quality compounds.
| Year | Paying families | ARR | Delta vs. A |
| ---: | ---: | ---: | ---: |
| 1 | 4,200 | $344k | +$24k |
| 2 | 8,500 | $748k | +$174k |
| 3 | 14,000 | $1.33M | +$440k |
### Scenario C: Full Strategy (Ontology + Loop + B2B Schools)
This adds a second revenue line on top of Scenario B: selling the same graph to schools as a teacher dashboard. B2C is direct-to-consumer (parents paying in the app); B2B is the school contracts. The table splits the two so you can see where growth comes from.
| Year | B2C ARR | Schools | B2B ARR | Total |
| ---: | ---: | ---: | ---: | ---: |
| 1 | $344k | 0 | $0 | $344k |
| 2 | $748k | 5 | $8k | $756k |
| 3 | $1.33M | 30 | $48k | $1.38M |
| 5 | — | 300 | $480k | about $3M |
### Cost of Inaction (3-Year)
Flipping the question around: what does staying in Scenario A actually cost us over three years? This table sums the revenue and time we forgo by not building the ontology and the loop.
| Foregone | Amount |
| --- | ---: |
| ARR uplift (A → B) | about $640k |
| B2B years 2–3 | about $56k |
| B2B years 4–5 (forward) | about $640k |
| Engineering time drag | about $170k |
| **Total three-year opportunity cost** | **about $860k+** |
## The Execution Plan
Phased delivery. Each phase has a kill gate. Total committed before the first real kill gate: 1.6 engineering weeks.
```mermaid
flowchart LR
P0["P0"]:::service --> P1["P1"]:::service
P1 --> P2["P2"]:::service
P2 --> P3["P3"]:::service
P3 --> P6["P6 (planner agent)"]:::node
P3 --> P7["P7 (personalization)"]:::node
P3 --> P8["P8 (schools)"]:::node
P0 -->|fail| ABORT["ABORT"]:::danger
P1 -->|fail| WIKI["wiki only"]:::danger
P2 -->|fail| NOLOOP["no loop"]:::danger
P3 -->|fail| OUT1["Outcome 1 only"]:::danger
```
| Phase | What | Eng-wk | Kill gate |
|---|---|---:|---|
| P0 | Plan discipline test | 0.1 | YAML front-matter on 3 recent plans by 2026-04-25 |
| P1 | Ontology spike (JSONL + extractors + agent) | 1.5 | Agent beats SQL on 20 support questions by 2026-06-01 |
| P2 | OutcomeEvent stream | 1.5 | 99% delivery for 7 days by 2026-06-20 |
| P3 | Moat experiment (10-bit regen A/B) | 3.0 | >=5pp correct_rate lift by 2026-08-31 |
| P7 | Personalized daily lesson | 3.0 | +3pp D30 retention by 2027-01-31 |
| P8 | Teacher dashboard pilot | 4.0 | >=2 of 5 schools convert to paid by 2027-04-01 |
Each gate authorizes the next spend. No gate, no money.
## What Could Be Wrong
Every strategy post should end with an honest "here is where I might be fooling myself" section.
1. **Ring 0 might be too small.** If the addressable market is 500k households, not 2M, then $10M ARR at monopoly share is $2.5M. Still a business, but not the same thesis.
2. **The content loop might not close.** If regenerated content is not meaningfully better -- if the LLM produces lateral variations rather than genuinely improved questions -- then correct_rate does not improve and moat #4 is dead. P3 is explicitly designed to catch this.
3. **Parents might not care about Quranic + MSA combined.** If the beachhead is actually two separate markets (religious parents who want Quran content vs secular parents who want MSA literacy), then the niche is even smaller than modeled.
4. **A funded competitor could brute-force the signal gap.** If someone raises $50M and acquires 100k users in 6 months, they accumulate outcome signals fast enough to match our loop. The moat buys time, not invincibility.
5. **We might not maintain the ontology.** If the schema drifts from reality and nobody updates it, the agent gives wrong answers, trust collapses, and the whole system becomes shelfware. This is why P0 is a discipline test, not a technical test.
## The Bottom Line
Thiel says dominate a small market. Munger says build something that compounds. Popper says make it falsifiable.
Applied to Arabic kids' ed-tech:
- **Small market:** Gulf parents of 4-8 year olds who want Quranic + MSA literacy.
- **Compounding:** the signal loop -- every child session makes the next content generation better.
- **Falsifiable:** if regenerated bits do not show >=5pp correct_rate lift by August 2026, the moat thesis is dead and we pivot to a simpler business.
The worst outcome is not that the thesis is false. The worst outcome is believing it without testing it. Every claim above has a number, a date, and a kill switch.
### How I Overhauled SEO, GEO, and AEO for My Personal Website
- URL: https://mohammadshaker.com/en/blog/seo-geo-aeo-improvements-for-mohammadshaker-com
- Date: 2026-03-26T00:00:00.000Z
- Tags: AI, nlp, react, design, Authored with an LLM, Engineering
Overhauling SEO, GEO, and AEO on a bilingual Next.js site with 300+ posts required structured data on every page type, dynamic OG images per locale, speakable schema for voice search, and an llms.txt file for AI crawler discoverability. This post documents every implementation decision, the trade-offs, and what made the biggest measurable difference.
#### Content
I recently spent a focused session overhauling the search engine optimization, generative engine optimization, and answer engine optimization of my personal website — mohammadshaker.com. This is a bilingual (Arabic/English) Next.js site with 300+ blog posts, academic publications, comparative studies, and interactive visualizations.
Here is what I did, why, and what I learned.
## What Are SEO, GEO, and AEO?
**SEO (Search Engine Optimization)** is the classic discipline of making your site discoverable and rankable by Google, Bing, and other search engines. It covers metadata, sitemaps, canonical URLs, structured data, and page speed.
**GEO (Generative Engine Optimization)** is the newer practice of making your content appear in AI-generated answers — ChatGPT, Google AI Overviews, Perplexity, and similar systems. It focuses on structured data that AI can parse, content formatting that makes extraction easy, and authority signals.
**AEO (Answer Engine Optimization)** targets featured snippets, People Also Ask boxes, voice search results, and knowledge panels. It is about structuring content as direct answers to questions.
## The Audit
I started by running three parallel audits — one for each discipline. Each audit analyzed the entire codebase: page metadata, structured data, content formatting, sitemap coverage, robots.txt, and AI discoverability signals.
### Key Findings
**SEO gaps**: The homepage used static metadata without canonical/hreflang. Seventeen section index pages had no descriptions. Tech blog posts had no hreflang in the sitemap. The 404 page redirected instead of returning a proper status code.
**GEO gaps**: Tech blog posts had zero JSON-LD structured data. FAQ content on the About page existed as invisible JSON-LD but was not rendered visually. Blog post tags mapped to generic "Thing" entities without Wikidata URIs. No author bio was visible on posts.
**AEO gaps**: No HowTo schema for tutorial posts. No visible breadcrumb navigation. Speakable schema only targeted h1 and excerpt. No People Also Ask targeting. No table of contents for long articles.
## What I Built
### Structured Data on Every Page
Every page type now has appropriate JSON-LD:
- **Blog posts**: BlogPosting with speakable, citations, Wikidata entity URIs
- **Tech posts**: BlogPosting + Breadcrumb (previously had nothing)
- **About**: Person, ProfilePage, Breadcrumb, FAQPage (now visible too)
- **Publications**: Breadcrumb, FAQ, ScholarlyArticle ItemList
- **Section pages**: ItemList + Breadcrumb on all 18 study sections
- **Category pages**: New — Breadcrumb + ItemList, with per-category RSS
- **Contact, Studies**: Breadcrumb schema
### Metadata Everywhere
- Converted 17+ static `export const metadata` to `generateMetadata` with locale-aware canonical URLs and hreflang alternates
- Added descriptions to 8 layout files
- Added description + canonical to all 17 chapter/poem/surah detail pages
### Content Enrichment
- **TL;DR summaries** on all 300 blog posts
- **FAQ sections** on 10 key technical posts targeting People Also Ask
- **Author bio** at the bottom of every blog post
- **Related posts** section using category + tag scoring
- **Table of contents** auto-generated from headings with active section tracking
### New Pages and Features
- **Search page** (`/search`) with client-side filtering and SearchAction schema
- **Category landing pages** (`/blog/category/[category]`) with structured data and RSS feeds
- **Per-locale RSS feeds** (`/en/feed.xml`, `/ar/feed.xml`)
- **Reading progress bar** on blog and tech posts
- **Dynamic OG images** generated per blog post with dark/light variants
- **Visible breadcrumb** navigation replacing back links
### Infrastructure
- **Sitemap ping** to Google and Bing on every production deploy
- **Lighthouse CI** workflow in GitHub Actions with performance/a11y/SEO thresholds
- **robots.txt** updated to block `/admin/` paths
- **Proper 404 page** instead of redirect
- **`llms-full.txt`** expanded with full article bodies for AI crawlers
## Technical Approach
The entire implementation was done in a single session using parallel subagents — multiple independent changes running simultaneously across different files. This allowed ~150 files to be changed across 13 commits without conflicts.
Key architectural decisions:
1. **Centralized schema generation** in `lib/seo/structured-data.ts` — all JSON-LD generators in one file
2. **Reusable Breadcrumb component** — one component used across blog, tech, search, and category pages
3. **Script-based summary generation** — a TypeScript script that derives summaries from excerpts for bulk updates
4. **File-based OG images** — using Next.js `opengraph-image.tsx` convention for automatic integration
## Results
After implementation:
- Every page has structured data (JSON-LD)
- Every page has proper canonical URL and hreflang alternates
- All 300 posts have TL;DR summaries
- 10 key posts have FAQ sections for PAA targeting
- Blog tags map to Wikidata entity URIs
- Search functionality enables the WebSite SearchAction schema
- Category pages create topical clusters for authority signals
- Per-locale RSS feeds properly tag content language
## What I Would Do Differently
1. **Start with the search page** — having it early would have made the SearchAction schema valid from the start
2. **Batch content changes more carefully** — some subagents made changes beyond their scope, requiring cleanup
3. **Test OG images locally** — dynamic OG images are hard to preview without deploying
## Frequently Asked Questions
### What is GEO (Generative Engine Optimization)?
GEO is the practice of optimizing content to appear in AI-generated answers from systems like ChatGPT, Google AI Overviews, and Perplexity. It focuses on structured data, content formatting, entity relationships, and authority signals that AI systems can parse and cite.
### How is AEO different from SEO?
AEO (Answer Engine Optimization) specifically targets direct answer formats: featured snippets, People Also Ask boxes, voice search results, and knowledge panels. While SEO aims for high rankings in search results, AEO aims to be THE answer that appears above all results.
### Do you need structured data for every page?
Not necessarily, but it helps significantly. At minimum, your homepage (WebSite), about page (Person), and content pages (Article/BlogPosting) should have JSON-LD. The more structured data you provide, the more context search engines and AI systems have about your content.
### GAIA's S-Curve of Agent Effectiveness
- URL: https://mohammadshaker.com/en/blog/gaia-s-curve-of-agent-effectiveness
- Date: 2026-02-13T00:00:00.000Z
- Tags: AI, agents, benchmarks, Authored with an LLM, Research
Agent effectiveness on the GAIA benchmark follows an S-curve. Plain LLMs plateau early because the failure mode is execution, not reasoning. Tool use drives the steep middle phase. The top systems now reach ~91–92% — matching human-level performance — but the remaining errors live in the long tail of edge cases.
#### Content
We can think of GAIA as a stress test for "general assistants" that must do what humans casually do all day: find the right source, read it correctly, combine a few steps, and give a precise answer. Not just "reason" in the abstract, but execute: browse, extract, verify, and finish cleanly.
When we plot agent effectiveness on GAIA against time and capability, it naturally forms an S-curve.
At the beginning of the curve, plain LLM behavior doesn't help much. We can write plausible text, but GAIA tasks punish plausibility. Without disciplined tool use, the system either can't reach the needed information or can't assemble it reliably. Improvements in prompting and base reasoning move the needle, but not dramatically, because the failure mode is execution, not eloquence.
Then the middle of the curve arrives and the slope gets steep. This is where tool use and orchestration show up: search, browsing, structured extraction, multi-step planning, retries, and basic self-checks. Once an agent can consistently do "find → read → compute → answer" instead of guessing, accuracy jumps fast. This is the part of the S-curve that feels like progress is suddenly compounding.
Finally we hit the plateau. Not because we stopped improving, but because the remaining errors live in the long tail. The last few percentage points aren't about doing the common case better. They're about not breaking on messy pages, not selecting the wrong source when multiple are plausible, not misreading a table in a PDF, not dropping a constraint halfway through, and recovering when an early step goes wrong.
On GAIA specifically, the human baseline is roughly in the low 90s. The best agent systems on the public leaderboard are now essentially there as well: the overall average is in the ~91–92% range. In other words, in GAIA terms, we're already in the top-right of the chart: phase three, near the asymptote.
Here's the mental picture:
```
Effectiveness (GAIA %)
100 | ________ human ~92%
95 | _____/
90 | _____/ we are here (SOTA ~91–92%)
85 | _____/
80 | _____/ steep gains: tools + orchestration
70 | ___/
60 | __/
50 | _/ early: LLM-only struggles on execution
40 | _/
30 |/
0 +---------------------------------------------------------
Phase 1 Phase 2 Phase 3
(no tools) (tools+agentic) (robust autonomy)
```
What changes once we're on that plateau is the nature of work that matters. Benchmarks become less about average score and more about reliability characteristics: variance, tail risk, and "cost to correct." The question stops being "can we solve this kind of task?" and becomes "how often do we fail in annoying ways, and how expensive is it for a human to catch and fix it?"
This also explains why Level 1 tasks look "basically solved" while harder levels still leak errors. The limiting factor isn't raw intelligence; it's robustness under ambiguity and messy real-world inputs. A system can be brilliant and still pick the wrong page. It can reason correctly and still extract the wrong number from a table. It can follow a plan and still silently drop a constraint.
So where are we in the chart? We're already at the part where gains come from boring, high-leverage engineering: verification loops, provenance discipline, better fallback strategies, and tight recovery when the first attempt goes off the rails. Once the mean is near-human, the differentiator is not "more IQ." It's fewer unforced errors.
### Reading Books at the Right Time
- URL: https://mohammadshaker.com/en/blog/reading-books-at-the-right-time
- Date: 2025-11-26T00:00:00.000Z
- Tags: personal, reading, Human-Written
On reading books at the right time. I first encountered The Black Swan by Nassim Taleb while doing my MSc in France in 2014, listening to it as an
#### Content
On reading books at the right time.
I first encountered The Black Swan by Nassim Taleb while doing my MSc in France in 2014, listening to it as an audiobook.
I hated the book. And I hated Nassim Taleb.
Seventeen years later, he became the single most influential author in my life. His work changed the way I think—and changed my life.
I’ve read all of his books, each at least four times. I just finished Fooled by Randomness for the fifth time—an exceptional book, and the one I gift most often.
Some books only matter when they meet us at the right moment.
I don’t know why. And I don’t know the formula. It’s just that if we are ready for the book, we’ll cherish it. If not, not.
We should just keep our mind open and our curiosity alive.
Salam, peace.
### What Would I Spend 10,000 Hours On?
- URL: https://mohammadshaker.com/en/blog/what-would-i-spend-10000-hours-on
- Date: 2025-08-06T00:00:00.000Z
- Tags: entrepreneurship, trust, bootstrapping, Human-Written
If I had 10,000 hours to invest, I'd split them between building products and building trust — not raising capital. Engineering is necessary but not sufficient; marketing and distribution compound just as hard. Bootstrapping wins because the incentives stay aligned. Skills that don't age — systems thinking, writing, shipping — are always the best bet.
#### Content
I wrote this back in 2021. I was thinking about where I'd invest significant time and effort.
I have a strong preference for building products. But I have a weakness — marketing and sales. I used to dismiss them as inferior to engineering work.
That dismissal stems from an engineering-centric mindset that undervalues humanistic skills like storytelling and trust-building. But the truth is: **we live by trust, by telling stories**. These qualities cannot be reduced to quantifiable metrics.
Successful product adoption requires people to trust the creator enough to purchase what they've built.
## Trust is evolutionary
Our ancestors lived in small villages. Human behavior was shaped around trust and personal relationships. People befriend those they trust, buy from friends, and marry into trusted circles.
Kevin Kelly's concept of "1,000 true fans" captures this perfectly. Bootstrapped, non-VC ventures require earning initial small groups of loyal supporters — progressing from 10 to 100 to 1,000 fans through localized, incremental wins.
## No VC
I have a firm stance against the venture capital model. I prefer to build ventures using personal capital, **the same way my grandfather did.**
### Agents Are Eating the Application Business Layer
- URL: https://mohammadshaker.com/en/blog/agents-are-eating-the-application-business-layer
- Date: 2025-06-25T00:00:00.000Z
- Tags: personal, AI, react, Human-Written, Engineering
LLM-powered agents are replacing the traditional business logic layer in software. Instead of hardcoded rules and workflows, agents decide at runtime which tools to call, which APIs to hit, and when to delegate. The familiar three-tier architecture is collapsing into a thin CRUD layer plus an agent reasoning layer.
#### Content
Lately, I have spent the first 1h each morning tinkering with agents, embeddings, RAG, multimodal RAG, LangChain, Haystack, etc ([www.deeplearning.ai](http://www.deeplearning.ai) has amazing 1-2hrs courses. You can digest 1 each morning.)
The more I tinker with the technology, the more I see how agents can literally change the whole way we design and ship software.
If I’m to write a prod-ready software from scratch, there’s a completely new way to do it with Agents. It’s simply this:
1. Write agents with function calling, tools and MCP links.
2. Thin CRUD layer on top of a storage layer (DB.)
3. Let agents talk directly to that thin CRUD layer and between each others.
That’s it. No fixed application or business logic. Each agent will decide which tool/function to use on top of the CRUD layer of a DB. And when to call other agents for help outside their domain (expertise or content/knowledge area)
Taking a step back now.
For the last two decades, we’ve seen most software stack layers get abstracted, outsourced, or minimized. Frontends went from handcrafted jQuery to SPA to a React runtime. Infrastructure went from hand-rolled Chef scripts to serverless YAML.
Every software out there is built around a familiar three‑tier model:
1. **Presentation** (front‑end or API)
2. **Application / business logic**
3. **Data and storage layer**
LLM‑powered agents collapse tiers 1 and 2 (and sometimes parts of 3) into a single reasoning layer. The agent decides what to show and what to call next from a list of tools (functions or other MCPs); the database becomes a very simple, schema‑enforced persistence tier. Think of it as _serverless, except the function is intelligent_.
Traditional applications center around deterministic business logic. Take:
if user.is\_premium:
apply\_discount()
Logic is gated by feature flags, permissions, pricing tiers—each wrapped in code paths and tightly scoped unit tests.
Take customer support. What was once a flowchart of conditionals, database lookups, and email dispatching is now a retrieval-augmented agent querying embeddings, asking clarifying questions, and summarizing previous tickets.
Business logic used to be the “meat” of an application. Now, increasingly, it’s becoming scaffolding for agents.
A table can help us understand:
Agents don’t need all the rigid business logic—only tools (APIs), memory and storage (DB), and context (retrieval). They decide _how_ to use them, not you or me. Your app becomes a toolkit. The agent decides what hammer to use.
**A shift from Logic to Language.** The shift is epistemic. We’re moving from codified logic to language-driven reasoning. From state machines to semantics. From “how does this function work” to “what does the good outcome looks like?”
The agent doesn’t need to know every rule. It needs to know when to ask, where to look, and how to act.
**This Changes Value Creation in Businesses.** This enables something like this:
Instead of a 4-step wizard, lead with outcome. Describe what this flow should lead to. Let the agent deduce the rest.
The job of a product manager shifts from defining flows to defining _outcomes_. This also changes the engineer’s work from defining fixed rules to defining _capabilities_. The job of the engineer shifts from building guardrails to exposing tools (capability) safely.
Not everything will be agentified. Critical paths—like payment processing, identity verification, or compliance checks—still demand determinism. But increasingly, those become thin wrappers. The bulk of the experience, the thing that _feels_ intelligent, adaptive, and helpful—will be the agent.
Final thoughts for a panicking industry. If you are an engineer interested in this, the only way to help us understand this shift is to tinker with it. Something you can do today is the following:
1. **Identify a narrow self-contained slice in the codebase** like a workflow (e.g., invoice approval).
2. **Expose CRUD endpoints** (REST or GraphQL). Keep them boring.
3. **Wrap the endpoints in tools** for your agent framework (Adept, AutoGen, LangChain, Haystack).
4. **Ship with a human‑in‑the‑loop** review step; measure approval rates.
5. **Iterate**: evolve guardrails, expand autonomy gradually.
I’m still tinkering with this around the edges and will have something more technical to share soon.
## Frequently Asked Questions
### What is agentic AI?
Agentic AI refers to artificial intelligence systems that can autonomously plan, execute, and adapt multi-step tasks without continuous human intervention. Unlike traditional AI that responds to single prompts, agentic systems maintain context, use tools, and make decisions across complex workflows.
### How do AI agents replace business logic?
AI agents replace traditional business logic by handling decision-making, routing, and orchestration that was previously hardcoded. Instead of writing if/else chains and state machines, developers define goals and constraints, and agents determine the optimal path through available tools and APIs.
### What is MCP in AI?
MCP (Model Context Protocol) is a standard for connecting AI models to external tools and data sources. It provides a unified interface for agents to interact with APIs, databases, and services, similar to how USB standardized hardware connections.
### Agents Are Eating the Business Layer.
- URL: https://mohammadshaker.com/en/blog/agents-are-eating-the-business-layer
- Date: 2025-06-25T00:00:00.000Z
- Tags: AI, agents, architecture, LLM, Human-Written, Engineering
LLM-powered agents are replacing the traditional business logic layer in software. Instead of hardcoded rules and workflows, agents decide at runtime which tools to call, which APIs to hit, and when to delegate. The familiar three-tier architecture is collapsing into a thin CRUD layer plus an agent reasoning layer.
#### Content
LLM-powered agents are fundamentally transforming software architecture. They are collapsing traditional three-tier application design into a new paradigm.
## The Traditional Stack vs. Agent-Powered Architecture
Historically, software followed this structure: presentation layer, business logic layer, and data storage. **LLM-powered agents collapse tiers 1 and 2 (and sometimes parts of 3) into a single reasoning layer.**
In this emerging model, agents become the decision-makers, selecting from available tools rather than following predetermined code paths. The database becomes a simplified persistence mechanism with a thin CRUD interface.
## The Architectural Shift
Rather than hardcoded conditionals and service pipelines, agents equipped with function-calling capabilities, tools, and MCP (Model Context Protocol) links handle decision-making autonomously. The business logic shifts from "how do we process this?" to **"what capabilities does the system need?"**
Consider customer support: what once required flowcharts of conditionals now involves retrieval-augmented agents querying embeddings and synthesizing previous interactions.
```mermaid
flowchart TB
subgraph OLD["Before: three tiers"]
p1["Presentation"]:::client
b1["Business logic (rules, pipelines)"]:::config
d1["Data store"]:::storage
p1 --> b1 --> d1
end
subgraph NEW["After: collapsed"]
p2["Presentation"]:::client
a2["Agent reasoning layer (tools + MCP)"]:::service
d2["Thin CRUD store"]:::storage
p2 --> a2 --> d2
end
b1 -. "collapses into" .-> a2
```
## Where to Start
For engineers wanting to explore this shift:
1. Start with a narrow workflow slice
2. Expose CRUD endpoints
3. Wrap them as agent tools
4. Implement human-in-the-loop review
5. Iterate on autonomy levels
The transformation represents an epistemological shift — from codified logic to language-driven reasoning.
### Voluntary Discomfort.
- URL: https://mohammadshaker.com/en/blog/voluntary-discomfort
- Date: 2025-06-16T00:00:00.000Z
- Tags: stoicism, growth, discipline, Human-Written, Philosophy
Discomfort and adversity, rather than comfort, catalyze personal growth and push individuals beyond perceived limitations.
#### Content
I survived my Master's program in France while nearly broke. Eating sparingly and living frugally. Rather than viewing this as hardship, I felt content and focused on the ultimate outcome — securing a well-paying software engineering position at Squla upon graduation.
## Discomfort as catalyst
Discomfort and adversity, rather than comfort, catalyze personal growth and push individuals beyond perceived limitations.
I advocate for deliberately creating manageable hardships as a practice: extended fasting, intensive workouts, accelerated learning, or rapid team turnarounds.
## Seneca knew this
Seneca advised:
> "Set aside a certain number of days...with the scantiest...fare, saying to yourself...'Is this the condition that I feared?'"
Soldiers train during peacetime to prepare for actual conflict. Similarly, individuals should toughen themselves proactively through self-imposed challenges rather than waiting for crisis to strike.
I haven't yet determined how to systematically create this practice or teach it to future children, but I recognize the necessity of doing so.
### Weekly: Drinking from Small Cups
- URL: https://mohammadshaker.com/en/blog/weeklydrinking-in-small-cups
- Date: 2025-06-11T00:00:00.000Z
- Tags: personal, AI, deep-learning, aws, reading, Human-Written
It's reading a deep \[hard\] books or learning something new for an hour after an exercise/running session.
#### Content
## **Something new I’m trying**
It's reading a deep \[hard\] books or learning something new for an hour after an exercise/running session.
Whenever I’m back from the gym, I shower, and I just sit down and read a thorough book for 30-60 minutes - fresh out of a refreshing session.
It’s the same thing I’m trying to do after waking up. Wake up, sit down and read a hard science book or learn something hard.
When our mind is calm and clear in the morning, it’s the best to fill it in with the most hardcore stuff, the most difficult. Same after an exercise - the blood is flowing and the mind is clearer.
It works.
### **Something I keep doing**
- Drinking tea from small cups pouring from a teapot, like we used to do back home in Syria. We used them when I was 3. Excellent in the morning with green teapot and a hard book.
- (You can still buy them online -- from the French manufacturer Luminar. Or Duralex.)

### **What I'm reading**
I’m reading multiple books on parallel. Here are some of books I've read recently.
- Technical/Engineering
- Ben Stopford on Kafka: **Designing Event Driven Systems** is a very enjoyable concise read.
- **AWS Certified Solutions Architect** (All available books. Read 5 so far. None is the best tbh.)
- What is ChatGPT Doing, Wolfram.
- Deep Learning courses on MCP, Agentic AI, RAG systems and MM-RAG. Just search www.deeplearning.ai and you'll find many. All under 2hrs. Attend 1 course on LangChain and you'll have many ideas to tinker with within 30 minutes.
- History and Philosophy:
- **History of Islamic Philosophy by Majid Fakhry** is an amazing read (language/readability is a bit rough - still an amazing read.)
- **The Cambridge History of China (Buckly.)** Good, not best.
- Daily Life in Ancient Mesopotamia, Nemet Nejat. Good, not best.
### **1 thing that is making me happy**
Reading autobiographies. I’ve listened Ford’s autobiography across two days on 4 long walks (via Audible, not that I like Audible.)
### **Art or Poetry I'm reading**
I'm reading Al-Mutanabbi poetry from a very old book with yellow pages from the old library of my father-in-law (RIP.)
### **Say thanks for 1 Person today**
And my father-in-law, who passed away on Sept, 2023.
### Don’t Take Anything for Granted.
- URL: https://mohammadshaker.com/en/blog/dont-take-anything-for-granted
- Date: 2025-06-04T00:00:00.000Z
- Tags: personal, Human-Written
Don’t Take Anything for Granted. One of the good things that I learned early in management of engineering functions is this: never force a ready-made
#### Content
Don’t Take Anything for Granted.
One of the good things that I learned early in management of engineering functions is this: never force a ready-made template that worked elsewhere onto a new team or function.
Yes, we should start from principles, but not from fixed practices.
Each company, product, priority, and team—along with the team’s skills and purpose—differs.
If all these vary, how can the same rule suit every company?
No single process fits every engineering team—whether Kanban, Scrum, waterfall, one-week sprints, two-week sprints, or project milestones.
We learned to walk by falling many times. Bottom-Up, Never Top-Down.
The soundest principle I found to work every time is to start small. Plan for the short term first, a 1 week amount of work, then proceed.
This rule is constant: start small, start local.
From there, change in cycles: spot a problem, propose a solution, implement, review, and begin a new cycle (call it a sprint, week, or epic—it must end so you can reassess).
This is bottom-up action. In each cycle the team discovers what works. Not the EM, the director, or the CTO, but the team itself. It adapts locally and iterates.
Thus everyone helps decide what each next cycle should be, which practices to adopt, and which to discard.
With enough cycles, teams discover the few processes that matter. Control only the essential variables; ignore the rest.
Keep policies to a minimum. Too many feel robotic and dictatorial. Preserve an open-market style of work.
Start small. And keep an open mind with min # of policies.
Hire mediocre people and you need to force a way. The principles above won't hold.
Hire great people and let them find their way. The principles above hold.
### (De)invest for the Long Term.
- URL: https://mohammadshaker.com/en/blog/deinvest-for-the-long-term
- Date: 2025-05-29T00:00:00.000Z
- Tags: Brain dump, Human-Written
In the engineering context, most advice around the topic of what to build/not to build comes from the following:
#### Content
In the engineering context, most advice around the topic of what to build/not to build comes from the following:
> Invest for the long term.
Which is absolutely true. The problem though is that most engineers, would view this as a license of investing in _everything_ since _everything_ is needed in the long term.
Another approach to counter this way of thinkings seems to be:
> Deinvest for the long term.
Which means asking the question:
> What are things that we should ditch and are NOT worthy in the long term?
This means that we must _not_ do anything that’s _not_ tied to our core value proposition in the long term.
Example: Building our own logging system? This is a solved problem. Absolutely not a core value prop in the long term for most startups or companies. That means we must not do our own but buy or use a 3rd party solution.
A simple principle that you can apply anywhere.
Salam, peace.
### Freedom Is Hard
- URL: https://mohammadshaker.com/en/blog/freedom-is-hard
- Date: 2025-05-21T00:00:00.000Z
- Tags: freedom, immigration, london, startups, Human-Written, Life
Why I chose London — tracing my journey from Damascus through France and Amsterdam, chasing freedom above all else.
#### Content
I chose London as my home. Let me trace my journey from Damascus, Syria through France and Amsterdam before settling in the UK in 2018.
## Why London?
I selected London — alongside Berlin — because it represents **the best place in Europe for startups and for taking risks**, with the US and China being obvious alternatives outside the continent.
## The timeline
- **2013**: Completed a five-year IT and AI degree in Syria
- **2014–2015**: Pursued an M.Sc. in Ubiquitous Computing and AI in France
- **2018**: Worked at Dutch ed-tech startup Squla in Amsterdam
- **2018–present**: Relocated to the UK with only a Syrian passport
## The freedom factor
The defining advantage was the UK's Exceptional Talent visa (now Global Talent visa), which grants five-year residency without employment restrictions. This contrasted sharply with Germany, the Netherlands, and France, where non-EU citizens faced constraints on entrepreneurship.
I've leveraged this freedom across multiple roles: employee at Neurofenix, founder of Almeta (later Alphazed), Head of Engineering at Noon, and CTO co-founder of SpatialX.
**Freedom is precious and hard-won; it should be cherished above all else.**
### Weekly: Good Friends.
- URL: https://mohammadshaker.com/en/blog/good-friends
- Date: 2025-05-14T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. Something I keep doing Training, in the gym. Reading in the morning. Long walks on weekends.
#### Content
Here we go again.
## **Something I keep doing**
- Training, in the gym.
- Reading in the morning.
- Long walks on weekends.
### **Something I am changing**
- Not eating fruits and sugar at the end of the day. This one is big for me. Given my 1 meal a day regimen, the suger-kick kicks in and I'll really crave some sugar. I started substituting chocolate with fruits (blueberries.) And now I'm trying to substitute sugar with nothing.
- Having my 1 meal a day before 8pm.
### **Best of what I watched**
[Webb images on space.](https://science.nasa.gov/mission/webb/multimedia/images/) I'm speechless.

### **What I'm reading?**
- Wealth of Nations, Adam Smith.
- Understanding Understanding, Richard Saul Wurman.
### **Something I’m remembering from the past**
My dad. And my mom.
### **Say thanks for 1 guy today.**
My friend, Hussam Sandouk. For pushing me to have better answers and clearer mind. Hussam is one of the smartest people I know. And I'm lucky to have friends who are way smarter than I am.
Being weak among the strong, makes me stronger. Thank you, friends.
### “How Can I Help?”
- URL: https://mohammadshaker.com/en/blog/how-can-i-help
- Date: 2025-05-07T00:00:00.000Z
- Tags: Brain dump, Human-Written
“How can I help?” — It’s a simple question, but one that’s carried me further than most job titles or credentials.
#### Content
“How can I help?” — It’s a simple question, but one that’s carried me further than most job titles or credentials.
We help others because it fills something within us—true. But real help isn’t about making yourself “feel” useful; it’s about “actually” being useful.
In business, asking “How can I help?” isn’t a performance. It’s a prompt for action.It’s an act of leadership—quiet, service-oriented, and rare.
Years ago, I read a tactic that opened my eyes:
- Find what your boss hates doing—or what they spend the most time on.
- Figure out how to take it off their plate.
- Document your plan.
- Share it with them.
- Act on it.
- Report back on it, weekly.
It sounds simple. It is. That’s why it works.
The best people reduce load. They show up with solutions, not problems. They subtract from the chaos instead of contributing to it.
The question scales. It works with your boss, with your peers, and with your team (or “direct reports,” though I’ve never liked the phrase). You can lead in every direction if you start from this place.
But—like all good habits—it’s easy to forget. Busyness creeps in, and the instinct to help dulls. This is a reminder to sharpen it. To make it a rhythm. At least once a week, ask:
How can I help?
It’s a better question than “Can I help?”
The latter invites a one-word answer.
The former invites collaboration.
### Knowing What I Can't Control.
- URL: https://mohammadshaker.com/en/blog/know-what-i-can-not-control
- Date: 2025-04-30T00:00:00.000Z
- Tags: personal, startups, reading, Human-Written
When should I adopt a new technology? When should I take risks? When should I play it safe?
#### Content
When should I adopt a new technology? When should I take risks? When should I play it safe?
It all ties to risk. Ten years after reading and rereading Nassim Nicholas Taleb’s work, I think I understand a bit more.
I must know what falls within my circle of control. More importantly, I must know what I cannot control.
This view of looking at everything from the prism of risk is echoed by the thinkers I admire—Hayek, Taleb, and Seneca. In economics, in daily life - and in what I try to do in tech.
12 years ago I learned the first step:
"Focus on what’s under my control and leave what is not."
10 years ago I learned to take it one step further:
"Focus on what’s under my control and render what’s not ineffective" (Antifragile, Taleb)
This is the essence of Taleb’s barbell strategy. In investing, keep 90 % of capital in ultra-secure assets and put the remaining 10 % in highly speculative bets with unlimited upside (and downside). However that 10 % fares, you won’t be ruined; your tranquility remains intact.
In tech and startups, the analogy is to design systems that are ultra-stable—using technologies the team knows and trusts—while reserving 10 % for bleeding-edge experiments. Juniors often disagree; seniors rarely do. Juniors get excited about every shiny tool without proof of its durability (a la the Lindy effect).
The barbell strategy doesn’t encourage timidity; it limits the cost of being wrong. I want to survive—and to do so wisely.
I’m not chasing hyper-growth;
I’m chasing sustainable growth,
in both
business
and
life.
### “I'm Not a Hypocrite.”
- URL: https://mohammadshaker.com/en/blog/im-not-a-hypocrite
- Date: 2025-04-22T00:00:00.000Z
- Tags: personal, startups, aws, Human-Written
I should only ask people to do things I will do myself. If I'm asking someone to do something that I won't do myself, then I am a hypocrite. Plain and simple.
#### Content
I should only ask people to do things I will do myself. If I'm asking someone to do something that I won't do myself, then I am a hypocrite. Plain and simple.
The Golden Rule by Hammurabi says: "Do unto others as you would have them do unto you.”
The Egyptian, the Indian, the Greek, the Romans, and the Persian; they all share a version of this Golden Rule.
**The inverse, “If you don't want “X” done to you, don't do “X” to someone else”**, is the Silver Rule\*: “Do not treat others the way you would not like them to treat you.”
The Silver Rule translates well in business: Never ask someone to do something you won’t do yourself. Never do to others what you don’t want them to do to you. Do to others what you want them to do to you.
This served me well; especially working in small startups early in my career. I should only ask people to do things I will do myself. If I'm asking someone to do something that I won't do myself, then I am a hypocrite. Plain and simple.
When I started my first two startups, we were bootstrapped and we were resource scarce. The team was a small team of fresh grad across the board: engineers or (a) marketer. That’s what we could afford. They were juniors who had great potential. None of the engineers has worked with proper cloud provider before (like AWS) for instance (10 years ago)
What made this team a killer is the attitude. We launched our first full webapp and a full working backend in 3 months. We all showed our skin in the game. We were all hungry to learn and none was a BSer.
## Skin in the Game
I’m still extra proud of working with such a team. I was completely engrossed with the work with everyone in the company. I wasn’t a CEO or a CTO or a CFO or a CMO. I was all of these. I worked with everyone side by side. I had skin in the game and I showed it.
I did everything. I programmed the first backend service. I designed our app and the website UI. I fixed bugs myself. I created, designed and posted IG posts just like any marketer. I answered customers emails and I reached out to influencers. I (cold) emailed school heads just like any sales person.
I did everything I would ask a designer, an engineer, a marketer, or a sales rep to do. I showed them that I’m there for them. And they showed me their best.
It wasn’t all rosy. I made **every** managerial mistake in the book. I got angry, I put someone down on public, I was impatient when I should’ve and patient when I shouldn’t have.
Though, the good thing that I did was that I told them I’m wrong when I’m wrong. I had regular 1:1s and reverse 1:1 weekly with all of them. I kept everything open and everyone can speak out his mind.
I simply learned management on the spot. By practice. Bottom up. Not leading from the front nor from behind. But side by side.
### Thanks
To teach is to do. And no one would’ve trusted me in hard times if I didn’t do that.
And hard times we had - just like any startup: limited runway, wrong targeting and a pivot within 1 year.
In a 3-year period we lost none on the team. And I believe this is because we trusted each other in bad times the same as in the good times.
People need to trust us for them to work with us. If we show them that we understand their pains, they would understand that we’re real colleagues - authentic and trustworthy.
These two traits are rare,
extremely rare, nowadays.
No one likes to work with someone they don’t trust. No one. If you’re a manager who can’t code, none of the engineers can trust you. If you’re a head designer who can’t design literally everything (as Massimo Vignelli would say: “To Design is to Design Everything”), none of the designers will trust you. If you’re a founder who’s undecided, no employee will trust you.
Skin in the game that is. And in the game; skin, sweat, tears and blood will be spilled. And the people we work with will reward us for seeing us sweat with them.
Salam, peace.
\* Skin in the Game, Nassim Taleb.
### \"Requirement\" Is an Empty Word.
- URL: https://mohammadshaker.com/en/blog/requirement-is-an-empty-word
- Date: 2025-04-17T00:00:00.000Z
- Tags: software engineering, product management, user stories, Human-Written, Engineering
"Requirement" is one of the most overloaded words in software engineering — it sounds precise but isn't. It forces a solution frame too early and closes off better alternatives before the problem is fully understood. Replacing it with constraints, goals, and user stories keeps the conversation open longer and leads to better outcomes.
#### Content
The term "requirement" is overused and ambiguous in software engineering contexts.
When someone declares something a "requirement" to solve a problem, it creates an illusion of clarity while actually narrowing the solution space unnecessarily.
## The problem with "requirement"
Labeling something a "requirement" carries implicit weight. It signals certainty and non-negotiability — especially when pronounced by senior team members, who may inadvertently treat requirements as commandments that shouldn't be questioned.
## What to use instead
I advocate for replacing requirements language with **User Stories and Narratives**. This approach keeps the solution space open and flexible.
My team at SpatialX implements this practice across all Product and Technical tickets in Asana.
## The takeaway
Use "requirement" sparingly and only when genuinely certain something is truly necessary. Otherwise, embrace more collaborative language that invites discussion and alternative approaches.
### The Current Expert Problem with ChatGPT
- URL: https://mohammadshaker.com/en/blog/the-current-expert-problem-with-chatgpt
- Date: 2025-04-09T00:00:00.000Z
- Tags: personal, Human-Written, Engineering
ChatGPT is a generalist, not a sniper — it answers confidently even when it shouldn't. This is dangerous for junior engineers who copy-paste without review and for PMs who treat its output as final. Used correctly, as a junior team member who needs guidance and guardrails, it is genuinely useful.
#### Content
I was talking with my friend Mehdi Zonjy about the following, and I thought it was worth sharing.
ChatGPT’s ability to generate responses across a vast range of topics is both a strength and a limitation.
It’s not a sniper—precise and targeted—it’s a shotgun: broad and often generic.
And that’s a problem.
Take its confidence in incorrect answers. Right now, ChatGPT is like a junior developer who never says: “I don’t know.” Or “I am not sure the solution I am providing is robust enough.”
It will give you an answer with confidence—whether or not it should, and whether or not the answer is correct.
This is extremely risky for new engineers entering the field, who often copy and paste solutions without fully understanding what’s inside. For production work this is not valid.
That said, ChatGPT is still incredibly useful—if you treat it like a junior team member who needs review, context, and guardrails.
As a senior engineer, you still have to guide it: nudge it toward specific design patterns, ensure it uses OOP when appropriate, or point out that a Lambda function has memory and execution limits, making its proposed solution invalid.
In Product Management, ChatGPT is good at surface-level tasks: generating user stories, drafting PRDs, summarizing customer interviews. That’s helpful, and in many ways mirrors a Junior PM—someone learning the ropes and able to produce artifacts quickly with the right direction.
But it can’t replace the judgment of a Senior PM—the person making trade-offs between business value and technical complexity, who knows when to ship a rough cut versus when to polish.
For example, if you ask ChatGPT how to prioritize features, it might suggest RICE or MoSCoW. But it won’t challenge the underlying assumptions.
That’s a bummer. It won’t ask things like: Are we even building the right product? What’s the risk of not shipping this?
It can’t sense the tension between short-term growth and long-term strategy.
That requires context, lived experience, political awareness, and an instinct for timing—things machines don’t yet have.
The danger is when people start treating ChatGPT like a senior.
When it comes to deep expertise, a human—a senior—is still necessary.
That will change. And I hope it does.
### Weekly: 3 Biggest Problems (3BP)
- URL: https://mohammadshaker.com/en/blog/weekly-3-biggest-problems-3bp
- Date: 2025-03-26T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here's a digest of my week. Something I want to keep doing Gym 5-6 times/week. Focus on 1 thing in each day. Not thinking about anything else.
#### Content
Here's a digest of my week.
## **Something I want to keep doing**
- Gym 5-6 times/week.
- Focus on 1 thing in each day. Not thinking about anything else. I do plan the next day at the end of each day.
- Sitting and Writing.
### **Something I'm learning.**
- Technical: How Figma does sync its events between the frontend and the backend. We're _not_ using something similar in SpatialX.
- Technical again. VectorDB.
### **Something I’m remembering from the past**
My father, a lot.
### **My current exercise regimen**
- 120 steps (on a stepper) in 20 minutes.
- 40-60 min weights training.
- No runs sadly this Ramadan.
### **My current day routine regimen**
- Up.
- Read for an hour.
- Work till 6.
- Gym.
- Iftar and sit with the family.
- Reading/working for 2hrs.
- Sleep.
### **What's working for you that you'd be crazy to change?**
- Focusing on 1 thing at a time.
- Asking myself what are the 3 biggest problems (I call it 3BP) that I want to solve: whether this is daily or weekly.
- Planning for the week ahead. Contrary to what most people think, [the idea is not to fill the calendar, but to clear the calendar when I plan.](http://mohammadshaker.com/2024/06/07/inverse-productivity/)
- Solve the biggest problem first.
### **What's not working for you, and you'd be crazy not to change?**
- Not reading enough
- Not writing enough
- Not sharing technical work and the great things we're doing technically at [SpatialX](http://www.spatialx.ai).
### **What makes me waking up enthusiastic lately? What can I do this week or month around this?**
- Our work at SpatialX really fires me up. Seeing it come together really makes one proud. Thanks to every single person in our small team.
That's it for the week.
### On Trusting.
- URL: https://mohammadshaker.com/en/blog/on-trust
- Date: 2025-03-19T00:00:00.000Z
- Tags: personal, startups, Human-Written
Writing this reminds me of when I started working on my own two startups seven years ago. I struggled to delegate work.
#### Content
Writing this reminds me of when I started working on my own two startups seven years ago. I struggled to delegate work. I wanted everything done to my standards—my quality, my ethics. In my mind, I was the best at what I did—“no one does it like I do.”
I think many capable engineers go through this. We start with zero trust in others. Slowly, we learn to trust people to do their jobs as well as we do—sometimes even better.
It’s a mindset shift—realizing that others bring different approaches and that we can actually learn from them. But the truth is, exceptional engineers are rare. The more we demand excellence from ourselves and others, the harder it becomes to trust that anyone else will meet that standard.
Trust isn’t given—it’s earned. No one deserves a free pass because of their title or past experience.
Trust is earned,
every time,
anew,
in every new venture we create or join.
Salam, peace.
### On Self Worth.
- URL: https://mohammadshaker.com/en/blog/on-self-worth
- Date: 2025-03-12T00:00:00.000Z
- Tags: personal, Human-Written, Life
Waiting for outside approval is one of the worst thing we can tie ourselves into.
#### Content
Waiting for outside approval is one of the worst thing we can tie ourselves into. We are weak if we want/strive/look-for outside validation, approval or appraisals.
If we do not have the self-worth from the inside, no outside validation or praise will suffice. If we simply follow the herd in “Fake it till you make it”, we are simply cheating ourselves.
“Don’t fool yourself and you’re the easiest man to fool.”
We are the ones who decides our self worth. We should be the ones who self-question, self-criticize (Popper), self-validate, and self-improve.
Most people tie in the idea of effort to the idea of self worth, praise or validation. In the \[imaginary\] realm of schools and universities, if we study well, we’ll pass. We’ll succeed. We’ll praised. In the real world, and in business, if we work hard, if we spend a lot, we may not succeed.
But that should not undermine our self worth.
Success or failure is not what determines our self worth. We do. Trying, doing, acting, tinkering is honorable.
Salam, peace.
### No SDETs and No QA Engineers.
- URL: https://mohammadshaker.com/en/blog/no-sdets-and-no-qa-engineers
- Date: 2025-03-06T00:00:00.000Z
- Tags: personal, AI, startups, Human-Written, Engineering
Never hiring QA engineers or SDETs is one of the best decisions I’ve made across multiple startups. With TDD, CI/CD pipelines, and AI-assisted testing, developers should own quality directly. Dedicated test roles create slow feedback loops, diffuse accountability, and signal that the team doesn’t fully trust its own engineers.
#### Content
One of the best decisions I’ve made across multiple startups is **never hiring QA engineers or SDETs** (Software Development Engineers in Test).
A hard rule, but always the right one.
**Before the Pushback: Two Important Notes**
1\. I’m not saying SDETs are _never_ valuable. Some of the SDETs I’ve worked with are excellent. But for most startups and companies, they are unnecessary. And sometimes _harmful_.
2\. The need for lean engineering teams is more evident than ever, thanks to:
- Proper engineering practices (DevOps, TDD, Automated CI/CD, UI testing, etc.).
- AI tools that enhance software quality (though not necessarily the best code quality—whatever “best” means for a company).
With that out of the way, let’s dig in.
**Who’s Responsible When Something Breaks in Production?**
If a bug appears in production and you have both engineers and QA testers, who’s responsible?
- With engineers only, the answer is simple: engineers.
- With engineers and testers, responsibility becomes shared—which means no one is fully accountable.
If we trust engineers to build quality software, why do we need SDETs? Hiring SDETs suggests a lack of trust in our engineers.
Ultimately, writing, testing, and ensuring code quality is the engineer’s job — not the QA engineer’s, not the SDET’s.
Sure, in highly regulated industries or complex release workflows, extra testing is necessary. But for most startups, this isn’t the case.
**The 2nd Order Effect of Having SDETs & QA Engineers**
Bringing in SDETs fundamentally alters the entire software development lifecycle (SDLC) and team dynamics across engineering, product, design, and even sales. Take the following problems, viewed from the engineers perspective:
**1\. “It’s Not My Problem” Mentality**
**Without SDETs:** If a bug hits production, **it’s the engineer’s responsibility**—no excuses.
**With SDETs:** "_Not my problem. The SDET didn’t catch it either._"
**2\. No End-to-End Ownership**
**Without SDETs:** Engineers **own** the full cycle—from ideation to deployment.
**With SDETs:** "_I don’t need to think about the feature end to end. Someone else will check my work later._"
This mindset kills innovation, accountability, and attention to detail (aka, craftsmanship.)
**3\. No Involvement in Releases**
**Without SDETs:** Engineers care about **how and when** things go live.
**With SDETs:** "_My work will be released by someone else—SDET, DevOps, or whoever. I just write code._"
This breaks the connection between engineers and the impact of their work.
**4\. Recurring Rollout Problems**
**Without SDETs:** Engineers take ownership of feature rollouts.
**With SDETs:** "_We keep running into deployment issues, but that’s not my concern. Someone else handles it._"
Firefighting never ends, and engineers never truly understand why things fail.
**5\. Best Practices Get Ignored**
**Without SDETs:** Engineers rely on **TDD, automation, and best practices**.
**With SDETs:** "_TDD slows us down. We don’t have proper tests anyway. Manual QA is faster!_"
The result? A fragile, unstable codebase.
**6\. No Dogfooding or Feature Flagging**
**Without SDETs:** Engineers **test their own features in real-world conditions**.
**With SDETs:** "_Dogfooding is for product managers. Feature flags are complicated and expensive._"
(This is one of the worst arguments I’ve heard from engineers and managers—fire them immediately.)
**A Simple Change. A Big Impact.**
At one startup, I eliminated all SDETs within two months. The result? Software quality improved. And Product and cross-functional collaboration (XFN) became stronger.
With SDETs gone, engineers owned their work from start to finish. Now, they had to communicate with sales, product, and design to understand why, how, and when features were built.
**What Changed?**
• TDD became a no-brainer. Engineers now saw its value in protecting their own code.
• Feature flags became essential—allowing engineers to test safely in any environment, _including prod._
• Engineers could test on production with product managers, enabling faster and safer releases.
• Dogfooding and gradual rollouts became the norm—ensuring the right user experience before launch.
**Final Thoughts**
By removing SDETs and QA engineers, we fix ownership, accountability, and software quality. Our engineers become product-driven, responsible, and independent.
And most importantly, they _care_. They _own_.
Salam, peace.
**Summary of the differences in a table format**
Below is a concise table summarizing the key differences and consequences when SDETs/QA engineers are present vs. when they are removed, based on the scenarios described:
| Area | With dedicated SDETs/QA | Without dedicated SDETs/QA | Consequence |
| --- | --- | --- | --- |
| Production bugs | Responsibility is shared between engineering and test roles | Engineers own the defect and its prevention | Accountability stays with the people who design and build the system |
| End-to-end ownership | Engineers can hand work to a later testing stage | Engineers follow the feature from idea through deployment | Product context and craftsmanship stay connected to implementation |
| Releases | A separate role may coordinate validation and launch | Engineers participate directly in rollout decisions | Delivery feedback reaches the builders faster |
| Rollout failures | Deployment problems can be treated as another team's concern | Engineers own rollout safety and recovery | Teams learn from operational failures instead of repeatedly handing them off |
| Engineering practices | Manual QA can substitute for automated checks | TDD, CI/CD, and automated UI tests become part of normal development | Quality controls run continuously and remain close to the code |
| Production learning | Dogfooding and feature flags can be delegated or skipped | Engineers use flags, gradual rollout, and dogfooding with product partners | Real-user feedback arrives earlier with bounded release risk |
### Weekly: My Father. And Hot Water.
- URL: https://mohammadshaker.com/en/blog/dad-and-hot-water
- Date: 2025-02-26T00:00:00.000Z
- Tags: Brain dump, Human-Written
Something I keep doing In the last 3 months I've been doing: Sleeping 8hours have a tremendous effect on my mood. Recently I'm making this non negotiable.
#### Content
## **Something I keep doing**
In the last 3 months I've been doing:
- Sleeping 8+ hours have a tremendous effect on my mood. Recently I'm making this non negotiable.
- Long walks for thinking, especially in the summer. Now it's winter and I'm trying to grab sunny days for long walks or traveling between places.
- Gym 5-6 times a week. 4 sessions weight training, w sessions as zone 2/VO2 max.
- Reading 1h+ a day (while walking, while in the gym and before bed.)
### **Best of what I watched**
Abeer Nehmeh Opera-style voice with this poem is magical.
### **Best of what I listened to**
**What I'm reading?**
I'm reading a mix of 2-3 books of the following in a day. The ratings below are mine and are WIP. **Ways of Seeing, Berger** is the book that is different from all others. A definite read.
- Understanding understanding, Richard Saul Wurman. Physical copy. Rating: 6/10.

- Ways of Seeing, Berger. Physical copy. Rating: 8/10.

- DK History: Illustrated. Physical copy. Rating: 6/10.

- DK Simply Explained: History. Digital copy. Rating: 6/10.

- DK Simply Explained: Economics. Digital copy. Rating: 6/10.

- Superconnect, Koch. Rating: 6/10.

### **Something I'm learning**
Given the list of books I'm reading, I'm learning a lot on Economics, Finance, Philosophy and design.
### Something new to try
On Sat and Sun I'm doing gym first thing in the morning before enjoying my day. This removes the guilt when having all the cheat food in these two days. [Fit with my mental model of **_sweating for it before I earn my food._**](https://mohammadshaker.com/2023/09/25/sweat-for-it-i-must/)
### Something I’m remembering from the past
My father anniversary is coming up next week.
I remember, my father and I, in the bath. And how he is teaching me how to clean myself. I was around 6 maybe. I remember the **_very_** hot water he would pour on top of my head with a yellow copper bowl (tasse in French, طاسة in Arabic.) Close, but never as fancy as this:

### 80/20 of my time for the week I went on..
Simply, 80% of my time went to:
- Reading
- Writing. A lot (documentation, not the enjoyable type of writing.)
Another way of looking at it is that 80% of the outcome was based on:
- Thinking session which took 4 hours.
- Working on items out of that session, which took another 6 hours.
So this is again 80% of outcome from around 20% of the effort (the time available in a week.)
Ramadan in coming soon, so Ramadan Mubarak.
RIP, my father.
### Sweat for It, I Must.
- URL: https://mohammadshaker.com/en/blog/sweat-for-it-i-must
- Date: 2025-02-19T00:00:00.000Z
- Tags: Brain dump, Human-Written
My best buy of this week was a Venchi icecream I had near Piccadilly in London.
#### Content
My best buy of this week was a Venchi icecream I had near Piccadilly in London. This was after I did my work for the day, went to the gym, and walked for 45 min. It's worth it. (Side note: if you are in London, [Darlish](https://darlish.com/) is the place to go if you like Pistachio icecream. Am a fanatic when it comes to anything \[Aleppo\] pistachio, mediterranean olive oil, mountain honey, figs and middle eastern berries.)
I long have this problem: I can't have, nor enjoy pleasures (even food) before I do the work.
This started back in university years where I would stay up to 2AM each day, after a full day fast, without eating. I would only eat around 2:15 AM after I finish everything I wanted for that day. I would only enjoy the pleasure of my 1 meal a day at 2:15 AM. (I was also teaching summer courses for junior university students and that what made summers days ever longer trying to prep for courses I teach.)
I don't think it's a problem though. I think it's healthy to say that you have to **_sweat_** for it, in order to earn your food (in Arabic Aiesh, عيش, which means "earn the living.") That's literarily what our ancestors have done when they chase animals and **_sweat_** for their food.
Sweat for it, I **must**.
Salam, peace.
### 23 Things I Learned in 2024.
- URL: https://mohammadshaker.com/en/blog/things-i-learned-in-2024
- Date: 2025-02-13T00:00:00.000Z
- Tags: Brain dump, Human-Written
Twenty-three hard-won lessons from 2024: do the hard thing first, focus on one thing for four hours before touching anything else, be stubborn about truth but weak in holding it, and always put yourself in rooms where you are the least capable person. Reflection without action is just noise.
#### Content
Reflection is one of the best tools that taught me to stay still, think and write.
Here are **Some Things I Learned in 2024**:
1\. Be involved in doing hard stuff, grand stuff, for an extended period of time. Do a hard thing every day. Do the hard things first thing in the morning. I feel good every day I do that.
2\. Be known for doing hard things. Write about them and chase invalidation (Popper’s fallibilism).
3\. Focus on one thing at a time, then go full throttle on that one thing. Do this daily for at least 4 hours.
4\. I should always mark my own path. Don’t rely on anyone. Be self-sufficient all the time.
5\. Am I less or more _libre_? I should always be more free with time, a _libre_. I should always own my own destiny.
6\. **Apprenticeship:** Always put myself in places where I can learn from a group of people way better than myself. Be where the Pros are.
7\. Be stubborn as much as I like, but self-critique and self-reflect. I should always look for truth (with Popperian fallibilism and Talebian P (Probability) vs. E (Payoff) understanding. More on this in another post).
8\. Be strongly founded, with strong views backed by strong diligent work that are weakly held. I can be proven wrong. Most people have weakly founded, wrong views that are strongly held. Debunk anyone doing this. Demand that from others, especially in business.
9\. Observe myself from a third-person POV, reflect, and say sorry when I should. _“Don’t fool yourself, and you’re the easiest man to fool.”_
10\. I’m the master of my mind. Do what I think is right. I planned to run a marathon in 2024 when I was running 0 KM in January with a problematic knee. I knew I couldn’t and shouldn’t do it. Doctors told me not to run. I ran 17 km in November. I could never have imagined doing this in January with the knee I had.
11\. Don’t waste my time on tatters and the small stuff. Avoid the negatives, and the positives grow by themselves (_via-negativa_). If I ditch social media and my phone, I naturally read more, have deeper and longer conversations with people, and enjoy dinner with family, etc.
12\. **Produce > Consume.** Keep this in mind all the time. Excessive reading without reflection and action causes atrophy in the brain. Think, retrospect, future-gaze, dream big, and then act. Don’t do one side without the other.
13\. Don’t try to milk every inch of everything. Some things should and must NOT be efficient (e.g., time with family, holidays, etc.).
14\. **Mind the incentives of people.** Doctors give us medicine even if we are well without it. That’s what they are paid to do.
15\. **Minimize risk for downsides:** Knowing what I should avoid is as important as knowing what I should chase (NNT).
16\. **Maximize risk for unbounded upsides:** Know that good things happen when I keep my mind open and free. _“You’ve got to be very careful if you don’t know where you’re going because you might not get there.”_ — Yogi Berra.
17\. **Follow the Lindy:** Things that have proven their fitness. In books, these are the classics (e.g., works of Aquinas, Marx, Russell, Popper, Galileo, Darwin, Imam Ali, Montaigne, etc.).
18\. **Go wide with specific deeps (T or V Shape).** Read from multiple disciplines but go deep in some. Be encyclopedic. Learn and read from adjacent domains at the same time. And read multiple of them together. Reading Russell, Popper, Hayek, Taleb (for a 5th time), and Wittgenstein at the same time actually changed my mind.
19\. **Days for time off (really off—without a phone or notifications, just a pen and paper) every 2-3 months actually change my perspective.**
20\. I can actually read while I walk. (I don’t enjoy Audible that much since I can’t highlight. I replaced that with the _Spoken Content_ feature on iPhone to read any iBook for me.)
21\. **Motion creates emotions.** That’s absolutely true. I feel amazing after every run (-1 run in 2024). If I’m blocked or down, I move, and I’ll feel good. Creativity spurs in motion. My best technical or product solutions came after a 20+ min walk.
22\. Spending time with loved family and close friends is a **bliss**.
23\. **Globalism as a way of thinking is for the average.** Most people are being molded into a global way of thinking. This is a rut I should always fight.
Salam, peace.
### I’m a Libre. I Hate Dependencies.
- URL: https://mohammadshaker.com/en/blog/im-a-libre-i-hate-dependencies
- Date: 2025-02-05T00:00:00.000Z
- Tags: technology, Human-Written, Engineering
All misery stems from dependence—whether in personal life or business. To be free, truly free, we must be libre.
#### Content
All misery stems from dependence—whether in personal life or business. To be free, truly free, we must be **libre**.
The word _free_ originates from the French _libre_, which itself comes from Latin. However, _libre_ carries a deeper meaning than just freedom—it represents liberty, not price.
For example, _libre_ software is free to use and free of restrictions, but not necessarily free of charge. It’s unfortunate that we often conflate the two concepts.
True freedom lies in eliminating dependencies. To be _libre_ is to be independent.
If I rely on a single income source, a specific technology, a particular provider, or even a single company, I will never be truly free. Dependencies restrict control over my own destiny.
The only way to attain deep, lasting freedom is to systematically break dependencies. Let me explore how this principle applies across different areas of life.
## Personal Freedom
At an individual level, it is possible to minimize dependencies. But complete independence is unrealistic and undesirable. Relationships, friendships, and communities provide meaning and support. However, we must avoid dependencies that strip away optionality. Clinging to a single option in life—whether a career path, or a financial stream—limits choices and creates fragility.
Partners are different. My wife reads this. (And I love her.)
### Business Independence
In business, reducing dependencies is crucial. A company should not be overly reliant on specific employees, technologies, or customers. **Optionality is key**:
- No single employee should be indispensable.
- No single technology should be irreplaceable.
- No single customer should hold disproportionate leverage over the company’s survival.
A business must be resilient by maintaining flexibility—ensuring that if one element fails, the entire system does not collapse.
### National Sovereignty
Independence is power. A country is only as strong as its ability to operate without a crippling reliance on others. Economic and geopolitical independence ensures stability. A nation with few alternatives is a fragile one.
### The Role of Optionality
If being _libre_ means eliminating dependencies, then **optionality** is the key to achieving it. The freedom to choose is the foundation of true independence.
This concept aligns with Milton Friedman’s idea that a free society must allow individuals to make their own choices. By maximizing optionality, we create an anti-fragile system that can adapt and thrive under uncertainty.
### Practical Applications
#### For Individual Contributors (ICs)
- Develop a broad skillset with deep expertise in areas of genuine interest (T-shaped skills).
- Avoid over-specialization in a single tool, framework, or industry.
- Ensure your value is transferable across different domains and market conditions.
#### For Managers
- Build teams that are not reliant on a single individual or technology.
- Hire generalists where possible—full-stack developers instead of ultra-specialized roles (unless the team is large enough to support specialization).
- Implement redundancy: Ensure that if one person leaves, the team remains functional.
- Apply resilience principles from engineering:
- Fault tolerance (handling failure without collapse)
- Risk mitigation (proactively addressing weaknesses)
- Redundancy (e.g., RAID storage systems)
- Graceful degradation (ensuring systems function at reduced capacity if necessary)
- Modularity (designing components that can be swapped or adapted)
- Introduce controlled chaos: Rotate team members, mix disciplines, and encourage flexibility to prevent single points of failure.
#### For Companies and Startups
- Hire adaptable employees who can take on multiple roles.
- Avoid relying on a single major customer (no 80/20 revenue dependency).
- Diversify product offerings to maintain flexibility.
- Ensure no single VP or executive holds too much control—organizational fragility is a risk.
### Conclusion
Freedom is not about rejecting all attachments but about **avoiding over-dependence**. Whether as individuals, companies, or nations, we must cultivate optionality to remain truly _libre_.
Salam, peace.
### Recommended Readings
For deeper insights on being _libre_, consider these books:
- _Fooled by Randomness_ – Nassim Nicholas Taleb
- _The Black Swan_ – Nassim Nicholas Taleb
- _The Open Society and Its Enemies_ – Karl Popper
- _The Law_ – Frédéric Bastiat
- _On Liberty_ – John Stuart Mill
- _Safe Haven_ – Mark Spitznagel
- _Human Action_ – Ludwig von Mises
- Any book by F. A. Hayek
### For the Love of Old Books
- URL: https://mohammadshaker.com/en/blog/for-the-love-of-old-books
- Date: 2025-01-21T00:00:00.000Z
- Tags: personal, Human-Written
Old book are normally called the timeless - what marketers call nowadays: the evergreen.
#### Content
Old book are normally called the timeless - what marketers call nowadays: the evergreen.
Like in Natural Selection, old books are books that proved their _fitness._ They proved their _ability to survive._
For me, and in no particular order, nor exhaustive, these books are:
- It would be the [The Mythical Man-Month](https://en.wikipedia.org/wiki/The_Mythical_Man-Month) or Cormen’s [Algorithms](https://www.amazon.co.uk/Introduction-Algorithms-Thomas-H-Cormen/dp/0262033844) or anything from [Paul Strassmann](https://www.google.com/search?sca_esv=4fbe7c869920a6f3&rlz=1C5CHFA_enGB883GB883&sxsrf=ADLYWIJdjHgrZ1lb2QCybdWUohEd20qJDw:1736767231714&q=The+business+value+of+computers+Paul+A.+Strassmann&stick=H4sIAAAAAAAAAONgFuLSz9U3MK0oLq9KUQKzzczN0kwrtaSyk630k_Lzs_UTS0sy8ousQOxihfy8nMpFrEYhGakKSaXFmXmpxcUKZYk5pakK-WkKyfm5BaUlqUXFCgGJpTkKjnoKwSVFicXFuYl5eTtYGQHnhs5obgAAAA&sa=X&sqi=2&ved=2ahUKEwiFmJ2PyvKKAxXG0wIHHaDOAcsQgOQBegQINxAG&biw=1512&bih=749&dpr=2).
- It would be Bruno Munari’s books.
- It would be the 60 pages book of Vignelli’s _Cannon_ or his \*Design is One. (\*Vignelli taught me how to look at any space, whether its a paper, an app or a building and know how to play with it.)
- It would be _any_ book from \*\*Edward Tufte (I can’t thank Tufte enough to release me from all the dogma of vulgar designs, anywhere.)

- It would be The Minard System or Brockmann’s Grid Systems.
- It would be Cicero or Cato on politics, humanity, and virtues.
- It would be anything from Bastiat, Hayek, Von Mises and Hazlitt in economics.
- It would be Russell and Wittgenstein in thinking, philosophy and history.
- It would be Howard Zinn’s books, Marx and Adam Smith.
- It would be Karl Popper for falsification and questioning myself in every belief.
- It would be anything from Imam Ali or Nassim Nicholas Taleb. They both taught me on being righteous, to always look for truth, to think for myself, to be unwavering on what I stand for.
- It would be ..
I should read a timeless everyday.
Salam, peace.
### Week 44: I Lose, Reality Wins.
- URL: https://mohammadshaker.com/en/blog/i-lose-reality-wins
- Date: 2024-09-26T00:00:00.000Z
- Tags: quick-read, startups, Human-Written
The meaning of success in school and university is the antithesis of reality. If we compare what success means in school or university VS real life at
#### Content
The meaning of success in school and university is the antithesis of reality. If we compare what success means in school or university VS real life at work, it’s completely wrong.
In school and universities:
> Effort == Success
This doesn’t translate well at all to business (especially startups) where:
> Effort != Success.
Effort doesn’t correlate one-to-one to success. You may work for a year. And no one buys your product or you are net negative or anything but success. Viola, Effort = 100%. Success = 0.
Real life punches you in the face. Customers punch you in the face. Reality wins, you lose. Reality won, I lost - 2 times before.
Metrics don’t move because we spent 30 days and 30K of employees salaries on an experiment. Metrics move because we provided a service people want. Until we find that formula, reality will always win.
### Week 43: You Can't Be Surprised Twice.
- URL: https://mohammadshaker.com/en/blog/week-43-you-cant-be-surprised-twice
- Date: 2024-08-15T00:00:00.000Z
- Tags: Brain dump, Human-Written
This is part 4 of answering a list of question. Part 1 is here. Part 2 is here. Part 3 is here.
#### Content

This is part 4 of answering a list of question. [Part 1 is here.](https://wordpress.com/post/mohammadshaker.com/4721) [Part 2 is here.](http://mohammadshaker.com/2024/06/06/r-not-all-weeks-are-created-equal/) [Part 3 is here.](http://mohammadshaker.com/2024/04/30/ready-week-yy-guesy-questions-part-3/)
You can find the full list of questions [here](https://guzey.com/questions/).
1. why did the japanese keep debating whether to surrender even after two nuclear bombings?
1. [Because it matters.](http://mohammadshaker.com/2024/06/06/r-giving-up/)
2. whose advice has been so correct in the past that you now just go and act on it?
1. My grandfather. Never work for anyone.
2. Nassim Taleb. Employment salary is the new age slavery.
3. why not ask them for a piece of advice?
- I should. And I always do.
4. is the failure condition clear?
- There’s this notion of always define your failure condition before getting into something. Like defining when you’ll close a business before starting a business and by when.
5. when was the last time you talked to someone you really admire?
1. My sister, Nuha.
6. what are you going to regret in a year?
1. I’ve just made a decision today so that I won’t regret it in a year.
2. A better questions: What are you NOT doing today, that you’ll regret in a year?
7. have you spent 1 minute today on your biggest goal?
- Yes.
8. have you spent 1 hour today on your biggest goal?
- Yes.
9. have you spent 10 hours today on your biggest goal?
- No.
10. [how to prepare for the times of change?](https://marginalrevolution.com/marginalrevolution/2023/03/existential-risk-and-the-turn-in-human-history.html)
1. Antifragile.
2. Stoicism and Seneca.
11. what’s the most common advice you give to people?
- Learn by mistakes. Learn anything by having a skin in the game. If you want to start a startup, do it with your money, develop it yourself first, etc.
12. do you follow it yourself?
- Yes I did. I bootstrapped two startups myself.
13. does the world need to be saved?
- If we don’t fight for our freedom, no one would give it. The world need to be saved from the greedy of us. Again. [What matters, matters.](http://mohammadshaker.com/2024/06/06/r-giving-up/)
14. what can you do in the next 60 seconds that will make you feel proud of yourself?
- Air squat for 30 times.
15. what are you wrong about?
- Many things - as I should be. Unless I try things out, I won’t be wrong about anything. It’s wrong not be wrong.
16. who’s consistently ahead of you?
- My sister, Noor.
17. why don’t you catch up?
- Each of us have his own path in life that he should walk. She’s ahead in her path. I hope am on track on mine.
18. how to overcome the second law of thermodynamics?
1. NA
19. how did zuck make a comeback?
- From the back.
20. what do you wish you could do but you’re absolutely confident you can’t?
- The question is stating something as a fact which I don’t agree with. That’s a self-destructing belief that I won’t ever have myself. There’s always hope. And Hope (Amal in Arabic) is my mother name.
21. how do you know?
- I’ve seen it happening before. I’ve seen it in real life. In practice, not in theory. Read Popper.
22. when was the last time you closed your eyes for a minute and asked yourself if you’re doing the right thing?
- Every week in the last 12 months.
23. what can you do all day long and feel great going to bed?
1. Exercise.
2. Read.
3. Building something.
24. do you like yourself when you look in the mirror?
- I’m good. Alhamdullilah.
25. does anyone believe they’re evil?
- Everyone does.
26. what do other people tell you about you that you are always surprised to hear?
- I can’t be surprised about one thing more than once, right? That’s an illogical question.
That's it. Questions are answered.
Salam, peace.
### Week 42: A Prayer, by Definition.
- URL: https://mohammadshaker.com/en/blog/week-42-prayer-by-definition
- Date: 2024-08-08T00:00:00.000Z
- Tags: quick-read, Human-Written
This is part 3 of answering a list of question (the full list is here.)
#### Content
This is part 3 of answering a list of question (the full list is [here](https://guzey.com/questions/).)
[Part 1 is here.](https://wordpress.com/post/mohammadshaker.com/4721) [Part 2 is here.](http://mohammadshaker.com/2024/06/06/r-not-all-weeks-are-created-equal/)
1. Why is it easier to do something for the other person than to do the exact same thing for yourself?
- Because we sometimes like the other person more that we like ourselves.
2. did the person actually tell you no or did you just assume they would?
- That’s the first good question here. At last.
3. who do you want to be less like?
- Hollow souls and trendy people.
- A son of a millionaire. I'm gratefully not.
4. what would you do if you were alone in the universe?
- Build a ladder to get out.
5. would you walk into a teleporter?
- Enough of teleportation already.
6. what makes you do the right thing?
- 10 years ago, I would say that I do the right thing because I follow my faith, the virtues and values I learned from parents, and customs of where I originated from.
- I respect all of the above (ala David Hume) but this is not completely true now. Now it's my _thinking_ mind first and formost.
7. when was the last time you prayed?
- Every day. But that is wrong. Since I'm not actually doing prayer correctly. And a prayer is a prayer done correctly - by its own definition (ala Wittgenstein.)
8. what are you proud of?
- All of my 3 sisters. And my mother. And my father. They all are better than myself in every aspect.
9. when was the last time you did something you’ve never done before?
- If it counts, taking a new way home, last week.
10. do you feel better or worse after spending time with the person you spend most of your time with?
- Nothing feels better than time spent with family or close friends. I love my family, close and far. And I love my close friends. We’re social beings.
11. what are you afraid to tell to your best friend?
- Hard choices, easy life. Easy choices, hard life.
12. why did you do something you regret?
- I let the heart lead the way, not the mind. I shouldn’t. Hard decisions, easy life. Easy decisions, hard life.
13. why not travel to your favorite city as often as you can?
1. Because I should try another city that may be my next favorite city.
14. am i wasting my time writing these questions?
- Yes. Because after writing my answers this far, I think they are mediocre questions without a connecting theme.
15. what have you promised to do and not done yet?
- I’ve promised myself financial freedom which I haven’t done yet. But I’m changing this now.
### Week 41: What Matters, Matters.
- URL: https://mohammadshaker.com/en/blog/week-41-what-matters-matters
- Date: 2024-08-01T00:00:00.000Z
- Tags: quick-read, Human-Written
We’re built into the notion of “never give up.” But only the unwise do NOT. Like myself for quite some time.
#### Content

We’re built into the notion of “never give up.” But only the unwise **do NOT**. Like myself for quite some time.
**Not giving up on what do not matters pushes us away from what actually does.**
If what matters no longer matters, then giving up is what matters. If a business doesn’t make sense and losing money because no one is buying. Stop. Do I want to be right or do I want to be stubborn?
On the other hand, if my relationship with a friend is going south, I’ll never give up. Because it matters and it will always matter.
What matters, matters. What doesn't, doesn't.
Salam, peace.
### Week 40: God and Fear.
- URL: https://mohammadshaker.com/en/blog/week-40-god-and-fear
- Date: 2024-07-25T00:00:00.000Z
- Tags: quick-read, Human-Written
What am I afraid of? No one comes to mind. People of my faith would tell me that I should be afraid of him, God with a capital G.
#### Content
What am I afraid of? No one comes to mind.
People of my faith would tell me that I should be afraid of him, God - with a capital G. But I’m not sure if that is what **the** God wants.
That where it ends this week.
Salam, peace.

## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Week 39: Not All Weeks Are Created Equal.
- URL: https://mohammadshaker.com/en/blog/week-39-not-all-weeks-are-created-equal
- Date: 2024-07-18T00:00:00.000Z
- Tags: Brain dump, Human-Written
This is part 2 of answering a list of question. Part 1 is here.
#### Content
This is part 2 of answering a list of question. [Part 1 is here.](https://wordpress.com/post/mohammadshaker.com/4721)
You can find the full list of questions [here](https://guzey.com/questions/).
1. **Why is the market capitalization of facebook, founded 15 months prior to yc by a 19 year old, larger than that of all yc companies combined?**
1. That's not the smartest question I've read recently.
2. **Why not take a pen and a piece of paper, go offline for 1 hour and resolve one of them?**
I do. And I’ve always loved the analog version outside the digital world. [Look at my notebook](https://mohammadshaker.com/2023/07/30/week-16-journalling/).
3. **What are the key questions?**
- Righteousness: Do I think what I’m doing is the right thing? I should ask this for every action I do.
- Growth Mindset: Where can I grow the most? Even if that means I’ll have to do things I'm not comfortable with or face my fears.
- Course-correct in smaller scale. I shouldn't wait for things to go completely sideways. Daily and weekly checkins are tools for that. [That's why I journal](https://mohammadshaker.com/2023/07/30/week-16-journalling/).
- Am I following the herd or the values that I believe in? I’m a muslim. And I believe I should do good in the world. That’s essential for me and for my identity. Do I follow my own virtues and values? Am I a hypocrite asking someone else to follow virtues, values, or customs I don’t follow my self? Daily and weekly assessments are essential to correct my faults. That’s why I try to [journal daily or at least once or twice weekly.](https://mohammadshaker.com/2023/07/30/week-16-journalling/)
4. **What do you want?**
- Do what I believe is the right thing to do. This is harder to do in practice that in theory. Some can do it. I can see my sister, Noor, has done it. And she is an inspiration.
5. **Is self eternal?**
- I can obviously not say. Since I can’t remember what’s been there before I was born (but who said it is), nor expect to know what will be there in the future.
6. **What’s slowing you down the most?**
- Staying up late.
7. **When was the last time you took 3 months completely off?**
1. Never.
8. **What’s your destiny?**
1. I can build
9. [**Is 1 week 2% of the year?**](https://nat.org/)
1. No. Not all weeks are created equal.

### Week 38: Outsmarting Customers
- URL: https://mohammadshaker.com/en/blog/week-38-outsmarting-customers
- Date: 2024-07-11T00:00:00.000Z
- Tags: quick-read, startups, Human-Written
This is one of the most things I hate: Outsmarting customers.
#### Content
This is **one of the most things I hate**: Outsmarting customers.
In pre product-market-fit (pre-PMF), it’s always good to ship something in 2 weeks, test with real users and let them tell you what’s wrong/right than waiting for a full cycle of 1-2 month of conducting research, designing and iterating 5 different versions, and shipping 1 at the end.
Nothing beats a user experience delivered in a user hand on a real device with a prod app. This is always turned out to be true VS trying to _**outsmart**_ the users - trying to design _**the**_ perfect feature. I’ve seen it failing again and again with startups.
Know when to get into action - I love people with **_a bias for action_**. And I absolutely detest the approach of “we should iterate on this for another week.” I don't buy "slow is smooth and smooth is fast." And again, especially for products pre-PMF.
Until a startup reaches PMF, stop trying to be smart. Get your PMF and do whatever you want. Everything changes then.

### Week 37: What Does Good Future Look Like?
- URL: https://mohammadshaker.com/en/blog/what-does-good-future-look-like
- Date: 2024-07-04T00:00:00.000Z
- Tags: quick-read, Human-Written
Another question I'm trying to answer. What does good future look like?
#### Content
Another question I'm trying to answer. **What does good future look like?**
It's very open ended. But here are top of mind - in no particular order.
A good future for me is:
- Great health.
- A loving wife.
- And a loving family.
- Close, few and deep relationships.
- And a wide network.
- Friends to laugh with.
- Worry-free sleep.
- Financial freedom.
- A hot beverage and book in a garden every morning.
- Slow prayers (what other people call meditations is equivalent to prayers in Islam.)
- A muslim family with kids who are adamant on doing good in the world and lead a purposeful life for themselves and others around them.
- Closer to my sisters (family) and my extended family.
- A free calendar.
- Doing good on daily basis, for family, friends, colleagues and customers.
- Doing things on purpose. Not by consensus.
- A Porsche (I'm human after all.)
- Runs under the sun.
- Fresh air every morning - outside cities.
Salam, peace.
### Week 36: Selfishness and Spirituality
- URL: https://mohammadshaker.com/en/blog/week-36-selfishness-and-spirituality
- Date: 2024-06-27T00:00:00.000Z
- Tags: quick-read, Human-Written
Another question I had to answer. What's my disteny?
#### Content
Another question I had to answer. **What's my disteny?**
I can build up to the destiny I want. But I cannot tell what it is. What I’m building for is my own greed for freedom - something completely selfish.
At the same time, I’m trying to get there by doing good in the world. Being a muslim affects a lot my views of what I want and what’s the destiny I want to have. I want to do good in the world.
The balance between being selfish - aiming for my financial and intellectual freedom - and doing good in the world, reminds me of Adam Smith’s view of how individuals, following their own benefit (greed), will lift the society’s overall wealth (and worth) and speed up market interactions.
Salam, peace.
### Week 35: Answering the Wrong Question
- URL: https://mohammadshaker.com/en/blog/week-35-answering-the-wrong-question
- Date: 2024-06-20T00:00:00.000Z
- Tags: quick-read, reading, Human-Written
I've recently came into answering this question: Why is democracy the best form of government? That’s the wrong question.
#### Content
I've recently came into answering this question:
**Why is democracy the best form of government?**
That’s the wrong question. But I'll try and answer it first: Who said it’s the best form of government and why this question is enforcing it? It shows lack of knowledge of history and how democracy came to be. One can see that Democracy _is not the **best** form_ of government but _the most **viable** form_ of government that we have in recent times. And that’s a big difference. Reading Popper, Von Mises, and most importantly, Bertrand Russel, will prove how wrong the question is.
Now, I've answered the question, here are my thoughts: **I've learned from Peter Thiel that I shouldn't always answer the question that's being asked since it may well be the wrong kind of question or the wrong question.** What I should do is that I should always answer the right kind of question and the right question - which entails that I should, sometimes (and most of the time these days), change the question being asked altogether.
This proves hugely important business, in management, in 1:1s and in asking what we should work on next.
Salam, peace.
### Week 34: God and Future
- URL: https://mohammadshaker.com/en/blog/week-34-a-bias-for-action
- Date: 2024-06-13T00:00:00.000Z
- Tags: Brain dump, Human-Written
I've come upon a post by Tim Ferriss. He mentioned a blog containing a list of 140 questions worth answering.
#### Content
I've come upon a post by Tim Ferriss. He mentioned a [blog containing a list of 140 questions worth answering](https://guzey.com/questions/).
And since I love questions, I just sat and answered them. Here are the first couple of them.
1. **What are you thinking about these days?**
1. Having and raising kids.
2. Life at startup VS scaleups VS big companies.
3. Short term pain (lower comp when starting something new) VS long term gain (freedom of a profitable successful startup.)
4. Working on something I enjoy VS something that’s of benefit to winder audience VS both.
5. Extended off-time to think.
2. **What happens to your consciousness when you walk into a teleporter?**
- Stays the same.
3. **What’s the most embarrassing cold email you’ve sent?**
- Nothing in mind lately. Which is a worrying sign (not being out of one's comfort zone is worrying.)
- I've sent many sales/intro/customer-support emails/whatsapp/linkedin messages in my startup days but I don't think they can qualify as embarrassing.
4. **Does time exist?**
- It just did.
5. **Does god know the future?**
Wrong kind of question again. Since it entails that we can understand the nature of _the_ God or _a_ deity. The question is asking me, a player from within the game, whether the game exists, and if it exists, whether the creator the game knows how the game works and ends. The rules within the game do not apply to the creator of the game. Our rules and notions of time, past and future, within the game itself are of no resemblance to rules and notions outside the game.
More to follow in part 2.
### Week 33: My Tribe (and Hires)
- URL: https://mohammadshaker.com/en/blog/week-33-my-tribe-and-hires
- Date: 2024-06-07T00:00:00.000Z
- Tags: quick-read, Human-Written
What kind of people I want to be like? Or what kind of a person I would like to be friend with?
#### Content
What kind of people I want to be like? Or what kind of a person I would like to be friend with?
I would like to be part of the tribe of people who **think** for themselves **while being humble enough that they know they might be wrong.** **They actively work on proving every idea they have wrong.**
People who don’t take anything on face value. They read. They think. They tinker. They act. People with strong founded beliefs - weakly held.
At least that’s the tribe that I try to be from.
And these are the same exact kind people that I want to work with. And should hire.
Salam, peace.
(If you're one of these people, please do send me your CV or LinkedIn profile at: mo@spatialx.ai)
### Week 32: Working at Intersections
- URL: https://mohammadshaker.com/en/blog/week-32-working-at-intersections
- Date: 2024-04-25T00:00:00.000Z
- Tags: Brain dump, Human-Written
Something I keep using Books app on iPhone. It's good. Books app on iPad. It's OK. Goodreader app on iPad. I hate it. I use it for pdf books/documents.
#### Content
## **Something I keep using**
- Books app on iPhone. It's good.
- Books app on iPad. It's OK.
- Goodreader app on iPad. I hate it. I use it for pdf books/documents. The UX is terrible. Though, it's the sunk cost (I paid for it) maybe that I keep using. If anyone has a really good pdf reader on iPad where I can mark highlights easily, I would love to hear about it.
- Use of my Garmin watch (I have the Fenix 7) to track all my activities. It's good and has a long battery life (around 2 weeks.) I use none of the non-health features (no notifications, or spotify) I don't have to think about it compared to Apple watch where I have to charge it every day.
### What did I love when I was a kid? I should do more of.
- Lego (the Syrian version.)
- Drawing.
- Visual Basic programming language.
- PC games in an MS-DOS like this Prince of Persia in the 90s when I was a very young kid.
### Something I love doing?
I should aim to be at the top 10-20% on two things that I love + highly valuable + rare. Working at the intersections of domains. Being the bridge. Take design + tech. This is very rare to be in 1 person. And I try my best to be at that. Take product + design + tech + AI, this is even more rare. An I'm trying my best to be that in [SpatialX](http://www.spatialx.ai).
### How can I make my better?
- Early sleep, early rise. It feels like cheating. Was successful in the last 3 months (apart from Ramadan.)
- I have 1 meal a day. And it's a struggle to eat late in the day and have a comfortable sleep. Eating this meal early in the day is another struggle for me. But I'm pushing to get this meal earlier by 7pm more often. My sleep definitely improved.
### **Best buy of the week**
[A Syrian Madlouka](https://www.google.com/search?q=%D9%85%D8%AF%D9%84%D9%88%D9%82%D8%A9+%D8%A8%D9%81%D8%B3%D8%AA%D9%82&tbm=isch&ved=2ahUKEwiGze6znr6DAxVRnycCHTDDABMQ2-cCegQIABAA&oq=%D9%85%D8%AF%D9%84%D9%88%D9%82%D8%A9+%D8%A8%D9%81%D8%B3%D8%AA%D9%82&gs_lcp=CgNpbWcQAzoECCMQJzoFCAAQgAQ6BggAEAcQHjoECAAQHjoGCAAQBRAeOgcIABCABBAYUPUCWMkOYO8PaABwAHgAgAHfAYgBtQmSAQUwLjMuM5gBAKABAaoBC2d3cy13aXotaW1nwAEB&sclient=img&ei=wcKTZcboGNG-nsEPsIaDmAE&bih=753&biw=1512&rlz=1C5CHFA_enGB883GB883) from Levant. If you know, you know.
### 1 thing that is making me happy?
Starting something new.
### 1 thing that is making me angry?
What's happening in Palestine.
### Week 31: Be the Weakest Person in the Room.
- URL: https://mohammadshaker.com/en/blog/week-31-be-the-weakest-person-in-the-room
- Date: 2024-04-18T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. Something I'm learning Getting more clear in my thinking and my writing.
#### Content
Here we go again.
## **Something I'm learning**
Getting more clear in my thinking and my writing. On slack, my aim is that any person reading any message I write, should have 0 questions of what I meant. If it's a group of people, it should be very clear what the topic of the message is, what the goal/outcome is, what my rational is, and what the next steps are. The structure of:
- The What
- The Why
- The How
seems to be the cleanest. Books that helped me before and I'm trying to reread again are:
- Revising Prose, Keith
- Several Short Sentences about Writing, Verlyn Klinkenborg
- On Writing Well, William Zinsser
- The Pyramid Principle, Minto
### **Something that changed my mind**:
Longevity.
The more I read about health and the story of Peter Attia in his book _Outlive_ around heart diseases, the more I think I should take exercise seriously. Heart problems are not a joke. And in my family it goes deep. I should act accordingly.
A recommended read on longevity.

### **Something I'm thinking of**
“Be the weakest person in the room.” This served me long way in my career. Whenever you don't feel like growing, go somewhere else where you do. We're the average of 5 people we surround ourselves with. These should be carefully selected.
### Something that's making me really happy
My sister is, after such a long journey, an official pathologist in the US. Congrats Nuha.
### Something I'm enjoying
Writing.
It settles the chaos in my mind. It forces me to think in rational logical steps in one axis. And in complete different dimensions in different axes. It tells me what I'm thinking. Why I'm thinking what I'm thinking. And how to think about what I'm thinking.
### Week 30: Produce > Consume
- URL: https://mohammadshaker.com/en/blog/week-30-produce-consume
- Date: 2024-04-11T00:00:00.000Z
- Tags: Brain dump, Human-Written
Something I'm thinking of Output Input. Produce Consume. The idea is to make sure that I'm generating more than what I'm consuming.
#### Content
**Something I'm thinking of**
Output > Input. Produce > Consume.
The idea is to make sure that I'm generating more than what I'm consuming.
Example: generating and producing work that one is proud of and can share with others IS BIGGER THAN reading what others share. This is the absolute case of social media where we consume what others have produced.
Examples: Getting real work done IS BIGGER THAN reading. This can work the other way around if reading is in the shape of \[producing\] ideas.
Other examples that can work both ways are: (think > read), (read > think.)
**What am I learning or interested in learning?**
Learning how to learn. Meta learning. More on this later.
**1 scary thing to try this month**
Trying something new every week. Even if it’s taking a new way home (on foot, of course.) This counts. Starting small makes the bigger change easier to tackle.
**Something that changed my mind**
Working outside for a day (or half). Change of environment lead to change of ways of thinking.
**Best of what I read**
> “Men do not make laws. They do but discover them. Laws must be justified by something more than the will of the majority. They must rest on the eternal foundation of righteousness.”
\- Calvin Coolidge
### Week 29: Reframing the Week
- URL: https://mohammadshaker.com/en/blog/week-29-reframing-the-week
- Date: 2024-04-04T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. Here's what's been in my mind recently.
#### Content
Here we go again. Here's what's been in my mind recently.
## 1 thing that is making me happy?
Feeling that I'm doing something of value.
I sometimes feel that I'm not using all of what I've learnt and that drains me. If we are talking about where I come from, the Middle East, specifically, Syria, I do feel that I'm not doing enough at all there or with what I've learned and worked on. I want to change that. I'm doing it.
### What's one thing I cant stop talking about?
I love talking about design and things with unparalleled craftsmanship (quality naturally flows and follows.) I love talking about ways to be antifragile in my [health (Peter Attia, Nassim Taleb)](https://peterattiamd.com/), [Minimum Effective Dose (MED)](https://level1fit.com/the-minimum-effective-dose-tim-ferris-greek-to-freak-workout/), and about F1 cars circa [2006 on Renault's mass dampers](https://www.autosport.com/f1/news/banned-f1-tech-renaults-confidence-inducing-damper-solution-4982787/4982787/#:~:text=But%20at%20the%202006%20German,avoid%20the%20risk%20of%20disqualification.):

(credit: [Motorsport](https://us.motorsport.com/f1/news/banned-renault-mass-damper-tech/4788123/))
In 2006, I was a fan of Renault and Fernando Alonso (still am.) It was my first full season that I watched. And I was so amazed by the F1 tech used. Like Renault's [Mass Damper](https://us.motorsport.com/f1/news/banned-renault-mass-damper-tech/4788123/).
### What am I learning right now?
For around a year now, I've been mainly focusing on:
- Politics, Books.
- Philosophy, Books.
- Economics, Books. Financial Freedom and [not being a dog. Being a wolf.](https://mohammadshaker.com/2024/02/29/week-22-never-a-dog-ever-a-wolf/)
- Design, Doing it myself currently.
- Selling what I create as an artisan. I'm not yet an artisan.
### **1 new question I'm asking myself**
**If I'm to spend only 4 hours working this week. What would I work on?** I'm trying to ask this at the beginning of each week to force myself to think about what is absolutely important. It reframes the whole week.
Salam, peace
### Week 28: Forcing Oneself to Make Money
- URL: https://mohammadshaker.com/en/blog/week-25-forcing-myself-to-make-money
- Date: 2024-03-28T00:00:00.000Z
- Tags: Brain dump, Human-Written
10 years ago I've read what Jason Fried (from 37Signals/Basecamp) wrote:
#### Content
10 years ago I've read what Jason Fried (from [37Signals](http://www.37signals.com)/Basecamp) wrote:
> \[...\] I encourage entrepreneurs to bootstrap instead of taking outside money. On day one, a bootstrapped company sets out to _make_ money. They have no choice, really. On day one a funded company sets out to _spend_ money. They hire, they buy, they invest, they spend. Making money isn’t important yet. They practice spending, not making.
>
> Jason Fried
[Jason Fried](https://world.hey.com/jason) and [DHH](https://dhh.dk/), changed how I think about business and engineering since. And they are the reason I _bootstrapped_ almeta and [Alphazed](https://www.thealphazed.com). And they still do. And they are the reasons I love working in startups, building tech and teams from scratch. And that's why I recently joined/founded yet another startup, deep-tech, deep-science, [SpatialX](http://www.spatialx.ai).
Jason and DHH are contrarians. No BS, no faff, no BS mgmt. The jest of their thinking is simple: You should own your destiny and your destiny should be based on your own work. You work for your own destiny, not for someone else's.
This, what [Skin in the Game](https://www.google.com/search?q=skin+in+the+game+nassim+taleb&rlz=1C5CHFA_enGB883GB883&oq=Skin+in+the+Game+nas&gs_lcrp=EgZjaHJvbWUqCggAEAAY4wIYgAQyCggAEAAY4wIYgAQyBwgBEC4YgAQyBggCEEUYOTIHCAMQABiABDIICAQQABgWGB4yCAgFEAAYFhgeMggIBhAAGBYYHjIICAcQABgWGB4yCAgIEAAYFhgeMggICRAAGBYYHqgCALACAA&sourceid=chrome&ie=UTF-8), is. Ala Nassim Taleb.
(Jason's ideas on [making money while you sleep](https://signalvnoise.com/posts/2794-how-to-get-good-at-making-money) and [forcing myself to make money](https://signalvnoise.com/posts/1985-making-money-takes-practice-like-playing-the-piano-takes-practice) are recommended reads. 37Signals [books](https://37signals.com/books) are short, simple and to the point. I've just reread them for a 3rd time.)
Salam, peace.
### Week 27: Skills That Don't Age
- URL: https://mohammadshaker.com/en/blog/week-27-skills-that-dont-age
- Date: 2024-03-21T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. Here's what's been in my mind recently.
#### Content
Here we go again. Here's what's been in my mind recently.
## **Something I'm thinking of**
How to create a set of skills that:
1. don't age,
2. and have limited to no competition,
3. are valuable,
4. and are rare
Like the [Japanese Artisan work on pottery](https://www.youtube.com/watch?v=i9bt1W4SyRI&ab_channel=BusinessInsider). This always has an audience. And buyers.
The same thinking goes for coding. We engineers do code. But with the rise of ChatGPT, this is becoming less and less valuable. It's no longer how fast you code or how you produce. It's how to use the tools. Or better, how to create them.
### **Best of what I watched**
[Alonso VS Hamilton in Canada 2023](https://www.youtube.com/watch?v=nofjrfMKmkc&ab_channel=FORMULA1)
### **Something I keep using**
My Vivobarefoot shoes.
Our foot shape looks like this:

But if I look at most shoes in the market, are pointed at top, trying to lock the foot. Compare left and right.

Given that I have a knee problem, I do my best to try everything I can to get my knee healthy again. I've bought a pair of Vivobarefoot (have 0 affiliation with them) 2 years ago and they are simply the best shoes I've ever had (along with the Nike Free Run with wide fit.)
They don't affect my knee per say, but they do make the heel stronger. They take the shape of the foot. I do feel like a duck wearing them - and that's fine. They are minimal shoes - meaning 0 cushioning. A bit pricy. But definitely worth the try.
### **Something that changed my mind**
Hands-on and hands-off style of management. Will write about this later.
### 1 thing I’m remembering.
> Problems worthy of attack prove their worth by hitting back
>
> **Piet Hein**
Salam, peace.
### Week 26: People I Love
- URL: https://mohammadshaker.com/en/blog/week-26-people-i-love-2
- Date: 2024-03-14T00:00:00.000Z
- Tags: Brain dump, Human-Written
Ramadan Mubarak everyone! Ramadan started on Monday and it's the best month of the year for me.
#### Content

## My day in Ramadan
Ramadan Mubarak everyone! Ramadan started on Monday and it's the best month of the year for me.
- 8:30: Wake up. Later than usual.
- 17:30: Wrap up work.
- 18:00: Break the fast with a small portion. That is a couple of dates, full glass of \[warm\] water, yoghurt/cottage cheese with olive oil and walnut. I take supplements after that: Omega 3 pill, D3 supplements, Glucosamine (for the knee)
- 19:00: Go for a run or Gym for an hour.
- 21:30: Eat my main portion.
- 11:00: Pray Taraweeh and read the Quran.
- 12:00: Have Suhoor with couple of fruits + yoghurt.
- 12:30: Sleep.
### Something I’m remembering from the past
I can’t ever repay my mother and my father.
### Something I succeeded at recently
Playing to an engineer strengths not to his/her weaknesses.
### People I look up to
- Nassim Taleb, NNT
- David Hansson, [DHH](https://dhh.dk/)
- Andrew Huberman
- Tim Ferriss
- A real Arab literate. Still searching.
### **What I'm reading?**
All books. All digitally for now. Following a rating scale without a 6:
- [Guns, Germs and Steel](https://en.wikipedia.org/wiki/Guns,_Germs,_and_Steel). Rating so far: 8. Though, I've heard terrible rating on this from Dominic Sandbrook from "The Rest is History" podcast.
- A Human Action, Mises. Rating so far: 8
- [Dao of Capital](https://www.amazon.com/Dao-Capital-Austrian-Investing-Distorted/dp/111834703X), Mark Spitznagel. Rating so far: 9
- The Spirit of the Laws, Montesquieu. Rating so far: 6
### Say thanks for 1 person.
People who I built my first startup Almeta and later [Alphazed](http://www.thealphazed.com) with. Thanks all. You're absolute champs.
### People I love
People who make me laugh when I'm crying.
Salam, peace.
### Week 25: Days of Future Past
- URL: https://mohammadshaker.com/en/blog/week-26-people-i-love
- Date: 2024-03-08T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. The Past Present This week mark the first year of my father passing. And I will leave it at that. 2023 was a pretty tough year.
#### Content
Here we go again.

## **The Past** Present
This week mark the first year of my father passing. And I will leave it at that. 2023 was a pretty tough year.
RIP my father.
### The Past and the Present
Ramadan is among us next Sunday/Monday. And it's the favorite month of the year for many, including myself. The whole day structure changes, the way I work changes and the way I exercise. The gym is far from where I live now and I have to compensate that with running.
I have plan. Let's see how it goes. The idea is to maintain muscle mass. Not increase it. If I accomplish that during Ramadan, I'll be happy.
### **Something I keep doing**
- Going to the Gym 3+ times a week. Running 2 times a week. I'm back to running and it feels amazing. More on this in the upcoming weeks.
- Using a stair climb machine to pump up my heart rate and conditioning **before** I start my gym session.
- Reading daily.
- Reading old books only, pre 1990s - the year I was born. There are exceptions.
### **Something I'm learning**
- Politics, law, and philosophy.
- Refreshing my mind on deep learning and computer vision.
### **Best buy of the week**
Nothing. I moved out. So I am minimizing my purchases.
### **Worst buy recently**
I've never tried GymShark. And I did Dec last year with a purple gym top. It was awful. I had it washed with other tops. All my tops and shirts are dyed pinkish now.
### Why do I like ?
A principal engineer at Noon who pushes me to phrase and think about every word. This is a double edged sword though.
### 1 thing that is making me happy?
Change.
Salam, peace.
### Week 24: Never a Dog, Ever a Wolf.
- URL: https://mohammadshaker.com/en/blog/week-22-never-a-dog-ever-a-wolf
- Date: 2024-02-29T00:00:00.000Z
- Tags: Brain dump, Human-Written
This week, it's not a question. It's a core idea I believe in.
#### Content
This week, it's not a question. It's a core idea I believe in.
You never create wealth by being employed. Ever. This is the same thing my grandfather taught me. He started from literally zero, working in a shop when he was 9 years old to deliver for his family (with no father) under the French mandate of Syria in the 1920s.

After working as a young boy (صبي شغل) in lots of shops with older men in the Old Damascus Souq of Alhamedyeeh (الحميدية), he opened his own small shop. This caught fire by kids playing with fireworks after couple of years.
Started from zero again. This time with a smaller shop. Over the years he moved a bigger shop that became most famous for everything from toothpaste, to paste, to brushes and everything in between. He did by himself. Doing it by himself. Losing and gaining with his own money. Total skin in the game. And total antifragility.
**In Nassim Taleb’s universe, which is our universe, modern salaried life is essentially slavery.**
In, **Skin in the Game**, Nassim uses the analogy of the [Dog and the Wolf](http://www.bartleby.com/17/1/28.html) from Aesop’s fables:

(Image from [draudreyt.com](https://www.draudreyt.com/post/2018/06/14/the-truth-about-dogs-and-wolves))
> A gaunt Wolf was almost dead with hunger when he happened to meet a House-dog who was passing by. “Ah, Cousin,” said the Dog. “I knew how it would be; your irregular life will soon be the ruin of you. Why do you not work steadily as I do, and get your food regularly given to you?”
>
> “I would have no objection,” said the Wolf, “if I could only get a place.”
>
> “I will easily arrange that for you,” said the Dog; “come with me to my master and you shall share my work.”
>
> So the Wolf and the Dog went towards the town together. On the way there the Wolf noticed that the hair on a certain part of the Dog’s neck was very much worn away, so he asked him how that had come about.
>
> “Oh, it is nothing,” said the Dog. “That is only the place where the collar is put on at night to keep me chained up; it chafes a bit, but one soon gets used to it.”
>
> “Is that all?” said the Wolf. “Then good-bye to you, Master Dog.”
Freedom should not be exchanged for comfort or financial gain.
**I shall be a wolf. Never a dog.** And that's my next move.
Salam, peace.
### Week 23: Montesquieu and the Spirit of Law
- URL: https://mohammadshaker.com/en/blog/week-23-montesquieu-and-the-spirit-of-law
- Date: 2024-02-22T00:00:00.000Z
- Tags: 100-words, Brain dump, Human-Written
Here we go again with thoughts of my week. Something I keep doing Reading a deep/thorough book for 30 min in the morning.
#### Content
Here we go again with thoughts of my week.
## **Something I keep doing**
Reading a deep/thorough book for 30 min in the morning. Currently it's: Montesquieu's **Spirit of Law.**
Montesquieu, the man behind the theory of **Separation of Powers**, had ever lasting ideas. This is implemented in many constitutions throughout the world.

### **Best of what I read**
- Human Action, Von Mises (Rating so far: 8)
- Diderot, Audible biography (Rating so far: 6)
- Dao of capital, Mark Spitznagel (Rating so far: 9)
- Safe Haven, Mark Spitznagel (Rating so far: 9)
### **Something I keep using**
- Audible. I use it in long walks.
- My backpack. Mandarina Duck. It's great and I've been using it for 5 years now. Would buy it again.
- Reminders app on iPhone as a planner. More on this later.
### **Best of what I watched**
Sami Yusuf's Beyond the Stars concert.
### **Something that changed my mind**
Putting focus on what I am good at and makes money VS what I love and doesn't necessary generate good money.
### Something new to try
I really like espresso (single shot) recently. Pure, clear, and simple. I like the bitterness.
### Something I’m remembering from the past
My dad and his prayers.
### Why do I dislike
Jordan Patterson. Never liked him. And his recent views are even worse.
### What's one thing I cant stop talking about?
I can't stop talking about something I want to learn about. This makes people eager to teach me more.
- Politics
- Law
- Nutrition, sport and health
- Tranquility (ala Seneca)
- Nassim Taleb (NNT)
- Antifragility (ala NNT)
- [Barbel Strategy](https://www.google.com/search?q=barbell+strategy+nassim+taleb&rlz=1C5CHFA_enGB883GB883&oq=Barbel+Strategy+Nassim+Taleb&gs_lcrp=EgZjaHJvbWUqCQgBEAAYDRiABDIGCAAQRRg5MgkIARAAGA0YgAQyDQgCEAAYhgMYgAQYigUyDQgDEAAYhgMYgAQYigUyDQgEEAAYhgMYgAQYigUyDQgFEAAYhgMYgAQYigXSAQgzNTUxajBqN6gCALACAA&sourceid=chrome&ie=UTF-8) (ala NNT)
### How can I make my days better?
One focus for the whole day. Or split a day into two and have 1 focus each. Doing it a lot recently and it's paying off.
### Say thanks for 1 guy today.
Just sent a message to my friends who I had great conversation with recently. I can't thank them enough on how kind they are. Thanks, my friends.
- [Mehdi Zonjy](https://www.linkedin.com/in/mehdi-zonjy-71bab5159/)
- Abd Nizam
- Hussam Sandouk
- [Zaher Wanli](https://www.linkedin.com/in/zaher-wanli/?originalSubdomain=de)
- [Hasan Sarhan](https://www.linkedin.com/in/hasan-sarhan-aa845a65/)
That's it for this week. Peace, Salam.
### Week 22: A Place I Call Home
- URL: https://mohammadshaker.com/en/blog/week-22-a-place-i-call-home
- Date: 2024-02-15T00:00:00.000Z
- Tags: Brain dump, Human-Written
Home. The place I was born. Because it’s the place I call: home. Many of my Syrian friends disagree and I will disagree even more.
#### Content
## Why do I like .. Home?
Home. The place I was born. Because it’s the place I call: home. Many of my Syrian friends disagree and I will disagree even more.

### Something I succeeded in doing
Increasing my calorie burn in strength sessions by doing the stepper first for 15 min. This will shoot up my heart rate all the way up to around 155 bpm range.

From there on, my heart seems to be kept at the higher range of 120 if I keep my exercise in sequence with minimal rest.
This works well when I’m doing supersets (i.e. overlapping sets from 2 exercises.) Supersets gives one muscle group an amble time to recover while the other group is under tension.
This is not an advise, nor sth I've read and I'm recommending. I just wanted to increase my calorie burn in strength sessions and that’s a nice trick I found to be working.
The result is: I jumped from average of 350 calorie/session to the range of 500s calorie/sessions.
### **Best of what I watched**
This is really good. Chip Conley interview with Tim Ferriss.
### **Best of what I'm read**ing
I am still very much enjoying the [DK's Big Ideas book series](https://www.dk.com/ca/promotion/big-ideas-series/) on politics, history and philosophy. It's good for general knowledge. The problem though is it lacks connectedness and depth.
One way to fix this is the following: Some DK books talks about the same person in different books - in history, in philosophy or in science. Reading DK History, DK Philosophy and DK Science books on parallel, encountering same person twice or thrice in two different books helps bridge the gab across topics - something the series don't do out of the box.

That's it for the week.
Salam, peace.
### Week 21: Invert Incentives
- URL: https://mohammadshaker.com/en/blog/ed-21-invert-incentives
- Date: 2024-02-08T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. Here's what's been on my mind recently.
#### Content
Here we go again. Here's what's been on my mind recently.
## A difficult conversation I had.
In management, I should position any feedback in terms of the benefit of the listener. Simply, minding incentives. This is essential for giving feedback as a manager (especially negative feedback.)
What people do (and sometime say) is based on their incentives. Understand incentives, and I shall understand what people want. And I should know how to approach them.
Ala. Charlie Munger: “**Show me the incentive, and I'll show you the outcome.”**

### Best of what I watched.
Tim Ferriss interview with my favorite author, Nassim Taleb. Although the podcast gives a very brief look of what Nassim has worked on or what his philosophy is, it’s still a good intro to Nassim from financial market POV.
[Podcast version here.](https://tim.blog/2023/09/07/nassim-nicholas-taleb-scott-patterson/) And youtube version [here](https://www.youtube.com/watch?v=18ZxCxN2ZMo&ab_channel=TimFerriss).
### **1 thing that is making me happy?**
- Journalling for 20 min in the morning.
- Reading for 40 min in the morning.
- Going weekly to a coffee place I like with Lamis, like [Arome in London](https://www.google.com/search?q=Arome+in+London&rlz=1C5CHFA_enGB883GB883&oq=Arome+in+London&gs_lcrp=EgZjaHJvbWUyBggAEEUYOTIICAEQABgWGB4yCAgCEAAYFhgeMggIAxAAGBYYHjIICAQQABgWGB4yCAgFEAAYFhgeMggIBhAAGBYYHjIICAcQABgWGB4yCAgIEAAYFhgeMggICRAAGBYYHtIBBzM0N2owajeoAgCwAgA&sourceid=chrome&ie=UTF-8&lqi=Cg9Bcm9tZSBpbiBMb25kb25I7_S736GvgIAIWhcQABgAGAIiD2Fyb21lIGluIGxvbmRvbpIBBmJha2VyeaoBQxABKgkiBWFyb21lKAAyHxABIhu7Pc9wDDXMeGgebZkJMMbpczMRavjOm-6--lUyExACIg9hcm9tZSBpbiBsb25kb24#rlimm=17013972700886924882), and ordering exactly the same thing.
- Having a structure for the day by planning my day, a day before.
- Blocking time for myself to **sit** and **think**.
- Working with people smarter than me.
### **1 thing that is making me angry?**
People who need high maintenance. And people who don't care.
### **What am I unwilling to feel?**
Stagnation.
Salam, peace.
### Week 20: 7 out of 10 Is Bad.
- URL: https://mohammadshaker.com/en/blog/ed-20-7-out-of-10-is-bad
- Date: 2024-02-01T00:00:00.000Z
- Tags: Brain dump, Human-Written
I've been doing it for the last year, keeping a small notebook, in my pocket, whenever I go. I like the physicality of the pen and paper.
#### Content
## **Something I keep doing**
I've been doing it for the last year, keeping a small notebook, in my pocket, whenever I go. I like the physicality of the pen and paper. Whenever I want to wait for someone or whenever I have 2 hours in the weekends, I just go into a cafe, sit down, open my small pocket notebook, armed with a pilot pen, and I simply sip coffee (or green tea) and I _think_. Analog. No digital. No distractions. Just myself and I, with a pen and paper.


### **A quote I'm living by**
> Its not what you say, it's what you do.
Another version is:
> Don't listen to what people say. Look at what they do.
This reminds me of [Nassim Taleb's idea of Skin in the Game](https://www.amazon.co.uk/Skin-Game-Hidden-Asymmetries-Daily/dp/0241247470). Which is a must read. Again, Nassim's ideas changed me 10 years ago. And they still do.
### **Why you shouldn't have a 7 out of 10 rating. For anything.**
Something I've learned in Neurofenix and from my friend [Dimitris Athanasiou](https://www.linkedin.com/in/dimitrisathanasiou/?originalSubdomain=uk), is that I shouldn't set a rating of 7/10. For anything.
7/10 doesn't tell you anything. It's misleading. It's neither too good nor too bad. If it's 6/10, it means OK and you shouldn't bother. If something is 8/10, it means it's good and you can proceed.
Used it ever since for watching movies, book ratings and recommendations, and, most importantly, hiring or feedback.
### **What I'm reading**.
Lately another set of books. These are lengthy, meaty and require brain power. All the following are linked in topics in a way. Between economics, policy making and society building (or destructing.)
I read them after a gym session or in the morning just after I wake up- when my mind is clear or the blood is flowing. Ratings are WIP since I haven't finished the books yet.
- The Constitution of Liberty, Hayek. Rating: 8/10

- A Human Action. Von Mises. Rating: 9/10.

- The Open Society and its Enemies, Popper. Rating: 8/10. It's amazing how different I view Plato after reading this one. Both his good side and the bad.

- Dao of Capital. Rating: 8/10.

### **Something I'm thinking of**
- What I will be doing when in my 40s, 50s, 60s, 70s or the 100s with the current pace of AI and the constant change. I'm not sure anyone would be that adaptable. So What's the constant in all of this? This is what I should be following.
- Taking extended time off. It's mind boggling how this simple act does affect my creativity while am off and _after_ I'm back. I know this because I've tasted this in the past. Off time is actually more important than _on_ time. It's where the creative thinking happens to cultivate a meaningful _on_ time. I should aim to take at least 1 week off every 2 months. Even 4 days (2 days off + the weekend) can do the trick if it's truly off (armed with a pen and paper.) That's something I **must** do regularly. The upside is just too great.
Salam, peace.
### Week 19: \"Where Is the Good Knife?\"
- URL: https://mohammadshaker.com/en/blog/ed-19-where-is-the-good-knife
- Date: 2024-01-12T00:00:00.000Z
- Tags: Brain dump, nassim-taleb, nnt, wittgenstein, Human-Written
“Where is the good knife?” If I’m looking for the good knife in the kitchen, that means I have other bad knives. I should throw those out.
#### Content
## Best Question I’ve Recently Read
**“Where is the good knife?”** If I’m looking for the good knife in the kitchen, that means I have other bad knives. I should throw those out.
This applies to anything:
- If I'm looking for the good shoes, it means I haver bad ones.
- If I'm looking for the good book, it means I haver bad ones.
- If I'm looking for the good Y, it means I haver bad Ys.
The Y can be anything: teams building, friends, relationships, supermarkets. Part ways with them, I should.
### The limits of our languages..

Side note on an adjacent topic that made me smile: A knife, should, by definition be a good knife. There shouldn't be a good or bad knife. A good knife is a knife. A bad knife is not a knife. I've never thought about this until I've read about Ludwig Wittgenstein. Another worthy concept is: [Wittgenstein's ruler](https://twitter.com/nntaleb/status/1082560527908384768?lang=en).
### Something I keep doing.
Intermittent Fasting. For me it’s eating 1 meal a day. This doesn’t have to be tied to a Keto or any other diet for me.
Intermittent fasting is something I’ve done for 12+ years now. It's basically a form of time-restricted eating. I only eat within a time window. The smaller the window, the better. For me:
- The window ranged from 2 to 7 hours over the years.
- I eat my main, and only, meal around 6pm (or later.)
- The meal size used to be big. In the last year I split it into 1 good-size meal followed by another portion 1 hour later. Much better for digestion, much better for sleep.
### Something I remembered recently.
[Mel Gibson’s Apocalypto](https://en.wikipedia.org/wiki/Apocalypto) is one of my all-time favorite movies.

### Week 18: Stay Still and Think
- URL: https://mohammadshaker.com/en/blog/stay-still-and-think
- Date: 2024-01-05T00:00:00.000Z
- Tags: Books, Brain dump, Reading, Weeklies, writing, Human-Written
On conifers vs angiosperm trees and their strategy of growth. From Dao of Capital. It blew my mind. I'm writing a piece on this.
#### Content
## **Best of what I read** **recently**
On conifers vs angiosperm trees and their strategy of growth. From **[Dao of Capital](https://www.amazon.com/Dao-Capital-Austrian-Investing-Distorted/dp/111834703X)**. It blew my mind. I'm writing a piece on this.
### **What makes me wake up enthusiastic lately? What can I do this month around this?**
- Seeing my VO2 Max going up. Very slowly. Training diligently has a second order effect. It's like sunk cost where I eat whatever I want, but 1 meal a day and no excess sugar. Train hard, eat whatever I like.
- Reading old, deep, thorough books.
- Tranquility. Stay Still and Think. Thinking sessions where I sit for 2-3 hours and ask my self questions, think about them and write my answers. Normally with a cup of coffee and a bottle of water.
- Writing. I'm enjoying the research part and "write to think" style of long-form writing I'm doing lately.
- Walking. Long walks under the sun.
### Best of what I watched
A style of Man's Search for Meaning from the Arab Palestinian Humam Yehya.
### Something I'm failing at.
- Early sleep, early rise. It feels like cheating. Though I'm yet to get the owl in me to sleep early and rise early.
- Eating early and being light going to bed. For years, I'm eating 1 meal a day. And it's usually after I finish work and a gym session at around 8. This makes going to bed with a big ball in my stomach really annoying sometimes. Still need to fix this. Will write about this when I do.
### Something that surprised me
I never knew such thing exists..
### **Something that changed my mind**
- Being stubborn from Nassim Taleb (NNT.) A blog to follow.
- Getting advice from experts from NNT. A blog to follow. A good thing I did in the last couple of years is never to believe anything without doing my own research around it.
### Week 17: I Think, Therefore I Laugh
- URL: https://mohammadshaker.com/en/blog/week-17-i-think-therefore-i-laugh
- Date: 2023-08-23T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. Here's what's been in my mind recently.
#### Content
Here we go again. Here's what's been in my mind recently.
## **Something I started doing**
The older I get, the faster time seems to go. This is true for everyone it seems. What makes time slows down is the anti-routine activities we do in our days. My day become _fatter_ (and more exciting) the more I do _not_ stick to my routine.
The problem is that I'm terrible not to stick with a routine. It drives my whole day off balance. The way I'm fixing it is by blocking half day a week and walk and work outside.
In these lovely summer day (which are not always lovely in rainy London), I'm going and working outside once a week in a cafe. A change of scenery should make us more creative - aside from the healthy benefit of walking to the cafe, before and after.
### **What I'm reading**


**The Road to Serfdom** and **Law, Legislation and Liberty**. Both by F. Hayek. Reading about socialism and economics. Hayek's writing style is simple, clear and engaging. I've come to know Hayek from my favorite write Nassim Taleb and he never fails.

**_F. A. Hayek_**
Hayek (Friedrich Hayek) and Mises (Ludwig von Mises) are two landmarks from the Austrian School of Economics. The more I read what they wrote, the more I understand what's happening on nearly everything around me. They are far, far ahead from anyone writing today on mostly everything - with the economy or daily life.

_**Ludwig von Mises**_
And I bet they both would've hated Bitcoin.
### **Something I keep doing**
Mixing strength training with Zone 2 exercises every week. Lately, my program is:
- 2 days strength training
- followed by 1 day Zone 2
- 1 day rest (sometimes I skip it)
- 2 days strength training
- repeat.
It keeps me balanced between being exhausted physically from strength training, and being exhausted mentally from a Zone 2 exercises.
Zone 2 exercises with constant watt on an Elliptical machine is a _good_ killer.
### **Best of what I watched**
The amount of creativity in Top Gear is immense. I'm rewatching the seasons over again. All of the 21 seasons. Living in the UK for 6 years now, I understand a bit more the British sense of humor. I like the self-inflecting jokes.
What I like more about it is how universal it is. How natural each episode flows. And, again, the amount of creativity they have in setting up every episode to be special on its own. It's funny, smart and informative - all in 1. For motor heads, this is an added bonus.
### **Best of what I read**
> **I don't understand bus lanes.** **Why do poor people have to get to places quicker than I do?**
>
> Again, by Jeremy Clarkson from Top Gear
This reminds me of a very enjoyable read by **John Paulos** in his book: **I Think, Therefore I Laugh: The Flip Side of Philosophy**. Easy read for quick wittedness.

Enjoy!
### Week 16: Journalling.
- URL: https://mohammadshaker.com/en/blog/week-16-journalling
- Date: 2023-07-30T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. Here's what's been in my mind recently.
#### Content
Here we go again. Here's what's been in my mind recently.
## **Something I keep doing**
Keeping my phone in the bedroom for at least the first 6 hours of the day. Lately I'm extending this to my whole working day. Have been doing it for a year now. And it's a home run for focus and being present with the work in hand.
### **Something I'm thinking of**
> **Hard Choices, Easy Life. Easy Choices, Hard Life.**
>
> Jerzy Gregorek
I keep thinking of this every now and then. And it keeps pushing me out of my comfort. Good reminder for me to keep thinking like this.
### **Something I got back to.**
Journalling. Journalling for 10 min in the morning. A simple 1 big pager should do it. I've done it on and off for years. And it always feel good when I am back at it. Same as walk or exercise. You never regret them. Same for writing. It's in simple 4 sections:
1. **Section 1**: A simple prayer (I came up with) at the beginning of the page.
2. **Section 2:** What I'm grateful for. Mention people name, places, experiences. (I find myself mentioning the weather far too much it seems. I love the sun. And it's always there in the list whenever it's sunny in London.)
3. **Section 3 (past/yesterday)**: what I did yesterday and how I felt.
4. **Section 4 (future/today)**: 3 questions:
1. What can make today an amazing day for me?
2. What's 1 thing that if done, I'll be satisfied with my day?
3. What I'm doing for what I want to become? A set of things I want to focus on like: I write them as “I'm X”.
1. I'm a writer → Share what I did last week, simple and concise. Learn how to be economical with your writing (my favourite book on this is [Revising Prose, Keith](https://www.amazon.co.uk/Revising-Prose-Richard-Lanham/dp/0321441699))
2. I'm a good servant of society → Send a friend a thank you note.
3. I'm a reader → Read: Confessions of an Advertising Man by David Ogilvy and The Republic by Plato.
4. etc.
Had this as my companion for journaling for the last couple of years.

### **Something new (and old) I'm doing.**
Same as last week, this is not new at all: I'm yet again, going back to how Syrians drink their tea, [zuhurat](https://en.wikipedia.org/wiki/Zhourat_shamia) زهورات شامية or mostly any drink. The whole idea of drinking a hot beverage is the experience you make around drinking the beverage. It's sensual as much it is actually basic, trivial and simple. In Syria, we drink hot beverages by pouring the drink in small cups. You drink the cup while its hot and pour another one if you like from the hot teapot. This way, the beverage is always hot because it's always from the teapot which is always under a cover (a simple towel does it.)

The whole experience around heating the tea (or zuhurat for me) in a teapot, waiting for it to simmer under the towel, then drinking it bit by bit, always hot, is reminding me of home. And how it felt when I was a kid.
### **Something I failed at**
Reading more. I'm averaging around 45 min a day lately which is not the best for me.
Reading is essential for me to clear my head and cool down. It makes me calm. It makes me think and it makes me wonder. I do it mostly while sitting on chair or in the sunny corner of my house - on a cushion on the floor just against the window.

I mostly read before work early in the morning and in the bed. In good days, it will be around 1-2 hours in the morning and 30 min in the evening.
### **Something I succeeded in doing**
At last, I'm seeing the fruits of training diligently. I'm front-squatting to weights I haven't done before at all. Weights that I was doing only 7 reps in 3 sets 2 months ago, now I can do 10 reps in 5 sets. Same for other body parts, though the improvements differ. It's making me committed.
### **Something I keep using**
Apple AirPods. A great example of great engineering and seamless design and user experience. I think they are one of the best things one can use for daily work. I became aware of the potential danger of wireless sets and whether they are completely safe or not. Need to investigate this more I think.
That's it for this week for me. Get at it.
### Week 15: Nonviolent Communication
- URL: https://mohammadshaker.com/en/blog/week-15-nonviolent-communication
- Date: 2023-07-03T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here's what's been in my mind recently. What I'm reading lately? NVC. Nonviolent Communication by Marshall B. Rosenberg.
#### Content
Here's what's been in my mind recently.
## **What I'm reading lately?**
NVC. Nonviolent Communication by Marshall B. Rosenberg. This book can win an award for the worst book cover and typography. But that's not the point.

At its core, NVC encourages individuals to observe **without judgment**, identify **feelings**, connect them to underlying **needs**, and make requests that are **clear**.
I now remember what my friend and colleague at Noon, [Andy Guest](https://www.linkedin.com/in/andy-guest-90a53015?miniProfileUrn=urn%3Ali%3Afs_miniProfile%3AACoAAAMahBoBqLnm20WrESHg6ySk9SPVskfaKoM&lipi=urn%3Ali%3Apage%3Ad_flagship3_search_srp_all%3BB9s7UAfzSjGXy2RZAZxqkg%3D%3D), was telling me around providing feedback in an offsite and how it actually matches 1:1 with NVC. And it's an absolute eye opener for me since. It's a gem for any manager or anyone who directly managing people.
The key components of nonviolent communication can be summarized as follows:
1. **Observation**: Describe what I observe without adding judgments. The data.
2. **Feelings**: Expressing my emotions given what I observed.
3. **Needs**: Connect my feelings with my needs, values, or desires.
4. **Requests**: Make an actionable request rather than demand. Requests should meet our needs and respect the needs of others.
It works in personal relations too. Instead of telling my wife: "You are wrong." I should say: "Given that I've seen you doing X, I feel sad. And that makes doing Y really hard for me."
I should always remind myself of this. It's much more peaceful and civilized in any situation.
### **Something new I'm trying**
Actually this is not new at all. But it seems that the older I get, the more I like what old people likes. Especially old medeteranians. And especially old Syrians. When I was a kid I hated famous singers of the time like Fairouz and Sabah Fakhri. Now they are my favorites. I rarely like any new kind of music.
### **Something I revisited**
I like the [NerdWriter](https://www.youtube.com/@Nerdwriter1). And I keep coming back to this exact video. It worth anyone's 6 minutes.
(The thumbnail also matches the title of this blog.)
### **Why do I like a **
This week is: **Why do I like walking?** 3 reasons:
1. Simply because I feel alive this way. I'm actually moving from one place to the other.
2. It gives me extended time to think. Think slowly and at my own pace. I won't be rushed and I won't be pressured to finish anything. I'm managing around 9.500 steps a day on average. I do most of these steps in 3 days a week.
3. It reminds me of my father. He used to walk after the morning prayers for around 2 hours in the morning (that means between 5AM to 7AM) and another 2 hours just before sunset. He's actually the person who's doing what [Huberman is suggesting](https://www.youtube.com/watch?v=t-ezOLT2Kv0&ab_channel=ChrisWilliamson) before Huberman was even born.
Here's my average daily steps during the last year. Naturally summer is the best for walking.

### **Some software engineering books that keeps popping up lately in my 1:1s with engineers**
For technical expertise, I think the basics are the most important, never a specific tech. So if I want to improve on sth, I sometimes just get back to the basics, to square 1. Don’t go and only read a book on React Native if you want to improve your tech skills. Read the basics that are applicable everywhere.
Here’re some of the books that helped me on the way:
**Good for going back to the basics:**
- Introduction to Algorithms by Thomas H. Cormen
- Algorithms: A Creative Approach Paperback – 1 Jan. 1989 by Udi Manber
**Design and architecture:**
- Patterns of Enterprise Application Architecture (Addison-Wesley Signature Series (Fowler))
- Design patterns : elements of reusable object-oriented software
**There are other famous books I didn’t quite like**:
- Effective Java by Joshua Bloch
- Clean Code by Robert Cecil Martin
- Clean Architecture: A Craftsman’s Guide to Software Structure and Design. by Robert Cecil Martin
There are other books like The Mythical Man-Month and The Pragmatic Progrmmer that I really like. But these are different from the coding/data structure books above.
Do you have favorites? Please share them via email directly to me and I'll happily add them here and mention you. (my email: mohammadshakergtr@gmail.com)
### Week 14: Shape Up
- URL: https://mohammadshaker.com/en/blog/week-14-shape-up
- Date: 2023-06-23T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. Here's what's been in my mind recently.
#### Content
Here we go again. Here's what's been in my mind recently.
## **Something I keep doing**
Being diligent in my exercise. Although I'm starting lately to overdo it, especially with weight exercises. The heavier you go, mindfully, and the more you improve, and the more you have to be mindful about rest it seems. I'm getting depleted more than I would like to. Rest matters the same as exercise.
Another way around this is doing lighter weight exercises in the off days. Not completely abandoning the exercise. Doing an active rest like long walks, easy pace is doing it for me.
Last week average: 5 hours strength training + 1 workout Zone 2 = 6 days in the gym.

### **Something I succeeded in doing**
VO2 max sessions that's not an agony to do. According to Peter Attia, a person should do a 4x4x4. Meaning:
- 4 min full effort,
- 4 min zone 2 effort.
- Repeat that 4 times (32 min total.)
Normally this is done on a rowing machine.
I just didn't want to do it like this in the last couple of weeks. It's boring. Am trying to make them fun. What I'm doing is 50 min workout and without 4 min rest in between (non scientific, just fun.) I'm doing the following:
- 5 min: elliptical. You can do treadmill with an incline. For me, it's easier on my knees on an elliptical.
- 40 min: Combination of
- Weighted sledge for 40 meters total. Forward and Backward. the backward side is for knees from the [KneesOverToesGuy](https://www.youtube.com/c/TheKneesovertoesguy).
- 20KG Kettlebell on each arm for 40 meters total.
- Repeat.
- 10 min: elliptical.
Am young. I should be able to handle it and have some fun from time to time. I want to feel life and alive. And I sure am when I am doing these sessions.
### **Something I keep buying**
Blueberries. I simply say to myself that Lamis, my wife, like them. And then I eat them all by myself. I think I also tell myself they are good antioxidants and that they're good for strength and the gym work I do. I find a reason to keep buying and eating them. No worries.
### **Something I keep using**
My Moleskine notebook. It's a bit pricy but it does the job. The quality of paper is superb and that matters if you draw or design on paper. I always choose the one with dotted papers.

I use them for design, reading and all my writing and thinking.

### **Best of what I watched**
I love [Aldaheeh](https://www.youtube.com/watch?v=BYLLXnSZqIU&t=885s&ab_channel=MuseumofTheFuture%D9%85%D8%AA%D8%AD%D9%81%D8%A7%D9%84%D9%85%D8%B3%D8%AA%D9%82%D8%A8%D9%84). I'm proud we have an Egyptian in the Arab world that has this impact on the youth - simply by explaining science in simple terms in 15 min. He's great. I've watched [this video](https://www.youtube.com/watch?v=BYLLXnSZqIU&t=885s&ab_channel=MuseumofTheFuture%D9%85%D8%AA%D8%AD%D9%81%D8%A7%D9%84%D9%85%D8%B3%D8%AA%D9%82%D8%A8%D9%84) on the Arabic Language as the Language of “daad.” Which we seem to get all wrong. It's not the language of “daad” after all. Recommended for any Arab.
A big side note: I should rethink this it seems. He's more of a presenter, than purely a science person. [His video on Keto diet](https://www.youtube.com/watch?v=fWrOwlvPWVQ&ab_channel=NewMediaAcademy) is purely misleading and unscientific. If you're a regular guy like me who is interested in exercise and diet and read a bit, you'll see [this episode](https://www.youtube.com/watch?v=fWrOwlvPWVQ&ab_channel=NewMediaAcademy) and be surprised. He got lots of things wrong. And he mixed between Keto diet and intermittent fasting.
### **Best buy**
If you're a runner or want a really good top for gym, I've [bought this by Nike](https://www.nike.com/gb/t/techknit-dri-fit-adv-long-sleeve-running-top-vwhgNk/DV4194-084) a while back, and it's a really good top. I've bought another one again. It's not as white as it appears to be. It's actually grey.

### **Something I'm learning**
I've just finished the first walkthrough in the **[Moral Foundations of Politics](https://www.coursera.org/learn/moral-politics/home/week/1)** course on Coursera. A really good intro on the subject. If you know any other good course on politics on philosophy, please do send me an email. Directly here: [mohammadshakergtr@gmail.com](mailto:mohammadshakergtr@gmail.com)
### 80/20 of my time for the week went on..
My work day is dominated by meetings lately. I'm trying to keep at least 1 day where it's completely free of meetings. I think if we, [at Noon](https://www.learnatnoon.com/), are to do something dramatically different in the EdTech space in particular and in tech in general, we should follow dramatically different approach of work VS other companies.
And I do think it starts with how we plan, execute, learn and iterate - closing the feedback loop for another cycle of plan, execute, and learn. We've made lots of changes in the last 1 year to plan properly and execute and iterate faster. But we still have long way to go.
We're starting to adopt to some uncommon practices in our industry. These are first principled by 37Signals/Basecamp. These guys got it right early, and still do. I read all their books 7-8 years ago and it effected all my management style, me starting a company, hiring and firing, and me adopting thinking and adopting different internal processes of how XFN teams should work. I'm in-progress of widely implementing lots of these practices at Noon. Give their [ShapeUp](https://basecamp.com/shapeup) book a read for anything tech/engineering/design/product related. Or any of their [other books.](https://basecamp.com/books)
They are very easy quick reads. But they are gems in a mired of insanity of engineering practices in the tech industry. I love these guys.

### I'm Not Different.
- URL: https://mohammadshaker.com/en/blog/im-not-different
- Date: 2023-06-04T00:00:00.000Z
- Tags: personal, alphazed, Human-Written
I think people are different. What works for others doesn't work for me. And I should accept that.
#### Content
I think people are different. What works for others doesn't work for me. And I should accept that. My way of work is completely different from other people. And that's OK. It's not about which way is better and which is not. It's simply about what works for each of us.
I hate hacks in general. And most people follow hacks. I don't want to. I prefer systems. Productivity _gurus_ always talk about the next _hack_ you can follow, the new tool you should try, and the new way you should _absolutely_ follow. I don't believe most of this. All previous words in italic are intentional.
I tend to work best, and produce best, when I focus on 1 thing, and 1 thing only for an extended period of time. Sometimes it's hours, sometimes it's months or years.
When I used to work for my own company, Alphazed, I used to set CEO vs CTO days. Meaning on a given day (or couple of days) I will only work on sales, PR and marketing. On another day I'll only work with the team on the tech side and drive technical decisions. I never mix the two in the same day. This was a great learning for me if I'm to fill both roles.
Working on and off for short period of time and on multiple things simply doesn't work for me. And that's OK. It seems that I've spent the last year fighting this in me and I shouldn't. I know my way work and that's great. I should always experiment of what's working and what's not. But that's it. That were it ends. What works, I adopt. What doesn't, I set aside. Time-box any experiment and I will be able to directly see when it's working and what's not. I should keep it as simple as that.
Am I different? I'm not. We are all different. And then we're not different at all.
### Week 13: Méditerranéen
- URL: https://mohammadshaker.com/en/blog/week-13-mediterraneen
- Date: 2023-05-25T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here's what been in my mind in the last week. Something I keep doing Reading in the sun, sitting on the floor. I don't want a mansion.
#### Content
Here's what been in my mind in the last week.
## **Something I keep doing**
Reading in the sun, sitting on the floor. I don't want a mansion. This is enough for me. It makes me extremely happy.

### **What's wrong with me?**
Recently I feel lost outside work. I feel myself focused on planning, and replanning, and planning again - not progressing on what I planned though. I wasn't like this. I was either ON or OFF. Recently I'm neither ON nor OFF. I'm not talking about my day job. I'm talking about the rest of my life and personal projects.
I was listening to [Tim Ferris conversation with Gary Keller](https://buff.ly/2Z9uJid) and the focus on one thing at a time. It's perfect timing for me. And it reminds me of my younger self. Anyone can listen to the 2nd half of it and get something out of it. It's good.
A quick thing I can do is limit my decision making to 1 major thing a day. Simply asking myself: If I am to do only 1 thing, and 1 thing only this day or this week, what would it be?
Simple question for a good day. More on this next week.
### **Something I keep doing**
Exercise. 6 days a week. If long walks (1.5h+) counts, then it's 7 days a week.
### **Best of what I watched**
Not sure yet, but [this course](https://learning.edx.org/course/course-v1:HarvardX+CalcAPL1x+2T2022/block-v1:HarvardX+CalcAPL1x+2T2022+type@sequential+block@59bcd650003849ebb170c5159b1451be/block-v1:HarvardX+CalcAPL1x+2T2022+type@vertical+block@24db8c070be4492388743848efd79962) by Harvard on applied Calculus seems mind boggling. It's _applied_ calculus. Not theory. All application. Exactly what I want to refresh the foundation of my math. Here's a screenshot of what you'll learn!

### **Best of what I read**
[This tech](https://www.theverge.com/2023/5/19/23729633/ai-research-draggan-manipulate-images-click-and-drag) the most impressive thing I’ve seen recently on AI. Make sure you play the videos.

### **Book I'm reading**
Whenever I read for Dieter Rams, I'm in awe of the excellence of his work. The simplicity and the effortless design he produced in Braun is unparalleled. Maybe only by Sony and Johny Ive by Apple (by which [Dieter Rams is actually an idol for Johny Ive](https://edition.cnn.com/style/article/dieter-rams-film-exhibition-style-intl/index.html).) I'm reading another book about him, [Less but Better,](https://www.amazon.co.uk/Less-but-Better-Dieter-Rams/dp/3899555252) which is structured as German with English translation side by side. Look at the shear simplicity - given the complexity involved.

### **Best buy of the week**
A really delicious watermelon. A simple delight.
### **1 thing that is making me happy.**
The summer. I'm a Syrian. A mediterranean. And I simply love the sun. My whole mood is way better in the summers. It's a daily gift for me.
### **1 thing that is making me angry.**
Feeling the lack of focus. Wrote above about getting back to basics. Will continue that on next week.
### **What did I love when I was a kid? I should do more of.**
Legos and drawing. And I'm trying to find away to get back at them because I still love them. Whether via a product or a game I still don't know. I've tried this before around 9 years ago with [this game, TheX](https://mohammadshaker.com/games/thex-on-android/).
### Something I Changed My Mind About.
- URL: https://mohammadshaker.com/en/blog/something-i-changed-my-mind-about
- Date: 2023-05-14T00:00:00.000Z
- Tags: quick-read, Human-Written
When dealing with people, I should sometimes aim for the heart, not the mind. Logic works, but only after building trust.
#### Content
When dealing with people, I should sometimes aim for the heart, not the mind. Logic works, but only after building trust. Especially at work, and especially with senior colleagues who have strong opinions. For me they are principal engineers, managers, or other directors. Wife may be included here too.
Unless they trust me, they wouldn't listen to what I'm saying. And unless I gain their heart and trust, I won't be able to influence them. To get anyone to listen, I first need to get them to trust that I'm on their side. That whatever I'm talking about or proposing is for their own benefit, not mine.
What I say doesn't always matter. How I say it matters more. I sometimes feel that I get too logical way too early in the conversation. I learned this the hard way and I need to remind myself of this. You can never get to people by logic alone.
### Week 12: I’m Remembering. My Father.
- URL: https://mohammadshaker.com/en/blog/week-12-something-im-remembering-from-the-past
- Date: 2023-05-14T00:00:00.000Z
- Tags: Brain dump, Human-Written
Here we go again. This is what's been in my mind for the last week.
#### Content
Here we go again. This is what's been in my mind for the last week.
## I’m remembering my father.
How I'm looking like my father, pictured above, both inside and out. I can see how I'm looking like him on the outside. And I can see how I've been doing things like him - especially for the last couple of years. He passed away 2 months ago. And it's good that I'm remembering him more often. I'm picking up his tenacity in doing things. I've never, ever, seen my father overweight in my entire life, and it seems that I've obesity-phobia by training 6-7 times a week in the last couple of months, and regularly for the last few years. I still indulge in feasts way too often, just like him. But I'll always balance that with very long walks or diligent exercise. Again, just like him.
### Best of what I read.
My friend, Zaher Wanli, sent me this photo. From 2011 it seems. It's appereantly from “The Magic of Thinking Big” (a good book.) A book from the small library I had back in my home in Syria. Indeed..
> Life is too short to be little - Benjamin Disraeli

### **Something I keep using**
My watch, a Garmin Fenix 7. Bought it at the beginning of this year. Had an Apple Watch for the last 4 years. The Garmin is good, but pricy. Thought it would be a bigger update. You have to be so hardcore about exercise for Garmin to make sense. Otherwise, stick with Apple Watch or sth else.
### **Best of what I watched**
“Air” on Amazon Prime. At last a good calm movie this year. A terrible bunch of movies release in the last couple of years that I lost hope on anything new.
### Something I succeeded in doing
Playing tennis without much pain. Don't know if it's because of this guy, [the KneesOverToes guy](https://www.youtube.com/c/TheKneesovertoesguy). Still too early to know. But basically I'm doing 10 min of backward walking on a treadmill machine before I start my exercise.
### Week 11: What I'm Pondering On.
- URL: https://mohammadshaker.com/en/blog/what-im-pondering-on
- Date: 2023-04-27T00:00:00.000Z
- Tags: Brain dump, Human-Written
I used to do this 8 years ago in Arabic. A digest of what I did in the last week that takes 2 min to read.
#### Content
I used to do this 8 years ago in Arabic. A digest of what I did in the last week that takes 2 min to read.
I thought of being this back now in English. Helps me keep a diary of my weeks.
## **Something I keep doing**
Exercise for 5 days a week. Trying to keep my walking AVG > 9000 steps a week. Setup is:
1. 1 session of Zone 2. I'm aiming to target 180 minutes of Zone 2 with 2 sessions a week atleast from now on. Via [Peter Attia](https://www.google.com/search?q=Peter+Atti&rlz=1C5CHFA_enGB883GB883&oq=Peter+Atti&aqs=chrome..69i57j35i39l2j0i20i131i263i433i512j0i131i433i512j69i60l3.145j0j7&sourceid=chrome&ie=UTF-8).
2. 1 session of VO2 Max. 30 minutes. 8 intervals of 4 by 4. Via [Peter Attia](https://www.google.com/search?q=Peter+Atti&rlz=1C5CHFA_enGB883GB883&oq=Peter+Atti&aqs=chrome..69i57j35i39l2j0i20i131i263i433i512j0i131i433i512j69i60l3.145j0j7&sourceid=chrome&ie=UTF-8).
3. 4 Strength workouts.
4. Long walks on Saturday and Sunday. >15,000 steps each.
5. Now the summer is coming I can ramp up more on the daily walks. Weather is nice. Better to have the breaks outside.
6. If I'm to combine Strength and Zone 2 workouts in the same day, I have to do Strength first. Via [Peter Attia](https://www.google.com/search?q=Peter+Atti&rlz=1C5CHFA_enGB883GB883&oq=Peter+Atti&aqs=chrome..69i57j35i39l2j0i20i131i263i433i512j0i131i433i512j69i60l3.145j0j7&sourceid=chrome&ie=UTF-8).
### **Something I keep using**
My strength working app, "Workout". It's good, not the best.

I hate the UX of this screen though. Can you figure out why?

### **Best of what I listened to**
Tim Ferris interview with Nick Kokoans, [Link](https://tim.blog/2020/05/17/nick-kokonas-2-transcript/). It's during Covid early days. I like Nick. His insights and antifragility showed as a restaurant business owner and experience-maker. Talks about his trail of thoughts managing one of the most famous restaurant, Alinea, with $300+ a table to selling $35 a meal. A really interesting conversation.
### **What I'm reading?**
A book on Arabic literature and history. الجامع في تاريخ الأدب العربي ل حنا فاخوري. My father-in-law recommended this. It's good. I like reading books in Arabic again.
Earlier this year I was reading Voyage of the Beagle and On the Origin of Species for Charles Darwin. He was particularly sad in his biography that he didn't spend more time on poetry, literature and the arts.
> “The loss of these tastes \[for poetry and music\] is a loss of happiness, and may possibly be injurious to the intellect, and more probably to the moral character, by enfeebling the emotional part of our nature.”
>
> ― **Charles Darwin,** [The Autobiography of Charles Darwin, 1809–82](https://www.goodreads.com/work/quotes/179213)
I'm trying to get back more into this. And it's an absolute joy to read something like this for Imru' al-Qays امرؤ القيس:

### **Something I failed at**
Fixing my knee to run again. Discovered [this guy](https://www.atgonlinecoaching.com/), the KneesOverToes guy by [Huberman](https://hubermanlab.com/)'s interview with Tim Ferriss. And I've started adding knee exercises to my weekly setup.
### **Best buy of the week**
I love [these VBall 0.5 Pilot pens](https://www.pilotpen.co.uk/en/v-ball-05-liquid-ink-rollerball-pen-fine-tip.html?master_product_id=804) since I was in high school. I got back at using them 5 years ago. I still do. I love fine tips and they are really good.
### **Something I'm learning**
Politics and Philosophy. An easy intros for these are DK's guides like [this one](https://www.dk.com/uk/book/9781405353298-the-philosophy-book/). Am also watching [Yale's course](https://www.coursera.org/learn/moral-politics) on the Moral Foundations of Politics. It's good. More on this next week.
### Quick Thoughts on Providing Feedback
- URL: https://mohammadshaker.com/en/blog/quick-thoughs-on-providing-feedback
- Date: 2023-01-27T00:00:00.000Z
- Tags: startups, entrepreneurship, Human-Written
Before I write a long feedback for a colleague or a direct report, I always try and reread some of the books or articles that helped me on the way.
#### Content
Before I write a long feedback for a colleague or a direct report, I always try and reread some of the books or articles that helped me on the way.
I’m trying to reread and finish these books before writing the full EOY feedback for [Noon's](https://noonacademy.com/) engineers and managers. Books are:
1. **How to Win Friends and Influence People** by Dale Carnegie
2. **Radical Candor** by Kim Scott
3. **Revising Prose** by Richard A. Lanham (A short book on writing. A gem.)
If you only have 15 min to spare and that is, then read chapter 1 on **Criticism** from **How to Win Friends and Influence People**. It really does help how to communicate a difficult conversation in a constructive manner.
A note I keep telling myself is:
> Feedback is not about the list of values and metrics of the company and how a person scores on each and what he should do. It’s about the way we _deliver_ the feedback. **How it _sounds, inspires, and changes_ the person.** It’s the same kind of feedback we ask ourselves to help us improve. People (engineers) owe us a detailed and inspiring review after a year’s work.
Pictured above is Hammurabi.
### On “CEO of the Year” Award
- URL: https://mohammadshaker.com/en/blog/on-ceo-of-the-year-award
- Date: 2023-01-23T00:00:00.000Z
- Tags: Startup, Human-Written
> As always, these posts are for me first and foremost to understand my own thinking. To get my thinking straight.
#### Content
> As always, these posts are for me first and foremost to understand my own thinking. To get my thinking straight. They weren’t intended for broader reach. Inspired by [Montaigne's](https://www.google.com/search?gs_ssp=eJzj4tDP1TdINkrOMWD04szNzytJzEzPSwUAQ1sGuA&q=montaigne&rlz=1C5CHFA_enGB883GB883&oq=Montaigne&aqs=chrome.1.0i355i433i512j46i433i512j46i512l2j0i512j46i175i199i512l2j46i512j0i512l2.1310j0j7&sourceid=chrome&ie=UTF-8) essays and pushed by other friends, I’m starting to share them. They represent my own views and my current state of mind.
CNN named [T-Mobile’s Mike Sievert CEO of the Year](https://edition.cnn.com/2022/12/26/investing/ceo-of-the-year-tmobile-mike-sievert/index.html). I don’t know the man, nor do I followup with what T-Mobile is doing. T-Mobile’s CEO may be truly one of the best. But best in what? And for how long?
CNN basically focused on **1 year period** - 2022. They may as well award him The Worst CEO of the Year award in 2023. No one knows when all T-Mobile strategy “investing in our customers” is flushed down the drain and replaced by “revenue at all costs.”
My problem is not with T-Mobile or CNN, but with the kind of focus by media on “the current”, “the immediate”, “the now.”
No one pays attention on the long term. On building things that matters for society for the long term. I should remind -but not kid- myself: **Great things take time**. **Great things, take time**.
I don't like many things about Amazon. But do I think that Jeff Bezos would have won that award the first 5 years after him starting Amazon? Not a chance.
The media mocked him back in the day. They called Amazon: “silly”, “Another middleman, and the stock market is beginning to catch on that fact.”

Back in the 1990s Bezos wouldn’t have been “hot enough” for CNN. They need eyeballs. And eyeballs follow what’s hot - **right now** and CNN and the media is the best orchestrators of What’s Hot **\- right now**. It should be happening _**now**, n_ot in 5 or 10 years. That’s the standard for CNN’s (and most of the media.)
I would love to see the “CEO of the Decade” Award. That’s what I’m interested in. But there is none.
## Ideas for Future Thinking and Writing
The same way the media raises someone to fame, they put that person back to drain. Same for startups. Sometimes they elate a startup, afterward they put it down. I should look at how the media was getting crazy about Theranos and many other startups when they were “hot” and compare that with their views now. It’s always the trends, the hypes, and the woos that get the attention. Something like [37Signals](https://37signals.com/) don't because they are quiet and they walk slowly.
### “I Will Not Manage You.”
- URL: https://mohammadshaker.com/en/blog/i-will-not-manage-you
- Date: 2023-01-06T00:00:00.000Z
- Tags: management, leadership, reading, Human-Written
No one likes to be managed. I don’t like to be managed. And I’m pretty sure anyone reading this doesn’t like to be managed.
#### Content
No one likes to be managed.
I don’t like to be managed. And I’m pretty sure anyone reading this doesn’t like to be managed.
The first thing I told every engineer at Noon when I joined back in April 2022 is this; “I will not manage you. And I won’t be your manager. I will be your coach.”
Basecamp’s manager of 1 idea is spot on:
> This means **we rely on everyone at Basecamp to do a lot of self-management.** People who do this well qualify as managers of one, and we strive for everyone senior or above to embody this principle fully. That means setting your own direction when one isn't given.
I like to be coached. And I love being a coach. A coach is always there on your side helping you improve. A coach simply tells you to:
1. Do more of X/Start doing X
2. Do less of Y/Stop doing Y
And then he moves away, watches, and coaches you again. He doesn’t manage every aspect of your work.
I believe that brilliant people like this. Because they know they have authority over their work. They know they have ownership of their work. And they know that they won’t be micromanaged. **They own their work.**
When I think of Noon, I don’t think we need managers at Noon. We need more coaches. Coaching can be in different shapes and sizes. Be it in a form of another engineer giving someone feedback on a PR on how to improve your code, an EM (or should I say a Coach?) discussing your career aspirations, or another team forcing more stringent quality standards.
The best managers I had in my career were coaches. They set expectations and then they move away. Thanks to Madhu Sivasubramanian for first teaching me this back in the day. And thanks to John Lusty for teaching me how to adjust my style when I moved into a director role, managing managers.
A coach is there for you. And we can always be coaches ourselves on a daily basis. I shall be a coach for someone every day. And I should be coached by someone every day.
\[_He_ is interchangeable with _She_ above.\]
## Thoughts to write about in the future
The word management is bemusing. It entails intervention. Nassim Taleb talks about iatrogenic **when a treatment causes more harm than benefit.** If someone is hired as a manager this entails “intervening.” To manage is to intervene. And I think this is nonsense. I shall write about this later.
### Aleppo, Traders, and Rotations (aka Hackamonth)
- URL: https://mohammadshaker.com/en/blog/aleppo-traders-and-rotations-aka-hackamonth
- Date: 2023-01-01T00:00:00.000Z
- Tags: startups, entrepreneurship, unity3d, Human-Written
Disclaimer: This is part of an article/post I wrote at Noon when introducing Hackamonth for 2023.
#### Content
**Disclaimer**: This is part of an article/post I wrote at [Noon](https://www.learnatnoon.com/) when introducing Hackamonth for 2023. If you like what you read and are interested in joining [Noon](https://www.learnatnoon.com/), please do reach out.
## My Two Sisters
My two sisters, doctors in the US, did rotations for around 2-4 years. And let’s say they are not the biggest fans. It’s tough. Long hours, stress, and some more stress. According to [Accumed](https://www.aucmed.edu/about/blog/what-are-clinical-rotations-in-medical-school#:~:text=Clinical%20rotations%20in%20medical%20school%20are%20assigned%20shifts%20at%20an,team%20discussions%20are%20common%20practice.), rotations are **assigned shifts for doctors at a healthcare site**. Once assigned to a site, doctors deliver supervised care individually and as a team. **Doctor rotations happen a few times across the year.**
### Introducing Rotations at Noon
We’re not doctors. And no one likes stress. And obviously, if you’re here, you’re not a doctor either (and maybe don’t want to be one.)
But, the idea of rotations is still valid. As an engineer, you’ll amass enormous knowledge by seeding into different teams for a limited time. The teams will also benefit from all the great engineering minds that come through it.
We want to grant engineers the opportunity to seed into a different team of interest, for a month, learn from them, and give them a new perspective, and back to the original team. Or stay with the new team if they love it so much,
> 💡 This can be the seed of moving to more sparse, self-directed, self-assembled teams at Noon —something like the Valve model with 0 managers and desks with wheels (a really great read is Valve’s “New Employee Handbook” [here](https://www.valvesoftware.com/en/publications) or this [blogpost](https://discover.hubpages.com/business/Valve-the-Billion-Dollar-Company-With-No-Managers)). Basically, we can move toward a model where we have a constant feed of product problems or OKRs and teams self-assemble around them, get them solved (delivered), and dissolved. Yet to assemble a new team around a new problem. We’re not there yet and we’ll take it a step by step.
[John Lusty](https://www.linkedin.com/in/johnlusty?miniProfileUrn=urn%3Ali%3Afs_miniProfile%3AACoAAABy2lcB0b35TKJZp2MCTvrQd-p6T7TXkFU) wants to brand this as Hackamonth. I simply prefer Rotations and I’ll take my chance of coining it as Rotations here as much as I please.
### Teams as Cities on a Silk Road, Engineers as Traders
The idea of rotations reminds me of Aleppo, a famous city in Syria - one of the oldest cities in the world and one of the most prominent cities on the silk road back in the day.

Silk Road and Aleppo. Aleppo is on the middle left under Turkey. Credits [Stanford](https://spicestore.stanford.edu/collections/geography/products/wall-map-silk-road). Damascus, where I’m from, is just below Aleppo, on the other extension of the Silk Road.
Here’s a snippet about Aleppo from [The History of the Silk Road](https://www.pilotguides.com/study-guides/history-silk-road/):
> Aleppo: A pivotal trade centre of the Silk Road, Aleppo's close proximity to the Mediterranean Sea and the Euphrates Valley ensured its importance as trade between East and West opened up. Aleppo was famous for its vast 13 KM bazaar, which served as an important trading centre for commodities and cultures.


Aleppo Souk, [Credits Trip Advisor](https://www.tripadvisor.co.uk/Attraction_Review-g295416-d324950-Reviews-Aleppo_Souk-Aleppo_Aleppo_Governorate.html). Note the old, very old ceiling.
### **We want engineers to be traders, not locals.**
We want to shake things in 2023, to dare to experiment more. It's OK to make mistakes that leads to better outcome. To learn from them. To iterate and be better.
> 💡 **What Aleppo did is setting itself on the path of maximum contact with the outside world. The people of Aleppo learned that the more you increase the surface area of contact with other traders, the more you benefit from trade. Same for engineering teams: the more a team increases its contact with new engineers the more it benefits from their experience, skills, and way of work.**
### What’s the purpose of Rotations?
Over time, we want to achieve this for product engineering:
- **More playful work and end-to-end ownership.** As an engineer, you don’t only have ownership of your work and delivery, but also ownership of **where you want to be** and **what problem you want to solve**. You know more about our product (e.g. learn how LJ and N+ should be tied closely in the student journey), generate more ideas (e.g. what’s a hot lead at Noon and how can we design a LJ for him/her in order to convert in N+) and learn about other technologies (nGraph, CI/CD pipeline, data pipeline, Android, iOS, etc.)
- **Placing Everyone at the Intersection of Passion and Talent.** It’s when people have the opportunity to work on something they’re passionate about and something they’re good at that the magic really happens. We want to give everyone the opportunity to work on something that they love, they’re talented at, and that the world needs. These are the ingredients for purpose and meaning.
- We want everyone to discover purpose and meaning in their work and their life. This is when work stops being a chore, stops just being a way to earn money, and becomes something more important, work becomes something that you want to do rather than something you need to do.
- If you want to do something for an intrinsic reason rather than need to do something for an extrinsic reason you’re going to enjoy it more and you’re going to care more about the quality of the product you create.
- **No top-down structure.** You have the choice of what problems you want to work on and what things of interest you want to pursue.
- **No single point of failure.** Currently, if a backend engineer moves to another team, the whole team collapses because no one is there to do backend work. Enabling rotations does enforce our engineering teams to remove this dependency on particular engineers. This increases redundancy in engineering teams and pushes the EMs to work hard on creating redundancy - a pillar for rapid and constant delivery for high-performant engineering teams.
### **What a Hackamonth at Noon is.**
At Noon, and over the longer run, rotations will be part of our product engineering process, not an ad-hoc, 1 off thing. Here’s what we’ll do:
1. The engineer will self-initiate the rotation. As an engineer, you decide **what** you want to learn and what team you want to hackamonth with.
2. We’ll take rotations on monthly basis. At the beginning of each month, we’ll open the poll for anyone who wants to do a rotation with another team. A hackamonth for an engineer to go to another team, work, and learn from them.
3. After the hackamonth is finished, the engineer goes back to his/here original team. If he/she likes the new team so much, he/she can continue with the new team.
4. First come, first served. The first engineers to submit their request on the poll will get their wishes granted.
### **What a Hackamonth at Noon is Not.**
We’re gonna evolve this program over the coming months. Some of the stuff we laid out here will change and we can experiment with that. Nothing is set in stone. We have to start somewhere.
_For now_, we’ll limit:
1. The total # of engineers who can change teams within any calendar month. 1 engineer per team.
2. The # of rotations per engineer per year. Limit it to 1 rotation a year _for now_.
3. Duration: 1-month hackamonth.
### When can I do my first Rotation?!
The new year - Jan 9th, 2023.
### أسبوع 10: رُخام
- URL: https://mohammadshaker.com/en/blog/%d8%a3%d8%b3%d8%a8%d9%88%d8%b9-10-%d8%b1%d8%ae%d8%a7%d9%85
- Date: 2021-08-20T00:00:00.000Z
- Tags: Brain dump, Human-Written
السلام عليكم ورحمة الله وبركاته، عودة بعد 3 أسابيع. لننطلق! --شيء أداوم على فعله رياضة ٤-٥ أيام في الأسبوع مشي لساعة كل يوم على الأقل،
#### Content
السلام عليكم ورحمة الله وبركاته،
عودة بعد 3 أسابيع. لننطلق!
---
**شيء أداوم على فعله**
- رياضة ٤-٥ أيام في الأسبوع
- مشي لساعة كل يوم على الأقل، بما أنني أعمل من المنزل :)
- تناول خضراوات طازجة فقط ضمن ٥ أيام. Cheat Days في اليومين الباقيين
---
**أسوء ما شاهدته**
- [Spirited Away](https://www.google.com/search?q=spirited+away&rlz=1C5CHFA_enGB883GB883&oq=Spirited+Away&aqs=chrome.0.35i39j0i433i512j46i433i512j0i512l2j0i433i512j69i65l2.2735j0j7&sourceid=chrome&ie=UTF-8): هذا الفيلم وكل هذه الضجة لماذا؟ لم أحبه أبداً.
---
**ماذا أقرأ؟**
بشكل رئيسي هذه الكتب الثلاثة:
- [Thinking and Deciding](https://www.google.com/search?q=thinking+and+deciding&rlz=1C5CHFA_enGB883GB883&oq=Thinking+and+Deciding&aqs=chrome.0.0i355i512j46i512j0i512l3j0i22i30l5.250j0j7&sourceid=chrome&ie=UTF-8): كتاب بأسلوب أكاديمي وبغلاف أكاديمي يذكّرني بكتب الجامعة التي لا تُقرأ (لسبب جيد.) هذ الكتاب معاكس تماماً للمقولة السوريّة: من برا رخام ... إلخ. كتاب رائع لحد الآن ولكنه سيأخذ الكثير من الوقت حتى ينتهي. هو مرجع ولهذا سأستمتع به قليلاً قليلاً. من تزكية [نسيم طالب](https://www.google.com/search?q=NAssim+taleb&rlz=1C5CHFA_enGB883GB883&oq=NAssim+taleb&aqs=chrome..69i57j69i59l2j46i433i512j0i512j0i433i512j0i512j69i60.2438j0j7&sourceid=chrome&ie=UTF-8)/Nassim Taleb.
- [Structures or Why Things Don't Fall Down by J.E. Gordan](https://www.google.com/search?q=Structures+by+J.E.+Gordan&rlz=1C5CHFA_enGB883GB883&oq=Structures+by+J.E.+Gordan&aqs=chrome..69i57j46i13j0i13.174j0j7&sourceid=chrome&ie=UTF-8): مهما كتبت الآن عن هذا الكتاب، لن أوفيه حقّه. كتاب عن الفيزياء. يجب أن أكتب عنه في منشور خاص لأنه أمتع كتاب علمي قرأته لحد الآن. مكتوب بأسلوب رائع، قريب لأسلوب [نسيم طالب](https://www.google.com/search?q=NAssim+taleb&rlz=1C5CHFA_enGB883GB883&oq=NAssim+taleb&aqs=chrome..69i57j69i59l2j46i433i512j0i512j0i433i512j0i512j69i60.2438j0j7&sourceid=chrome&ie=UTF-8). من تزكية Elon Musk.
- [As Little Design as Possible](https://www.google.com/search?q=As+Little+Design+as+Possible&rlz=1C5CHFA_enGB883GB883&oq=As+Little+Design+as+Possible&aqs=chrome..69i57j46i512j0i512l2j0i22i30l6.139j0j7&sourceid=chrome&ie=UTF-8): هذا الكتاب موافق لمقولة: من برا رخام ... إلخ. على قدر محبّتي لـ [Dieter Rams](https://www.google.com/search?q=Dieter+Rams&rlz=1C5CHFA_enGB883GB883&oq=Dieter+Rams&aqs=chrome..69i57.193j0j7&sourceid=chrome&ie=UTF-8) على قدر عدم محبّتي لكاتبة هذا الكتاب. كتاب غال لا يستحق ثمنه ولا يُحق بـ Dieter Rams أبداً. سأكتب في منشور لاحق عن فلسفة Dieter Rams في تصميمه لمنتجات Braun في 1970s and 1980s.
- بالمناسبة Dieter Rams من الأشخاص الذين أثّروا عليَّ كثيراً فيما يخص التصميم السلس، Simple, minimal في كل التطبيقات التي عملت عليها.

---
**تطبيق استخدمه بكثرة**
- [Loomly](https://www.loomly.com/)
- تعرفت عليها لجدولة المنشورات على جميع صفحات تطبيقاتنا. وجدتها، لحد الآن، أفضل من Hootsuite أو Buffer.
- [News Feed Eradicator](https://chrome.google.com/webstore/detail/news-feed-eradicator/fjcldmjmjhkklehbacihaiopjklihlgg?hl=en)
- ولا أروع. بما أنني أدخل فقط للفيسبوك لمداركة ماذا يجري في صفحات تطبيقاتنا، فأنا لا أريد أن أرى أي شيء من الـ Feed. ولهذا هذه Google Chrome Extension التي تجعل الـ Feed ناصع البياض. فقط ادخل على الفيسبوك واختر أنت ماذا تريد فعله، وليس العم مارك. ببساطة تظهر لك الصفحة هكذا.
- يمكنك استخدامها لأي منصة تواصل اجتماعي أخرى. ويمكنك إيقافها أو تفعيلها متى أردت.

---
**Shameless Promotion / شيء نجحت بتحقيقه**
- مرة أخرى هذا Shameless Promotion.
- سعيد بوصولنا لنكهة خاصة لكل تطبيق من تطبيقاتنا على حدا ولمنشوراتنا ولتصميماتنا لكل تطبيق على حدا. لحد الآن [تطبيق أمل](https://alphazed.page.link/get-the-app) هو التطبيق المتاح للجميع. هذه مثال عن منشوراتنا في آخر أسبوعين. شكراً لروضة وعلا لمساعدتي في ذلك.
- مثلاً، صفحة [تطبيق أمل](https://www.instagram.com/amal.the.multilingual/) التالي مغايرة عما سيلي في تطبيق ثريا القرآن


- وهنا [تطبيق ثريا القرآن](https://www.instagram.com/thurayya.quran/)


- يمكن معرفة تطبيقاتنا الجديدة على صفحة [Alphazed](https://www.facebook.com/TheAlphazed) الرئيسية وعلى [الموقع](https://thealphazed.com/) قريباً لكل التطبيقات معاً.
- - [تطبيق أمل: طفل يتعلم 3 لغات في 3 أشهر](https://www.facebook.com/amal.the.multilingual/?__cft__[0]=AZUVXlrofSqb1OtSIMmY3YaUyfdn2rdaKASJmuGrKqbqM7QpGM4HuR7oLOXSANeGmJYOt-eFOw1RlWOuuboAYYF6ZHrShN0UHV2yT6yITtTgZpbQcA6HS1JCogIahc-tAiINmVyZxpqbbcSyr-3YbyH23mpkNwsSVpTMpi3a8ek75Q&__tn__=kK-R) لتعليم تعدد اللغات عند الاطفال، ٤-٨ سنوات، وهو متاح الان للجميع!
- [ثريا القرآن: تلاوة القرآن للفتيات والفتيان](https://www.facebook.com/thurayya.alquran/?__cft__[0]=AZUVXlrofSqb1OtSIMmY3YaUyfdn2rdaKASJmuGrKqbqM7QpGM4HuR7oLOXSANeGmJYOt-eFOw1RlWOuuboAYYF6ZHrShN0UHV2yT6yITtTgZpbQcA6HS1JCogIahc-tAiINmVyZxpqbbcSyr-3YbyH23mpkNwsSVpTMpi3a8ek75Q&__tn__=kK-R): لتعليم الاطفال القرآن الكريم للأطفال ٤-٨ سنوات. سيصدر قريباً.
- [مونتيسوري ألفازد: تطبيق ما قبل المدرسة تبعاً للمنهاج العربي البريطاني](https://www.facebook.com/alphazed.montessori/?__cft__[0]=AZUVXlrofSqb1OtSIMmY3YaUyfdn2rdaKASJmuGrKqbqM7QpGM4HuR7oLOXSANeGmJYOt-eFOw1RlWOuuboAYYF6ZHrShN0UHV2yT6yITtTgZpbQcA6HS1JCogIahc-tAiINmVyZxpqbbcSyr-3YbyH23mpkNwsSVpTMpi3a8ek75Q&__tn__=kK-R): مونتيسوري للأطفال الصغار ٢-٥ سنوات. سيصدر قريباً.
---
**سؤال لك، قارئ المنشور**
بصراحة، الكتابة هنا ضمن Wordpress كارثية تماماً باللغة العربية من ناحية التنيسق والأريحية. أفكر في أن أنقل منشوري الأسبوعي من هنا إلى Newsletter يمكن لأي كان الاشتراك بها على البريد الإلكتروني هل تحب ذلك؟ إن كان جوابك نعم. فقط أرسل لي على mohammadshakergtr@gmail.com أو ضع نعم في التعليق.
---
**عرض عمل!**
هذا عرض عمل سريع. إن كنت في سوريا، فنحن بحاجة إلى شخص Junior/Intern للعمل معنا ضمن نطاق التصميم والسوشال ميديا ضمن [Alphazed](http://thealphazed.com). ستعمل معي عن قرب لتوليد محتوى السوشال ميديا وتصميمها ونشرها. خبرة سابقة في Illustrator/Photoshop غير مطلوبة، لكنّ حبك للتصميم مطلوب. إذا كنت جاداً فسنعلمك هذه البرامج. يجب عليك فقط أن تكون "شغيل" و "حويص" و"لا تحتاج إلى الملاحقة"! العمل حوالي ١٠-٢٠ ساعة في الأسبوع. إذا كنت مهتماً، أرسل لي بريداً إلكترونياً بعنوان Junior/Intern Designer على mohammad@thealphazed.com مع ذكر ٣ أمور تحبها في نفسك وطريقة عملك! إذا كان لديك نماذج أعمال سابقة ولو بسيطة سيكون ذلك أفضل.
---
أراكم!
### أسبوع 9: غش وذئاب وتسويق
- URL: https://mohammadshaker.com/en/blog/%d8%a3%d8%b3%d8%a8%d9%88%d8%b9-9-%d8%ba%d8%b4-%d9%88%d8%aa%d8%b3%d9%88%d9%8a%d9%82
- Date: 2021-07-30T00:00:00.000Z
- Tags: Brain dump, Human-Written
السلام عليكم ورحمة الله وبركاته، عودة بعد أسبوع. جزء من كتابتي هنا هي لإعادة نفسي للكتابة العربية السليمة التي بدأت بفقدانها من بداية دراستي في الجامعة
#### Content
السلام عليكم ورحمة الله وبركاته،
عودة بعد أسبوع. جزء من كتابتي هنا هي لإعادة نفسي للكتابة العربية السليمة التي بدأت بفقدانها من بداية دراستي في الجامعة وبعدها من السفر خارجاً.
لننطلق!
---
**شيء أداوم على فعله**
أكل بدون سكر صناعي أو أي شيء معلب. فقط لحم وخضراوات وفواكه طازجة.
عندما قرأت لميس (زوجتي) منشور الجمعة الماضية قالت: "Ha، شو مشان!" ولذلك يجب أن أوضّح أنه يوجد Cheat day يوم السبت أو الأحد (I'm only human after all)
---
**أفضل ما شاهدته**
**عندما تقف عاجزاً لجمال الآية الكريمة من سورة النور:**
> **الله نور السماوات والأرض مثل نوره كمشكاة فيها مصباح المصباح في زجاجة الزجاجة كأنها كوكب دري يوقد من شجرة مباركة زيتونة لا شرقية ولا غربية يكاد زيتها يضيء ولو لم تمسسه نار نور على نور يهدي الله لنوره من يشاء ويضرب الله الأمثال للناس والله بكل شيء عليم (35)**
أفضل ما شاهدته هو شرح معناها من الشيخ الشعراوي [هنا](https://www.youtube.com/watch?v=2E6ir_Loxro). الأفضل من ذلك هو ما نتخيله كلُّ منا بنفسه في عقله أثناء قراءتها وفهم معناها.
---
**ماذا أقرأ؟**
أحاول في السنتين الماضيتين ألّا أقرأ أي شيء نُشِرَ حديثاً. (لماذا؟ حديث لوقت آخر.)
- [The Marketing Imagination Book by Theodore Levitt](https://www.amazon.co.uk/s?k=The+Marketing+Imagination+Book+by+Theodore+Levitt&ref=nb_sb_noss): كتاب من الـ 1983. قرأته السنة الماضية بسرعة وأحاول إعادته ببطء هذه السنة. الكتاب رائع ولا يتكلم عن التسويق فقط. وإنما عن نطاق الأعمال، المنتجات، ربطها بعلم النفس، إلخ. على الرغم من قدمه فهو يطرح الأساسيات. أفضل بكثير من الكتب الترند مؤخراً. الكتاب ليس بطويل.

هذه صفحة رائعة عن المنتجات الملموسة والغير ملموسة Tangibles/Intangibles للمهتمين:

- [Of Wolves and Men by Barry Lopez](https://www.amazon.co.uk/s?k=of+wolves+and+men&adgrpid=71931884158) لأول مرة أقرأ كتاب مثل هذه النمط. عن الحيوانات، والذئاب تحديداً. أول 10 صفحات مذهلة. قرأت حوالي الربع. أسلوبه علمي قصصي ريبورتاجي استكشافي بحثي غريب. لم أقرأ أسلوب كهذا من قبل. مازال جيداً. سأكتب أكثر عندما أنتهي.

---
**تطبيق استخدمه بكثرة**
[Zwift](https://www.zwift.com/) على الأيباد مع الدراجة في المنزل. يمكنك أن تقوم بحرق حوالي 500 كالوري في 40 minutes.
---
**شيء اشتريته وأحببته**
التفاح الأخضر الحامض قليلاً. عدت إليه من جديد.
---
**مذيبات الوقت لهذا الأسبوع**
فترة الاختبار لكل التطبيقات التي سنطلقها قريباً في [Alphazed](https://thealphazed.com/). ولكنها جزء من العمل.
---
**لغة أتعلّمها**
الألمانية مرة أخرى على [Duolingo](https://www.duolingo.com/). حالياً حوالي الـ 400 يوم متتالي مع قيامي بالغش لمرة واحدة عندما قمت (Shamefully) بالدفع لتبقى سلسلة انتصاراتي في التطبيق. كان ذلك في اليوم حوالي 240.
---
**شيء أخفقت به**
النوم. أظن أنه من الأفكار المغلوطة التي كانت لدي هي أن أتجنب النوم قدر الإمكان. تَغَيَّر ذلك في الفترة الأخيرة: أحاول النوم ٨ ساعات كمتوسط ودوماً وأبداً. لست سوبرمان. أكون في أفضل حالاتي الذهنية والنفسية عندما أنام لفترة ٨ ساعات. أكون منتجاً، مبدعاً، مُجِدَّاً مع مزاج رائع عندما أنام لـ ٨ ساعات كمتوسط.
عندما كنت في الجامعة كنت من المهووسين الذين ينامون حوالي الساعة ٣ أو ٥ صباحاً كل يوم ولفترات قليلة. هذا خاطئ وربمّا لأنني كبرت في العمر. ولكن المكاسب التي أحصل عليها عندما أنام لـ ٨ ساعات أفضل بكثير من عدمها. يوجد دوماً أيام عمل كثيفة ولكن هذا لا يعني أن تكون هي المعيار (Norm) هذا الأسبوع كان عبارة عن صعود وهبوط في مجمل الأيام. السبب سأذكره في موضوع أعم في الأسبوع القادم: عن النوم، الرياضة والعمل والطعام.
---
**Shameless Promotion / شيء نجحت بتحقيقه به**
كلامي التالي تقني بحت. فأعتذر مسبقاً من غير المهندسين.
بصراحة لست أنا وحدي. ولكن فريقنا في ألفازد الصغير جداً قام في فترة أقل من شهرين من الانتقال من تطبيق إلكتروني واحد إلى تصميم Backend و Frontend تدعم 5 تطبيقات معاً. أعتبر ذلك إنجازاً تقنياً رائعاً لنا في ألفازد مع كل تقنيات الـ DevOps التي نستخدمها لضمان أفضل جودة لتطبيقاتنا. بصراحة فخور بذلك. شكراً لكل من روضة وبشر وأحمد وآمال وعلا لجعل ذلك حقيقة. مرة أخرى هذا Shameless Promotion. يمكن معرفة تطبيقاتنا الجديدة على صفحة [Alphazed](https://www.facebook.com/TheAlphazed) الرئيسية وعلى [الموقع](https://thealphazed.com/) قريباً لكل التطبيقات معاً.
---
أراكم الأسبوع المقبل!
### أسبوع 8: عودة للكتابة بعد ٤ سنوات
- URL: https://mohammadshaker.com/en/blog/3497
- Date: 2021-07-23T00:00:00.000Z
- Tags: Brain dump, Human-Written
السلام عليكم ورحمة الله وبركاته، غياب لـ ٤ سنين. لم أصدّق ذلك. آخر منشور لي كان منذ ٤ سنوات وشهر! أنا الآن في الثلاثين من عمري والعمر كما يقال يبدأ
#### Content
السلام عليكم ورحمة الله وبركاته، غياب لـ ٤ سنين. لم أصدّق ذلك. آخر منشور لي كان منذ ٤ سنوات وشهر! أنا الآن في الثلاثين من عمري والعمر كما يقال يبدأ بالمرور بشكل أسرع كلما كَبُرَ الشخص. الـ ٤ سنين الماضية كانت الأكثر سرعة في حياتي. **أبرز 5 أحداث خلال الـ ٤ سنين الماضية** ١- تزوجت بالفتاة الرائعة لميس ٢- حصلت على وسام الموهبة الاستنثائية Exceptional Talent من بريطانيا، المملكة المتحدة وانتقلت إثرها من أمستردام إلى لندن ٣- غيرت قطاع عملي من التعليم إلى قطاع الصحة ومن ثم إلى التعليم مرة أخرى ٤- قمت بإنشاء أول شركة خاصة لي وفشلت ٥- قمت بإنشاء ثاني شركة خاصة لي مع فريق صغير رائع. تبدأ الآن بالنجاح www.thealphazed.com هناك العديد من الأشياء الأخرى التي اختلفت وتغيرت. كل سأحكي عنه في وقته. لم أكتب أبداً خلال السنين الماضية غير مذكرات لنفسي على الورق. **سأحاول الآن أن أبداً مجدداً بالكتابة ولكن بدون أي ضغوط أو مواعيد أو التزامات.** سأكتب عندما أحب وأتوقف وانقطع عندما لا أحب. هدفي فقط نشر ما أفعله خلال اليوم، أو الأسبوع أو الشهر كل فترة مع أصدقائي. كل من يأتي الأونلاين هنا في موقعي هو صديق لي. سأستخدم نفس أسلوبي في آخر مرة كتبت فيها قبل ٤ سنين مع بعض التحديثات في كل مرة. لنبدأ بما حدث خلال الأسبوع الماضي ولحد الآن. **شيء جديد أجرّبه** عدم تناول السكر أو أي شيء معلب. أتناول فقط الفواكه، الخضار الطازجة أو اللحوم. **شيء أتذكّره من الماضي** أمي <3 رحمك الله **شيء أداوم على فعله** روتين يومي صباحي جديد منذ حوالي الشهر. هو: ١. قراءة لـ ٣٥ د أو ساعة ٢. تمرين تنفس لـ ٣ دقائق ٣. تمرين شد stretching لـ ٥ دقائق أكتاف و Hamstrings ٤. قراءة صفحة من القرآن الكريم ٥. سؤال يوم. من [QA a day](https://www.amazon.co.uk/Day-5-Year-Journal/dp/0307719774/ref=sr_1_1?dchild=1&keywords=QA+a+day&qid=1627049349&sr=8-1) ٦. صفحات الصباح: ٣ صفحات كتابة مذكرات. لأي شيء في عقلك ولكن على الورق. سأحاول أن أذكرها بدقة في المرات القادمة **شيء أفكر به مطولاً** - ربي. - من أين أنا ولأين أنا. - عملي في هذه الدنيا ولهذا قمت بإنشاء [Alphazed](http://www.alphazed.com) لتعليم الأطفال لتكون عملي في هذه الحياة أمام ربي - الفرق، وخاصة في أوروبا، بين الإسلام بذاته كـ دين وبين المسلمين أو من يطلقون عن أنفسهم بالمسلمين. **شيء خجلت به من نفسي** عدم فهمي لجل آيات القرآن الكريم بالرغم من أني أرددها منذ الصغر. مشكلة هي في النشأة المسلمة مع التركيز على حفظ القرآن أكثر من فهم آياته. حتى عند فهم الآيات فالتركيز يكون على تعليم ظاهر الآية وليس مضمونها **أفضل ما شاهدته** - مرة أخرى عن الفرق، وخاصة في أوروبا والغرب، بين الإسلام بذاته كـ دين وبين المسلمين أو من يطلقون عن أنفسهم بالمسلمين. فيديو رائع لخطبة جمعة [لنعمان علي خان](https://www.youtube.com/watch?v=K7qQdRKeQJw) **أفضل ما سمعته** عدت لسماع البودكاست بعد انقطاع لـ ٢-٣ سنوات. مؤخراً قمت بالعودة إلى بودكاست [Tim Ferriss](https://tim.blog/podcast/). متاحة على الأندرويد و Apple **ماذا أقرأ؟** اقرأ الكثير من الكتب معاً كالمعتاد. اليوم أنهيت كتاب Blue Ocean Strategy الذي يتكلّم عن نطاق الـ Business وكيف يمكن للشركات الخوض في قطاعات سوقية جديدة ومنتجات جديدة بما يطلق علي Blue Ocean Strategy. كتاب جيد. أولى ٩٠ صفحة فيه جدأ جميلة ومخالفة لغيره من الكتب. في المنتصف يتحدث عن مواضيع تخص الشركات الكبيرة أكثر من الناشئة (Team Management, Vision Alignment across organization ,etc.) وفي القسم الأخير تذكير وإعادة أمثلة. الأكثر إفادة هي الـ ٩٠ صفحة الأولى. **تطبيق استخدمه بكثرة** [أمل - Amal: تطبيق تعليم اللغة العربية للأطفال](https://play.google.com/store/apps/details?id=com.alphazed.amal) ببساطة لأنه أول منتج لنا في ألفازد وبالتالي أقضي الكثير من الوقت عليه وفحصه وتحسينه إلخ **شيء اشتريته وأحببته** تطبيق Zwift App للبسكلة (Biking) مع ربطة مع [Wahoo Kickr Core Smart Trainer](https://uk.wahoofitness.com/devices/bike-trainers/kickr-core-indoor-smart-trainer) أقضي حوالي النصف ساعة عليه يومياً (عندما لا تؤلمني ركبتي.) رائع بحق حتى الآن. عليه منذ حوالي الشهرين. المشكلة أنه يجب عليك شراء الدراجة المنزلية أولاً (يمكنك شراء أنواع بسعر أرخص من Wahoo Kickr) **مذيبات الوقت لهذا الأسبوع** تويتر Twitter: بدأت باستخدام تويتر السنة الماضية. لا أطيق الفيسبوك وأظن أن تويتر أفضل بكثير بغض النظر عن أن كلاهما مضيعة كبيرة للوقت. **80/20 من وقتي لهذا الأسبوع ذهبت على..** 20%: روتينات يومية كالقراءة، رياضة إلخ 80%: عمل **لغة أتعلّمها** الألمانية مرة أخرى **مشروع أعمل عليه** أظنك تعلم الآن أنني أعمل على مشروع Alphazed مع كل تطبيقاته نصف ساعة انتهت! أراكم الأسبوع المقبل!
### Elasticsearch Out of the Box Use Cases
- URL: https://mohammadshaker.com/en/blog/elasticsearch-out-of-the-box-use-cases
- Date: 2020-08-09T00:00:00.000Z
- Tags: arabic, engineering, genre-classification, ml, nlp, research, visualization-libraries, Human-Written
Elasticsearch ships with NLP-friendly features that most teams underuse: phrase-based did-you-mean suggestions, completion-based autocomplete, fuzzy matching, and built-in text analyzers. This post surveys those out-of-the-box capabilities and how they apply directly to Arabic and multilingual search applications.
#### Content
Elasticsearch is a distributed, RESTful search and analytics engine capable of addressing a growing number of use cases. As the heart of the [Elastic Stack](https://www.elastic.co/products/) (ELK), it centrally stores your data so you can discover the expected and uncover the unexpected.
In this post, we're investigating some features and out of the box use cases for ElasticSearch in the field of NLP.
## Search Enhancement Features
ElasticSearch provides us with a sort of cool stuff to enhance our end-user search experience.
### You Complete Me
Effective search is not just about returning relevant results when a user types in a search phrase, it's also about helping your user to choose the best search phrases.
#### Did you mean ...?
Elasticsearch has a [phrase-suggester](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-suggesters.html#phrase-suggester) that can correct the user's spelling after they have searched.
Phrase-suggester selects entire corrected phrases weighted based on n-gram language models. It's able to make decisions about which tokens to pick based on co-occurrence and frequencies.
#### Suggestions **while you type**
[Completion-suggester](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-suggesters.html#completion-suggester) can make suggestions while typing. Giving the user the right search phrase before they have issued their first search makes for happier users and reduced load on your servers, these suggestions can be e.g. existing tags.
### Lenient Searching
A fuzzy search is one that is lenient toward spelling errors. To give an example, you can find `Levenshtein` when searching for `Levenstein`.
Fuzzy searches are simple to enable and can enhance “recall” a lot, but they can also be very expensive to perform.
Fuzzy matching isn't always the right tool for the job, oftentimes imprecise matches can be found through other techniques. The [phonetic analysis plugin](https://www.elastic.co/guide/en/elasticsearch/plugins/current/analysis-phonetic.html) contains a number of tools for approximating matches, such as the [Metaphone](http://en.wikipedia.org/wiki/Metaphone) analyzer, which finds words that sound similar to other words. For instance, if what is required is making sure words like `run` and `ran` are both considered equivalent. Alternatively, for checking misspellings, N-gram analysis [as described in this short tutorial](https://web.archive.org/web/20140209084956/http://exploringelasticsearch.com/book/searching-natural-language/searching-non-word-text.html) can run quite a bit faster at query time, depending on the dataset.
## Basic NLP Tasks
These are basic NLP tasks, typically, running as a part of the preprocessing step before applying more complicated NLP analysis on the data.
### Text Processing
ElasticSearch has over 20 built-in language-analyzers including ones for Arabic.
What can an analyzer do?
- Tokenization
- Stemming
- Stopword removal
Refer to [here](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-analyze.html) for the Analyze API overview, for language-specific analyzers refer to [here](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-lang-analyzer.html#analysis-lang-analyzer).
Also, the analyzer can be customized, for instance, by providing your own stopwords list.
### Language Detection
Detecting languages is a so-called "solved" NLP problem. So no need to reinvent the wheel over and over. When you already have ElasticSearch up and running, you can simply install one of its plugins like [this](https://github.com/jprante/elasticsearch-langdetect) one. And you can always provide a custom plugin.
## Advanced NLP Task
These are higher-level tasks that require more complicated setup and strategies, let's discover how can they be handled by ElasticSearch.
### Text Classification
Text classification is a task traditionally solved with supervised machine learning. The input to train a model is a set of labeled documents. The minimal representation of this would be a JSON document with 2 fields: "content" and "category".
Traditionally, text classification can be solved with a tool like [SciKit Learn](https://scikit-learn.org/0.19/datasets/twenty_newsgroups.html), [Weka](http://weka.wikispaces.com/Text+categorization+with+WEKA), [NLTK](http://www.nltk.org/book/ch06.html), [Apache Mahout](https://mahout.apache.org/users/classification/twenty-newsgroups.html), etc.
This task can be solved in a much simpler way with Elasticsearch, providing the same described input:
You just need to execute 4 steps:
1. Configure your mapping ("content" : "[text](https://www.elastic.co/guide/en/elasticsearch/reference/current/text.html)", "category" : "[keyword](https://www.elastic.co/guide/en/elasticsearch/reference/current/keyword.html)")
2. Index your documents
3. Run a [More Like This Query](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-mlt-query.html) (MLT Query)
4. Write a small script that aggregates the hits of that query by score
The MLT query is a very important query for text mining.
How does it work?
It can process arbitrary text, extract the top n keywords relative to the actual "model" and run a boolean match query with those keywords. This query is often used to gather similar documents.
So, all you need is to just run an MLT query with the input document as the like-field and write a small script that aggregates score and category of the top n hits.
With the Elasticsearch approach training happens at index time and your model can be updated dynamically at any point in time with zero downtime of your application. If your data is stored in Elasticsearch anyway, you don't need any additional infrastructure. With over 10% highly accurate results you can usually fill the first page. In many applications that's enough for a first good impression.
### Recommendation Systems
Recommendations and search are two sides of the same coin. Both rank content for a user based on “relevance” the only difference is whether a keyword query is provided.
By translating the problem of recommending content to a user into a search problem for users' implied interests, we can base our recommender system on a search ElasticSearch.
For further information, you can refer to our previous post that tackles this use case in detail.
### Social Media Monitoring
Thanks to the digital social revolution, opinions that used to be bottled up within confined media channels or the four walls of one's private life are out there for all to see.
Wouldn't it be intriguing to get insights into this public sentiment towards e.g. your brand, products, hot topics, etc. ?
We would like to get a view of the public sentiment expressed towards a specific entity, which requires three main tasks:
- Tracking the data streamed from Twitter.
- Analyzing this data stream to get the sentiment and maybe semantic features.
- Visualizing the analysis results.
To this end, [ELK Stack](https://www.elastic.co/products/) (Elasticsearch, Logstash, Kibana) can be in benefit.
The world's most popular open-source log analysis platform, instead of being used to ingest log files, can be fed with the tweets using Twitter's streaming API. On top of the aggregated data, we can create a series of graphic visualizations that best depict the Twitter trends.
[Logstash](https://www.elastic.co/logstash), the "L" in the "[ELK Stack](https://www.elastic.co/products/)", is used at the beginning of the log pipeline to ingest and collect logs before sending them on to Elasticsearch for indexing. Log analysis the most common use case, but any type of event can be forwarded into Logstash and parsed using plugins.
In the context of this task, it would be configured to deal with the Twitter stream using [this](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-twitter.html) plugin. After setting up Logstash you can configure it to track specific keywords in the tweets.
After a while of receiving the feed from Twitter, a larger pool of data from which to pull should have existed.
We can then begin to use the Kibana to search for the data we're looking for.
If we're tracking public sentiment regarding our company’s brand, for example, we could query the brand name itself and check the correlation with sentimental expressions.
Once we have narrowed the available down to the information that interests us, the next step is to create a graphical depiction of the data so that we can identify trends over time.
As an example, we can create Mentions Over Time visualization, showing mentions of the entities we're tracking over time.
Another example is creating a map depicting the geographic locations of tweets.
And these are just simple examples of what can be done with your Twitter data in Kibana!
## Conclusion
ElasticSearch is a cool search engine provides us with many simple to use features to supplement our search engine with many enhancement features.
Some NLP tasks such as syntactic parsing require deep linguistic analysis. For this kind of tasks, Elasticsearch doesn't provide the ideal architecture and data format out of the box. That is, for tasks that go beyond token-level, custom plugins accessing the full-text need to be written or used. But tasks such as classification, clustering, keyword extraction, measuring similarity, etc. only require a normalized and possibly weighted Bag of Words representation of a given document, ElasticSearch can be in benefit.
The integration of ElasticSearch with the other ELK stack components namely: Logstash and Kibana is a strong data mining and analysis stack.
## Further Reading
1.
2.
3.
4.
5.
6.
7.
8.
9.
10.
### Automatically Extracting Valuable Content from News Streams.
- URL: https://mohammadshaker.com/en/blog/automatically-extracting-valuable-content-from-news-streams
- Date: 2020-02-21T00:00:00.000Z
- Tags: almeta.io, ideas, ml, nlp, research, Human-Written
A news content aggregator pulling from 50+ sources needs more than a firehose — it needs a pipeline that scores, filters, and ranks articles by quality. The main signals are readability, informativeness, and source reliability. Combining these lets you surface the 5% of articles worth reading and suppress the rest.
#### Content
Almeta News, as a content aggregator app from ~ 50 sources - to the time we wrote this blog post - always aims to provide its users with the best quality pieces to read.
Rather than an army of content watchers and editors, Almeta is looking forward to developing the best algorithms to review content automatically, looking for indicators of quality, assessing a content’s placement.
This post is a part of our research efforts seeking the best content ranking methodology.
In this post, we're trying to determine the most effective indicators of content quality. Depending on human experts' points of view, famous competitors' plans, and scientific literature. We're as well suggesting ways for measuring these indicators automatically.
## What Constitutes Quality Content?
In this section, we're listing the most important quality content indicators, derived from human experts' recommendations to the Authors. With a brief description of how to apply these factors technically.
**1.** **Quality Content is the valuable content in the user's point of view**
What my audience really cares about is what they care about, not what I care about.
a. Cultural context
Culture is what ultimately drives people and their decision-making process. Thus, we need to make sure that the top content is relevant to our target segment culture.
b. Geographical context
This, on one hand, is related to the cultural context. On the other hand, mostly, people are more likely to be interested in what affects them directly. Thus, geographically-based content filtering may be effective to determine what a user probably cares about.
c. Personal context
People are different, what's considered interesting to someone, doesn't have to be for others.
**Technical Application**
- Studying our target segment and focusing on the types of content that attract them.
- Detecting the user's location. Filtering the content geographically, which can be considered as an automatic content classification problem. Providing the user with cultural and local preferences.
- Getting to know each user alone, by the means of:
- **Profile**: users can provide some basic information about their career, education, etc. and other additional information about their hobbies, feelings, etc.
- **Interests**: let the users choose topics, keywords, etc. they are interested in following.
- **History**: learn about the user from his behavior. What kind of content he liked, commented on positively or negatively, shared, viewed, blocked, followed, etc.
- Making the content searchable letting the user find what he exactly needs in a point.
**What kinds of recommendations to make in this context?**
- Recommend content/interests that are similar to the one the user is interested in.
- Recommend content/interests based on the user's activity; things he's not following but based on his recent activity seems he's interested in.
- Recommend content/interests similar users are interested in.
- Recommend content that helps the user improve his career, makes him feel better, and an endless list of ideas!
The recommendations here may require a mix of classification, clustering, recommendation and even rule-based methods.
**2. Quality Content is made in the way the audience likely want to consume it in**
It's not all about the substance of the content, but the structure of it, too. For instance, does an audience want to read long narratives to glean what happened at a conference?
**Technical Application**
Give the user the choice to decide about the structure of the content he wants to read e.g. sorting the content according to their lengths.
**3. Quality is determined by the audience**
If they find it useful – which we can tell from the data and feedback – it may likely be quality.
**Technical Application**
Saving data and statistics for each article and favor the ones having the best feedback.
**4. Quality Content is usually actionable**
A general rule, content that has practical use is naturally interesting—or at least more interesting than content that is not.
Favor the content that the readers can take and use in real life. After consuming the content, readers walk away with advice, steps, or insight they can apply to a job they need to get done.
**Technical Application**
This can be implemented as a kind of automatic content classification. Where the data can be gathered from content like Wiki How.
**5. Quality Content is Novel**
The ideas discussed in the content should be novel! Not rehashing the same concepts over and over again.
Also, the novelty can be looked at from the user's point of view; the novel content is the one new and interested to the user.
**Technical Application**
We need to implement a kind of novelty detection algorithm. One of the proposed solutions is the online clustering of the content, where the most novel content forms a new cluster.
**6. Quality Content in Trendy**
To make the content more interesting to the current audience, we can leverage the power of recent trends. Trends can be thought of both in a worldwide or local setting.
**Technical Application**
One of the ways to detect trends can be also using online clustering. Where trends form emerging clusters in a period of time.
**7. Quality Content is Fresh**
The importance of a piece of news changes over time. Although old content can be interesting too, considering the freshness factor aims to keep the user up with the latest updates, and omits the stale stories.
**8. Quality Content is Accurate**
Accuracy builds trust with readers.
**Technical Application**
Concluded from the tips recommended to the authors to think about when mulling over the issues of content accuracy, the following two factors can be considered to detect the accuracy:
- **The number of links to other sources and content**: The more the author can back up and substantiate what he's writing about, the more trusted his content will become.
- **Consider who the content is linking to**: Are they a trusted and authoritative source? Linking to other quality websites will earn more trust from the readers.
**9. Quality Content is Trusted**
People tend to consume content from sources they trust more.
**Technical Application**
- Let users decide which sources to follow.
- Score the sources over time, according to the quality, amount, and frequency of the content they provided, also taking the popularity into the consideration.
**10. Quality Content is Short and Pointed**
There is nothing better than a brief, to-the-point content that is filled with information. So don't focus on word count. Longer content does not mean better content. Quality content is detailed and covers the topic from many angles. But it’s tight. Not a word is wasted.
**Technical Application**
- Automatic readability scoring.
- Read time estimation.
- The length of the article.
- How diverse are the sentences from the centroid of the content.
- And maybe more...
**11. Quality Content is Engaging**
A content that engages the reader.
**Technical Application**
Based on tips recommended for the authors to make engaging content, we may consider the following factors:
- **The number of questions**: engaging content includes questions that make readers reflect on how they can implement the knowledge you provided.
- **Have an important and promising introduction**: Most people probably decide within the first few sentences if the post is worth reading.
- In the simplest setting, this can be implemented as a kind of automatic classification, where the model determines how much an introduction is relative to the "good introductions" class.
The dataset can be gathered automatically using real introductions from famous trusted sources as positive examples, while choosing random sentences for negative examples, or for smoother results using introductions from known lower-quality sources for negative examples.
**12. Quality Content Avoids Balanced View**
Quality content is a content that considers the strengths and weaknesses of alternatives and selects the most correct.
**Technical Application**
Measuring the opinion bias in the content.
**13. Quality Content Provides good communication**
Vision is the strongest human sense, so people are naturally drawn to visually appealing pieces. Whether the content uses pictures, videos, or diagrams, they can help illustrate the author's point.
Quality content is often highly visual and easy to consume.
## How Do Giant Content Providers Do This?
A deeper look at how some famous content providers are surfacing valuable content.
### Google News
Forget PageRank, Google’s news service doesn’t rely on the same algorithm used by “regular” Google, of which PageRank is a part of. Instead, Google News taps into its own unique ranking signals. This was denoted by *Josh Cohen*, the business product manager of Google News in one of his interviews on November 24, 2009.
During this interview, Cohen showed us what's under the hood of Google News ranking methodology.
#### **So, How Does Google News Work?**...
Google News structure news into story clusters; a group of individual articles that are all on a given angle to a particular news event.
**What causes an individual article to be the lead item in a particular story cluster?**...
Various factors are involved, *Cohen* said:
**1. Freshness, local relevancy & originality**
> Is there original content? The timeliness. Coverage of recent developments? The relevancy to the cluster at hand. In some cases, is there local relevance? Is there content from a local source with local content?
**2. Publication reputation**
How an individual article ranks within a story cluster is further influenced by the reputation that its publishing source carries within Google News.
> What’s the volume of publication of original content in a given category? If you look at Bloomberg and Reuters, they may have hundreds of original articles in business. That’s a pretty good indication of the quality of that source for that category. Compared to sports, there’s not that much original content [and so they might not have as much authority for when ranking sports stories].
**3. Measuring Clicks**
> what users are clicking on from the results they see. You understand who are trusted sources for users. If you go to a given cluster for Google News, you’d expect the first story to get more clicks than the second and so on.
**4. Textual Content Counts, Tool**
> Your URL, title and body are three components you can look at. If you’re weak any one of those, it puts additional weight on other categories.
Moreover, Google News has various “editions” for different countries, such as Google News UK versus Google News US. Each edition has its own particular blend of signals it uses to rank news content. Furthermore, each section within a Google News edition (such as Entertainment versus Sports) also uses its own unique blend of ranking signals.
### Flipboard
At Flipboard, they deliver over 100,000 stories per day, and here are some factors they consider to surface such good quality stories:
**1. Ranking sources**
A team of humans determines the editorial quality of a source, and then something called the domain ranker comes into play. Built for spam detection, the ranker allows the team to favor sources with known track records, who themselves follow time-honored journalistic principles. Who’s ranked and how is carefully guarded and continually reviewed.
**2. Incorporating signal from as many people as possible**
While the ranker does make it harder for stories from the long tail to surface on Flipboard, there’s another filter that influences what you see: the user satisfaction score. A set of signals that indicate how engaged people are with a piece of content, the score is a proxy for quality.
**3. Clustering stories for multiple perspectives.**
They claim that they surface the plurality of sources and voices they have by story clustering. Story clustering is an algorithmic technique they use to pull together stories from different sources on the same topic.
> Not every cluster might actually have stories with truly unique viewpoints —machine learning just isn’t there yet— but the structure gives us a framework to offer balance.
>
> Flipboard Team
**4. Attributing for context.**
All stories on Flipboard have author, publisher and/or curator attribution so the user can see where it comes from and make his own informed decisions about the person’s inherent biases.
### Medium
The ethos of Medium is inherently democratic; it seeks to give a voice to people who have something interesting to say, even if they don’t have thousands of Twitter followers, an active blog or friends in the right places. Medium is built to reward content for its quality, not for the pedigree or popularity of the author.
So while Medium allows anyone to publish pretty much anything, it works hard to guarantee that visitors only see the good stuff.
The website’s ever-evolving algorithm that determines post ranking considers a variety of factors. *Ev Williams*, who co-founded Medium with Biz Stone, explained:
> What we’re doing is ordering things by our best guess of the relative quality/interestingness of the different items—according to the people who have seen them… It’s not a direct popularity ranking. It takes in a variety of factors, including whether or not a post seems to actually have been read (not just clicked on) and whether people click the “Recommend” button at the bottom of posts. The ratio of people who view it who read it and who read it and recommend it are important factors, not just the number. (This is an attempt to level of the playing field for those who don’t already have large followings and/or a penchant for writing click-bait headlines.)
On the other hand, Medium’s algorithm prioritizes quality over the date something was published in.
Medium also allows users to personalize their experience.
## How Do They Do this in the Literature?
In terms of white papers, just a few countable works considered this problem.
[1, 2, 3] proposed ranking algorithms for news information, finding the most authoritative news sources and identifying the most interesting events.
All of the proposed algorithms share the following properties:
**1. Ranking the news sources:**
Important News articles are Clustered. An important news story is probably (partially) replicated by many sources. For instance, consider a news article n originated from a press agency. The measure of its importance is also expressed by the number of different online newspapers which replicate n, this means that the (weighted) size of the cluster formed around n is a measure of its importance.
**2. Mutual Reinforcement between News Articles and News Sources**:
We can assign different importance to different news sources according to the importance of the news articles they produce. So that, a piece of news coming from “Washington Post” can be more authoritative than a similar article coming from say “ACME press”, since ”Washington Post” is known for producing good stories.
**3. Time awareness**:
The importance of a piece of news changes over time. We are dealing with a stream of information where a fresh news story should be considered more important than an old one.
## Conclusion
In this post, we discussed a variety of factors to consider for surfacing the valuable content from a large amount of aggregated data.
We gathered these factors looking in:
- Experts recommendations for authors to write quality content.
- The methodologies followed by giant content providers.
- Scientific literature review.
How to combine these factors in a meaningful way depends on your application. You may find a way to develop a weighted sum combining all of them, or simply provide some of them as separate features.
So pick your factors, put your plan, and let your customers enjoy valuable content!
## References
[1] Del Corso, Gianna M., Antonio Gulli, and Francesco Romani. "Ranking a stream of news." *Proceedings of the 14th international conference on World Wide Web*. 2005.
[2] Mahour, Bhavana, and Akhilesh Tiwari. "A Ranking Algorithm for News Data Streams." *International Journal of Computer Applications* 94.6 (2014).
[3] Trajkovski, Igor. "Pagerank-like algorithm for ranking news stories and news portals." *International Conference on ICT Innovations*. Springer, Heidelberg, 2013.
## Further Reading
1.
2.
3.
4.
5.
6.
7.
### Abstractive Summarization in Underresourced Languages
- URL: https://mohammadshaker.com/en/blog/abstractive-summarization-in-underresourced-languages
- Date: 2020-01-29T00:00:00.000Z
- Tags: nlp, Human-Written
Abstractive summarization for low-resource languages is harder than extractive summarization because it requires generating new text, not just selecting sentences. Morphological complexity and the scarcity of training data compound the difficulty for languages like Arabic. Transfer learning from high-resource language models is the most practical path forward.
#### Content
The increasing amount of text data in the digital age calls for methods to reduce reading time while maintaining information content. The process of summarization achieves this by deleting, generalizing or paraphrasing fragments of the input text to create a more conscious version. Summarization methods can be categorized into single or multi-document and extractive or abstractive approaches. In contrast to the single document, the multi-document setup can utilize the fact that in some domains like news articles there are different sources describing the same event and thus these articles hare a lot of similarities. Extractive methods solely rely on the words of the input and e.g. extract whole sentences from it. Abstractive approaches, on the other hand, are rarelybound toany constraints and they have gained a lot of traction recently due to current advances in Deep learning and seq2seq models.
**In this article, we will concentrate our discussion on *extractive* Vs *abstractive* approaches**.
On one Hand, extractive summarizes have several benefits:
- They are typically easier to implement, as in the simplest form the extractive summarizer can be seen as sentence ranking model,
- In practice, these models are very robust and can generalize to other domains easily mainly because their training step (if there is any) does not put any restrictions on the domain or language,
- Finally, most of these approaches are either unsupervised or does not require any training and thus are helpful for under-resourced languages like Arabic
However, on the other hand, the best available models for summarization in terms of performance are abstractive summarization that relays on very large deep learning models with million of parameters. These models have much better performance than their extractive counter-parts. However, these models suffer from 2 major problems:
- Deep learning models and **Seq2Seq** models, in particular, are data-hungry, they require large training sets of parallel text-summary where these sets are usually built in a manual way, these datasets are scarce in western languages let alone underresourced languages like Arabic.
- Furthermore, these approaches struggle when working with text from domains other than their training set. And the options for model adaptation in this task is rather limited (in comparison with other tasks tat utilizes Seq2Seq architecture like machine translation or automatic speech recognition)
In the case of Arabic language for instance [in our previous article](/blog/can-you-measure-a-text-informativeness-using-its-summary) we have already explored the available freely available summarization corpora and unfortunately, there is no Large training set to enable neural abstractive models.
In this article, we will explore several methods in which we can implement an abstractive summarizer in an underresourced language where no large parallel data is available.
**We try to tackle this task in 4 different ways:**
1. Easily building large datasets to support Seq2Seq supervised models
2. Using non-neural abstractive summarizers
3. Using unsupervised neural summarizers that require no annotated data
4. Using other miscellaneous approaches
## Building Datasets for Neural Models
One option to allow the implementation of abstractive summarizer is building a large enough dataset to enable the training of such models. However, the process of manually creating such a huge can be extremely costly. However again, if there is a way to generate such a dataset in an automatic or semi-automatic manner in the under-resourced language, abstractive summarizers can be easily created.
The only method we found that tackled a similar task is reported in [1]. In this article people at Google Brain used the references of English Wikipedia articles as input and trained a Seq2Seq model to generate the actual article from the references. The goal of this article was not summarization in its own right but rather to test the neural models in real settings.
Apart from the complexity of parsing all the references in Wikipedia articles, this approach might be hard to generalize to other languages either due to the limited number of references. Also, for example, in languages like Arabic, a good portion of references are not Arabic but English.
## Non-Neural Approaches
Research in abstractive summarization goes beyond current neural methods and some early Non-neural methods exist in this section we will outline some aspects of these methods.
In nearly all of these non-neural approaches the summarizer can be decomposed in 2 steps:
1. Content Selection
2. And Surface Realization
Content Selection aims to select a subset of the candidate phrases extracted from the text for inclusion in the final summary. Typically, subject to length constraints.
Some methods relay on heuristics for instance, the model in [2]
heuristically selects the candidate phrases most frequently mentioned
for an aspect.
However, the preferred method is Integer Linear Programming (ILP). ILP can be used to optimize an objective function subject to a set of linear constraints. When applied to content selection, the objective function is a weighted sum of a set of binary variables. Each variable represents a candidate phrase and has the value 1 if and only if ILP decides to select it for inclusion in the final summary. The weight associated with each variable indicates the importance of the phrase. Authors in [3] for instance, estimate the salience of each candidate phrase based on its position and its grammatical role in the input document and use the salience score as its weight. The linear constraints encode length constraints. e.g. one constraint limits the number of words in each sentence in the summary. *The key advantage of employing ILP for content selection is that the decision is made jointly based on all phrases.*
On the other hand, Surface Realization aims to combine the candidates selected in content selection using grammatical/syntactic rules to generate a summary. Tthis part usually includes complicated *Natural Language Generation* steps. And while some tools like [SimpleNLG](https://github.com/simplenlg/simplenlg) can be utilized for the English language, work in NLG on other languages is still extremely limited especially for under-resource languages like Arabic.
### Some Approaches
On of the Prominent early approaches are Template-based *methods. Template-based methods are motivated by the observation that human summaries of a given type (e.g., meeting summaries) have common sentence structures, which can be learned from the training set and encoded as templates. Then given an input document, a summary can be generated by filling the slots in the best-fitted templates learned for this type of documents.* Template-based methods typically consist of three steps:
1. Learning the templates from the human summaries
2. Extracting important phrases from the input document
3. Generating a summary based on the filled templates.
For example, in [4], they propose a template-based method for meeting summarization:
1. In step 1 (template learning), a template is first generated from a sentence of each human summary in the training set by replacing each Nominal Phrase (NP) in the sentence with a blank slot that is labeled with the hypernym of the NP’s head using WordNet. Then, these templates are clustered based on their root verbs.
2. In step 2 (keyphrase extraction), the important phrases for each topic of the input document are extracted and labelled with their hypernyms.
3. Finally, in step 3 (generation), the templates with the highest similarity to each topic of the meeting are selected. Then candidate summary sentences can be generated by filling each template with matching labels. With a sentence ranker is trained to rank the generated sentences in each segment. The highest ranked sentence for each topic segment will be selected for inclusion in the summary. Finally, The selected sentences are sorted by the chronological order of the topic segments in the input document.
Other early methods encode the text in a graphical way like event semantic link networks (ESLN) [5]. In this approach, given an input text, a graph is constructed where each node corresponds to an event mentioned in the input text. An edge between two nodes encodes the semantic relation between the corresponding events. After graph construction, ILP can be applied to this network to perform content selection (i.e., selecting a subset of nodes for generating the summary.) Then from such intuitive representation, simple NLG can build the final summary.
### Pros and Cons
**Pros:**
Most of these methods either use unsupervised machine learning or are fully knowledge-based systems *meaning they require very little to no training data.*
**Cons:**
Many of these methods relay on other NLP steps like key-phrase extraction and syntactic parsing this dependence causes several issues:
1. Propagation of errors from sub systems to the summarizes
2. Most of these per-processing steps are challenging in their own right and can have a low performance on under-resourced languages
3. The integration of these subsystems is usually done manually with a lot of tinkering
4. The complexity of such systems could impact their speed
5. The performance of such models are much lower than their neural counterparts
## Unsupervised Neural Approaches
These approaches aim at securing the gains of using a neural model, the simplicity and the performance without the main problem of data shortage. Following are some of the methods reported in the literature:
Most of these methods relay on some form of auto-encoders while we
present an introduction to Auto-encoders here, detailed study of
auto-encoders is beyond the limits of this article yet you can follow
[this
series](https://towardsdatascience.com/auto-encoder-what-is-it-and-what-is-it-used-for-part-1-3e5c6f017726) for a simple intro with code, or read [this
article](http://ufldl.stanford.edu/tutorial/unsupervised/Autoencoders/) if you wish to better understand the math behind
auto-encoders.
An autoencoder is a type of artificial neural network used to
learn efficient data codings in an unsupervised manner. The aim of an
autoencoder is to learn a representation (encoding) for a set of
data, typically for dimensionality reduction. The basic architecture
(shown ing the figure below) is similar to the encoder-decoder
systems with one main difference, Autoencoders (being Auto) use the
same text for input and output of the model.
In the case of abstractive summarization the auto-encoders are trained to shrink the size of the input text there are several ways this is accomplished:
In [6] the first-ever unsupervised abstractive summarizer is introduced, in this article the goal is performing multi-document summarization on customers reviews. Where given a set of customer reviews the model returns a single review that is supposed to include all the information of these reviews. The model is depicted in the following figure, it is composed of 2 Auto encoders the first is used to learn encodings of the input reviews, these representations are then averaged and used in another autoencoder that produces the final summary.
The whole model is trained in a single run to optimize both the Average Summary Similarity and the Auto-encoder Reconstruction Loss. The results are far from optimal. Yet, they are comparable to the other abstractive approaches. The following figure compares the results of this unsupervised model with a strong extractive baseline. The code of this publication is [available here](https://github.com/sosuperic/MeanSum)
Another Approach is reported in [7] In this work, the authors train an auto-encoder to learn representations of the input text, but use a technique from [8] to restrict the length of output of the auto-encoder to a predefined length, the authors report their results of the annotated Gigaword dataset which is one of the most famous corpora in English and use ROUGE as their metric, if you are not familiar with ROUGE [consult this article](https://rxnlp.com/how-rouge-works-for-evaluation-of-summarization-tasks/#.XdFmcNFS97A). The authors report a rouge score of 23.41, which compared to the score of 39.11 reported in the state of the art model on this task.
Other unsupervised abstractive summarizers include [9]
## Other Related Tasks
In this section, we outline some tasks that we believe are close enough to abstractive summarization yet they fall outside the scope of this article.
### Sentence Compression
Sentence compression is a paraphrasing task where the goal is to generate sentences shorter than given while preserving the essential content. A robust compression system would be useful for mobile devices as well as a module in an extractive summarization system. The Largest data set for this task was built automatically by [10]. The data is generated from news articles by utilizing the first sentence S and the header H. First they apply several filters on S and H to exclude articles with grammatically or semantically problematic headlines (click-bait stuff). Next, the headline is used to create a more condensed version of the first sentence. It is possible to extend such automatic to other languages like Arabic, yet the extension is not trivial.
### Cross-Lingual Abstractive Summarization
In [11], [12] the authors suggest nearly the same idea, basically, they start with a large summarization dataset usually found in English language, then they do a round trip translation as follows:
- To build the training data for a given underresourced language X, the articles of the corpora is translated from English to language X using a neural machine translation system (google translate for example). This results in noisy translation in language X, next, this noisy translation is re-translated back to English, in the same way, resulting in even more noisy articles.
- For Training the model is trained to use the noisy round trip translations articles as input and output the original clean English summaries.
- In deployment: for a given article in language X, the text is first translated to English using the same neural translator and then the summarizer can use this translation to generate English summary.
This can allow for summarizing articles in under-resourced languages, however, the summaries are in English which is not optimal for most cases.
### Automatic Text Paraphrasing
One of the main shortcomings of extractive summarization in real applications is the legal constraints where is it not allowed to copy large chunks of the article directly as this will be a violation of intellectual property. However, it is allowed to display text if it is a processed version of the article. This means that adding a paraphrasing component on top of a strong extractive summarizer can solve the legal issues if the extractive summarizer is good enough.
There is a large body of research on text paraphrasing, sentence compression is considered a paraphrasing task, another prominent paraphrasing task is style transfer.
## Conclusion
In this article, we tried to provide a very wide and very shallow overview of the implementation of abstractive summarization in under-resourced languages. This research only lists some of the approaches that we believe are promising in implementing an under-resourced abstractive summarizer. And that implementing any of the aforementioned approaches must be preceded by further "deeper" research. Following are some of the highlights we can have from our research:
- The main issue when moving to abstractive instead of extractive summarization is fear of non-truthful summary where the model would start generating random stories related to the whole topic instead of summarizing the actual article. This phenomenon is usually associated with neural models and to a smaller degree with non-neural approaches.
- We have not come across many research articles in the task of automatically or semi-automatically building training set for abstractive summarizers with the exception of [1] and [11] some of these approaches are applicable but we believe that further research in this area can be of value.
- The non-neural models while being general and usually domain and language independent are complex to implement and usually have little to no improvement on their extractive counterparts
- Although the reported performance of the unsupervised neural models is rather low, since some of them include code bases it might be easy to test them as an off the shelf component.
- In our point of view, the usage of paraphrasing component can be a great addition to our current system that uses extractive summarization and would usually require much lower time to build.
## References:
[1] P. J. Liu *et al.*, “Generating wikipedia by summarizing
long sequences,” *ArXiv Prepr. ArXiv180110198*, 2018.
[2] P.-E. Genest and G. Lapalme, “Fully abstractive approach to
guided summarization,” in *Proceedings of the 50th Annual
Meeting of the Association for Computational Linguistics (Volume 2:
Short Papers)*, 2012, pp. 354–358.
[3] L. Bing, P. Li, Y. Liao, W. Lam, W. Guo, and R. J. Passonneau,
“Abstractive multi-document summarization via phrase selection and
merging,” *ArXiv Prepr. ArXiv150601597*, 2015.
[4] T. Oya, Y. Mehdad, G. Carenini, and R. Ng, “A template-based
abstractive meeting summarization: Leveraging summary and source
text relationships,” in *Proceedings of the 8th International
Natural Language Generation Conference (INLG)*, 2014, pp. 45–53.
[5] W. Li, L. He, and H. Zhuge, “Abstractive news summarization
based on event semantic link network,” 2016.
[6] E. Chu and P. Liu, “MeanSum: a neural model for unsupervised
multi-document abstractive summarization,” in *International
Conference on Machine Learning*, 2019, pp. 1223–1232.
[7] R. Schumann, “Unsupervised abstractive sentence summarization
using length controlled variational autoencoder,” *ArXiv Prepr.
ArXiv180905233*, 2018.
[8] Y. Kikuchi, G. Neubig, R. Sasano, H. Takamura, and M. Okumura,
“Controlling output length in neural encoder-decoders,” *ArXiv
Prepr. ArXiv160909552*, 2016.
[9] M. T. Nayeem, T. A. Fuad, and Y. Chali, “Abstractive
unsupervised multi-document summarization using paraphrastic
sentence fusion,” in *Proceedings of the 27th International
Conference on Computational Linguistics*, 2018, pp. 1191–1204.
[10] K. Filippova and Y. Altun, “Overcoming the lack of parallel
data in sentence compression,” 2013.
[11] J. Zhu *et al.*, “NCLS: Neural Cross-Lingual
Summarization,” *ArXiv Prepr. ArXiv190900156*, 2019.
[12] J. Ouyang, B. Song, and K. McKeown, “A robust abstractive
system for cross-lingual summarization,” in *Proceedings of the
2019 Conference of the North American Chapter of the Association for
Computational Linguistics: Human Language Technologies, Volume 1
(Long and Short Papers)*, 2019, pp. 2025–2031.
## Frequently Asked Questions
### What are underresourced languages in NLP?
Underresourced (or low-resource) languages lack sufficient annotated training data, pre-trained models, and NLP tools compared to languages like English or Chinese. Most of the world's 7,000+ languages are underresourced, including many Arabic dialects, African languages, and indigenous languages.
### How can summarization work without large datasets?
Low-resource summarization uses transfer learning from high-resource languages, cross-lingual pre-training (mBART, mT5), data augmentation through back-translation, and zero-shot or few-shot approaches. Multilingual models trained on many languages can generalize to unseen low-resource languages.
### What is abstractive vs extractive summarization?
Extractive summarization copies important sentences verbatim from the source text. Abstractive summarization generates new sentences that paraphrase the content. Abstractive methods require more sophisticated language generation but produce more natural, human-like summaries.
### An Initial, Failed Solution for the Event Detection Task
- URL: https://mohammadshaker.com/en/blog/an-initial-failed-solution-for-the-event-detection-task
- Date: 2020-01-29T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Our first Arabic event detection system combined TF-IDF vectors, NER features, and timestamp proximity, then clustered articles sequentially against a 1,400-article ground truth spanning 120 events. It failed — and the failure analysis revealed which feature combinations actually helped and why the similarity threshold was the breaking point.
#### Content
In this post, we are trying to validate our initial solution for the event detection task.
If you're not familiar with the task you can refer to our previous post about [“How to” Event Detection in Media using NLP and AI](/blog/event-detection-in-media-using-nlp-and-ai).
For applying event detection in news articles we are planning to do the following:
- Represent each article as a vector of expressive features.
- Feed the vectorized articles into a sequential clustering model to aggregate the ones talking about the same event.
In this post, we're solving the problem of event detection in the Arabic language.
## **Features Effectiveness**
In this section, we're trying to measure the effectiveness of different features in representing the articles.
Our methodology based on the existence of a ground truth dataset of news events (i.e. news articles aggregated under the events they’re reporting). Hence, effective representation for the articles in the vector space is a one that represents the articles reporting the same event close to each other, while they are far from the articles reporting other events.
### **A Ground-Truth Events Dataset**
In a previous effort, we’ve collected a ground truth events dataset which contains ~1,400 news articles over ~120 different events. The dataset was collected following a methodology described in a previous post [Building a Test Collection for Event Detection Systems Evaluation](/blog/building-a-test-collection-for-event-detection-systems-evaluation)
### **Measuring the Quality of the Problem Representation**
We measured the effectiveness of representation in two ways.
#### **Numerical**
For each representation we
want to validate:
1. Represent the articles in the ground-truth dataset with this representation.
2. Measure the overlap between the events according to this representation.
Representations that cause a higher overlap between the events are worse.
To measure the overlap between the events we used **silhouette score** which is a measure of how similar an object (i.e. article) is to its own cluster (i.e. event) (cohesion) compared to other clusters (i.e. events) (separation). The silhouette ranges from −1 to +1, where a high value indicates that the object is well matched to its own cluster and poorly matched to neighbouring clusters. Values near 0 indicate overlapping clusters.
#### **Visual**
Moreover, for visual exploration we plot the following:
- All the articles after
projecting them into 2D space using t-SNE.
- The events, where each event
is the mean of all of its articles vectors also projected using
t-SNE.
You may notice events that are not associated with a group of points this can be explained as the articles of this event are scattered around and this representation couldn’t capture their similarities.
### **The Explored Features**
We explored features that may capture the answers of the 5W1H questions related to an event.
- **Who**: this feature can be captured by the output of a NER model, it’s specifically the persons and maybe the organizations mentioned in the article.
- **Where**: this can be also captured by the locations extracted using a NER model.
- **When**: we can use the publish date of the article to capture when the event took place.
- **What/Why/How**: are more complicated, we decided to capture them by the TF-IDF representation for the article content. Another option was to use word clusters which can refer to something similar to topics, where each word in the article is replaced by its cluster number, then the resulted content is vectorized using TF-IDF too. Check out our previous post about Arabic Words Clustering Using word2vec.
- **Other features**: multi-words terms which are simply frequent 2-grams and 3-grams. We believe that these features can also capture something related to the answer of who/where and also when.
We experimented with these features and their concatenations.
### **Results**
The results of applying the previously discussed features on our ground truth event detection dataset.
#### **Expressiveness Validation**
To validate that the chosen features are expressive as we think; we visualized the top-weighted features for each feature type in a word-cloud:
**Uni-Gram**
**Bi-Gram**
**Tri-Gram**
**Pe**ople
**Locations**
**Organizations**
#### **Silhouette Score**
The following table summarizes the results of some of our experiments
| **Feature** | **Filtering** | **Silhouette Score** |
| --- | --- | --- |
| TF-IDF (Uni) | max\_df=0.1, min\_df=10 | ~0.11 |
| TF-IDF (Uni + Bi) | **Uni**| max\_df=0.1, min\_df=10 **Bi**| max\_df=0.5, min\_df=50 | ~0.09 |
| TF-IDF (Uni + Tri) | **Uni**| max\_df=0.1, min\_df=10 **Tri**| max\_df=0.5, min\_df=50 | ~0.08 |
| TF-IDF (Uni + Bi + Tri) | **Uni**| max\_df=0.1, min\_df=10 **Bi**| max\_df=0.5, min\_df=50 **Tri**| max\_df=0.5, min\_df=50 | ~0.08 |
| TF-IDF (W2V clusters) | | ~0.08 |
| TF-IDF (Uni) + TF (PERS + LOC + ORG) | **Uni**| max\_df=0.1, min\_df=10 **PERS**|**LOC**|**ORG**| min\_df=5 | ~0.11 |
| TF-IDF (W2V clusters) + TF (PERS + LOC + ORG) | **PERS**|**LOC**|**ORG**| min\_df=5 | ~0.11 |
| TF-IDF (Uni) + TF (PERS) | **PERS**| min\_df=5 | ~0.06 |
| TF-IDF (Uni) + TF (LOC) | **LOC|** min\_df=5 | ~0.03 |
| TF-IDF (Uni) + TF (ORG) | **ORG**| min\_df=5 | ~-0.02 |
- **max\_df**: is a filtering threshold based on document frequency e.g. max\_df=0.1 means omitting the features that appeared in more than 10% of the docuemts.
- **min\_df**: is a filtering threshold based on document frequency e.g. min\_df=5 means omitting the features that appeared in less than 5 docuemts.
- max\_df & min\_df threshold values were chosen by manually inspecting the produced features by the vectorizers.
- **Why to use TF-IDF with W2V clusters?** Simply because very frequent clusters tend to be stopwords or at least not topic-related.
**Observations**: The silhouette scores -even the best ones- are not promising as they are too close to zero which means that the events are highly overlapped with this representation.
#### **Projection**
We’re not going to show all the projection plots but two as examples.
Following is the projection of the TF-IDF (Uni):
Here are two of the highly overlapped events (31, 90):
In **English**
In **Arabic**
The Lebanese prime minister Saad Al-Hariri asked for international help after the struggling to contain forest fires.
تواصل سلسلة من حرائق الغابات في لبنان، ورئيس الوزراء الحريري يطلب مساعدات دولية لاحتوائها
More than 100 wildfires have erupted in the forests of three Syrian governorates, most of them have been put under control by the Syrian Civil Defense Forces
نشوب أكثر من مائة حريق في غابات ثلاث محافظات سورية، استطاعت قوات الدفاع المدني السوري السيطرة على معظمها
Such events were supposed to be split after using the NER features according to their locations.
Following is the projection of TF-IDF (Uni) + TF (PERS + LOC + ORG):
The events are obviously more scattered and overlapped than the previous one, which explains why this representation Silhouette Score it lower.
Although the new representation using NEs could split 31 and 90 events according to their locations (one in Syria and the other in Lebanon), another problem emerged as 90 and 34 overlapped:
In **English**
In **Arabic**
The Lebanese prime minister Saad Al-Hariri asked for international help after the struggling to contain forest fires.
تواصل سلسلة من حرائق الغابات في لبنان، ورئيس الوزراء الحريري يطلب مساعدات دولية لاحتوائها
Protests in Lebanon over plans to impose new taxes
احتجاجات واعتصامات في شوارع لبنان، احتجاجاً على فرض ضرائب جديدة على الشعب اللبناني
#### Experiments With The Date Feature
We decided to use the publish date of the articles to model the time aspect of the events. Since date features have a very different meaning than the TF-IDF ones it’s probably not a good idea to combine them in one vector. However, we tried with doing so.
There are [many ways to handle time data representation as machine learning features](https://datascience.stackexchange.com/a/2370), but none of them seemed to be suitable for our problem. Here are some of our experiments:
- Date as three features (day, month, year): As these features ranges are much bigger than the range of the TF-IDF weights, the date features omitted the TF-IDF features effect, moreover the day feature has the bigger effect as it’s the most changeable one with high values e.g. articles published in day 20 from any month tend to be close to each other despite their different content or even different publish month.
- Date as a [timestamp](https://www.programiz.com/python-programming/datetime/timestamp-datetime): we tried to represent the date in one feature as a timestamp to reduce the effect of the day feature alone. However, again this feature range is much bigger than the range of the TF-IDF weights, which omitted the TF-IDF features effect.
- Date as a normalized timestamp: we normalized the timestamp to reduce its effect. However, it did not affect the results anymore.
### To Wrap up
The chosen representation seems not able to characterize the problem perfectly:
- High overlap between our
ground truth events.
- NEs don't play any role,
although they represent most of the information about the event.
- Dates can't be added to the
representation, although it's a significant characteristic for an
event.
## Training The Model
let's try to validate our assumptions towards the event detection features representation, by training and evaluating a model using them.
### **Evaluation Metrics**
The evaluation metrics for clustering given a ground truth include:
**Homogeneity score**: A clustering result satisfies homogeneity if all of its clusters contain only data points that are members of a single class. Examples:
Split
classes into more clusters can be perfectly homogeneous:
True
labels = [0,
0, 1, 1], Predicted
labels = [0,
0, 1, 2]
Score
= 1
Clusters that include samples from different classes do not make for homogeneous labelling:
True
labels = [0, 0, 1, 1], Predicted labels = [0, 1, 0, 1]
Score
= 0
**Completeness Score:** A clustering result satisfies completeness if all the data points that are members of a given class are elements of the same cluster. Examples:
Non-perfect labelling that assigns all classes members to the same clusters are still complete:
True
Labels = [0, 0, 1, 1], Predicted labels = [0, 0, 0, 0]
Score
= 1
If
classes members are split across different clusters, the assignment
cannot be complete:
True
labels = [0, 0, 1, 1], Predicted labels = [0, 1, 0, 1]
Score
= 0
**V-Measure Score:** The V-measure is the harmonic mean between homogeneity and completeness, thus perfect labelling is both homogeneous and complete. Examples:
Labellings that assign all classes members to the same clusters are complete be not homogeneous, hence penalized:
True labels = [0, 0, 1, 2], Predicted Labels = [0, 0, 1, 1]
Score = 0.8
Labellings that have pure clusters with members coming from the same classes are homogeneous but unnecessary splits harm completeness and thus penalized as well:
True
labels = [0, 0, 1, 1], Predicted Labels = [0, 0, 1, 2]
Score
= 0.8
Hence, we chose V-Measure as our evaluation metric.
### **Clustering Algorithm**
We performed Birch which is an online learning algorithm that constructs a tree data structure with the cluster centroids being read off the leaf. Birch has an important hyper-parameter that plays a critical role in its final results which it the **threshold,** in Birch the radius of the sub-cluster obtained by merging a new sample and the closest sub-cluster should be lesser than the threshold. Otherwise, a new sub-cluster is started. So, setting the threshold value to be very low promotes splitting and vice-versa.
### **Training and Evaluation**
In each experiment, the Birch model was fed by the articles gradually according to their published order, where each fed-batch represents the articles published along a day.
The following tables summarize our experiments which involve two kinds of features:
TF-IDF (W2V
clusters) + TF (PERS + LOC + ORG)
| **Birch Threshold** | **V-measure score** | **n\_clusters** |
| --- | --- | --- |
| 2 | 0.6591082932535002 | 900 |
| 2.5 | 0.6388812411242022 | 706 |
| 3 | 0.6157702294699662 | 523 |
| 3.5 | 0.5842678797976664 | 362 |
| 4 | 0.5387831995084763 | 255 |
| 4.5 | 0.4744330956061701 | 162 |
TF-IDF (Uni)
+ TF (PERS + LOC + ORG)
| **Birch Threshold** | **V-measure score** | **n\_clusters** |
| --- | --- | --- |
| 2 | 0.6643862201200933 | 916 |
| 2.5 | 0.6388584348941666 | 723 |
| 3 | 0.6170028933666231 | 548 |
| 3.5 | 0.5881280950330579 | 378 |
| 4 | 0.5472522486363269 | 262 |
| 4.5 | 0.49996199074765435 | 178 |
**Observations:** Given that we have ~120 events in our ground truth, a number of clusters like 900 is huge enough to approximately put each article in a separate cluster which is definitely not a good clustering. However, V-measure gave it the highest score, thus we couldn't be confident in this evaluation score and we checked the resulted clusters manually. The best clusters were produced by TF-IDF (W2V clusters) + TF (PERS + LOC + ORG) with a threshold=3. However, these clusters are still messy. This can be revert to and prove the fact that the used features can't characterize this problem perfectly.
**Examples of the produced clusters**:
## **Conclusion**
Our initial solution for event detection does not seem to be the perfect way to solve this problem. Now we need a new road map to continue on, which will be proposed in our future posts.
### Building a Test Collection for Event Detection Systems Evaluation
- URL: https://mohammadshaker.com/en/blog/building-a-test-collection-for-event-detection-systems-evaluation
- Date: 2020-01-29T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Evaluating an event detection system requires a labeled test collection — but building one for Arabic news means resolving annotation disagreements, defining event boundaries, and selecting a corpus that reflects real news diversity. We detail our methodology for constructing a 1,400-article benchmark across 120 events.
#### Content
Before we start, if you're not familiar with the Event Detection task in NLP you can refer to our previous post on this topic [here](/blog/event-detection-in-media-using-nlp-and-ai).
So you've built a system to detect events in the media… now what?
While building a system is a key step, how the system performs on real-world data has equal importance. We need to know whether it actually works and if we can trust its decisions.
So.. we need to evaluate our system before putting it in use.
Evaluation is a highly important step in the development of any system type as it allows the judgment on the system quality and assures that it meets its targeted goals efficiently providing the customers with a fulfilling reliable service.
In this post, we're discussing a technique to build a test corpus for evaluating an event detection system.
## What Does It Mean to Be a Successful Event Detection System?
Event detection systems usually deal with a stream of data (e.g. news articles when working in the traditional media domain).
The performance of an event detection system is measured in terms of its efficiency in performing its main functionalities:
- Locating similar news stories in the data stream.
- Detecting the new news stories in the data stream.
## How to Judge an Event Detection System Performance?
Judging an event detection system can be done using a test collection.
A test collection typically consists of:
1. A set of documents (e.g. news stories).
2. Events representing information needs.
3. Relevance judgments (i.e. labels) that specify which documents are relevant to the topics.
## How to Build a Test Collection for Event Detection Systems?
When constructing a test collection there are typically a number of practical issues that must be addressed.
### **What Items Should be Selected to Create The Document Collection?**
To construct a test collection, a dataset has to be acquired first. This document collection should be a good representative of the domain in which the systems will be applied.
### **How Should a Suitable Set of Events be Generated?**
Previous works elect to choose significant (i.e., popular) events to be the core of the test collection.
This has several advantages:
- Rich content (i.e., relevant stories) is more available for popular events.
- The popularity of the topics might help annotators to do more consistent judgments due to their probable familiarity with the significant events.
*How to determine significant events in a predefined period of time?*
- A set of events that took place in that period is determined with the help of [Wikipedia’s Current Events Portal](https://ar.wikipedia.org/wiki/%D8%A8%D9%88%D8%A7%D8%A8%D8%A9:%D8%A3%D8%AD%D8%AF%D8%A7%D8%AB_%D8%AC%D8%A7%D8%B1%D9%8A%D8%A9).
- The event set is filtered through manually searching for articles on popular news websites such as Aljazeera and CNN… using several manually-crafted queries per event. Keep those events that have been discussed by at least # news article.
### **How Should The Events be Expressed?**
Each event requires getting a set of potentially-relevant documents that constitute its **Judgment Pool**.
Previous studies showed that query variations for a given event are strong in producing a diverse document pool. Therefore, in order to diversify the documents pool, researchers adopted a process that manually crafts a list of keyword and phrase queries for each event by performing an interactive search on news websites. This would ensure wide coverage of the event aspects.
The crafted queries can then be used to search a documents-collection using an off-the-shelf retrieval engine to retrieve the potentially-relevant documents that constitute the judgment pool.
### **How Many Events are Required for Obtaining Reliable Evaluation Results?**
**A minimum of 50 events** should be included in the test collection to ensure reliable. That is the number typically used in [TREC (Text Retrieval Conference)](https://en.wikipedia.org/wiki/Text_Retrieval_Conference) and stated to be reliable in practice in literature [1].
### **Do The Events Represent a Diverse Enough Set of Information Needs?**
A range of events should be selected with varying characteristics to test the systems under a range of settings.
### Document-Event Relevance Assessments
After identifying the potentially-relevant documents, we need to collect relevance judgments.
Relevance judgment considers answering the question of: **is this document talking about this event?**
After this judgment, we'll get a set of clusters each is relevant to a specific event. Each event cluster contains stories cover different aspects of that event.
#### Novelty Judgments
Relevant documents for each event are distributed into clusters of “semantically-similar” documents; each cluster has a set of relevant documents considered to carry the same information. The determination of those internal clusters is considered a second layer of labels on top of the relevance judgments. The instructions for this labeling-level task have been shared by [1] [here](https://reemsuwaileh.github.io/EveTARNovelty/training.html).
## Event Detection in the Arabic Language
To the best of our knowledge, there is no prior work that constructs an Arabic-only test collection as a primary contribution. Therefore, researchers working on Arabic event detection systems had to construct their own test collections to conduct their experiments.
The only work that we are aware of is [1]. However, they considered building a test collection which consists of tweets rather than news articles.
## Usability
Event detection task requires relevance judgments for each event. This inherently enables the test collection to support the Ad-hoc Search (AS) task as well. Ad-hoc search is the typical search task in IR in which an ad-hoc
query (representing a topic of interest) is issued at a search system which is required to retrieve a ranked list of documents that are relevant to that topic over a collection of documents.
## Conclusion
A typical test collection comprises three major components: a document collection, a set of topics, and a set of judgments per topic. There are many challenges in the test collection construction process. First, the document
collection should be a good representative of the domain in which the systems will be applied. Second, the events should be carefully designed to represent real-world information needs. Third, the document-event pairs
to be judged should be carefully selected and the collected judgments should be consistent to achieve reliable evaluation.
## References
[1] Hasanain, Maram, et al. "EveTAR: building a large-scale multi-task test collection over Arabic tweets." *Information Retrieval Journal* 21.4 (2018): 307-336.
### Initial Genre Classification Experiments
- URL: https://mohammadshaker.com/en/blog/initial-genre-classification-experiments
- Date: 2020-01-29T00:00:00.000Z
- Tags: arabic, genre-classification, ml, nlp, research, Human-Written
Automatically classifying Arabic news articles by genre — politics, sports, business, science — lets a news aggregator route and rank content intelligently. We describe initial experiments using NLP features extracted from a corpus of Arabic news articles across major outlets, evaluating multiple classification models and reporting where genre confusion is highest.
#### Content
The ability to filter your news feed based on the genre is a critical component of any news aggregator, users would usually want to read sports or political news only not just the most recent or hottest news. In this post, we will explore in great details our initial genre classification system.
## Let's start with the.. ***data***
In the following experiments, we used an in-house data set.
The data set is composed of 190307 HTML document crawled from the following domains [Aljazira, Alarabia, Aljadeed, RT Arabic, BBC arabic].
For each of the documents we tried to extract the following features:
- title: title of the article
- text: the actual text of the article
- keyphrases: using either meta info or keywords set by the author
- summary: some anchors use the inverted pyramid style and their lead paragraph can be used a basic extractive summary
- genre (political, sports, … ) extracted from the URL, the meta-tags or in sometimes the HTML body
- URL
- domain
### Data cleaning
After the information is extracted only articles that contain text (regardless of the size) is preserved (many of the pages represents tags, infographics, ...) leaving us with 113526 articles.
Next, we removed any articles that have text but to which we could not extract the genre automatically this reduced the number of documents to 104205. The distribution of the articles among the news anchors is shown in the following graph:

However upon manual inspection, we found that several articles specifically from Aljadeed had very little text in them, the following figures show the distribution of articles’ lengths by words without any normalization among some of the news anchors. Note that in rt and BBC articles nearly 20% of the articles have length below 50 words and 10% of Aljazeera and Aljadeed articles follow this phenomenon. These articles were mostly breaking news, one-liners or reports and info-graphics. We decided to discard any article with a length lower than 50 words, the resulting set accounted for 95140 articles.




Next, we started filtering based on genre, many of the articles had uninformative genres like years in case of Alarabia and numerical values in case of BBC. The new filtered set contained 94953 articles. The distribution of the genres among anchors is shown below.



Note that all the anchors have extremely splintered distributions with the exception of Aljadeed that have way too little genres. note that in the BBC there are several articles with the same genre but with different formatting example “science-and-tech” vs “scienceandtech” these will be resolved once we unify the genres across the whole dataset.
Aljazeera graph was omitted because it was extremely splintered mainly because the genres in Aljazeera have hierarchical structure example:
- اقتصاد > قضايا
- اقتصاد > مؤسسات وهياكل
- الأخبار > استطلاع رأي
- الأخبار > الاقتصاد
- الأخبار > تقارير وحوارات
### Data Mapping
It is important to find a unified mapping among these different genres that correspond to the same content, we developed a manual mapping that unifies all of these different genres based on a single labels set the categories of this mapping is shown below:
- uncategorizable: these are genres that have no clear topic they include editor choices, videos and reports and similar stuff, or have an extremely low number of article to have its own category like hajj
- News: these are news articles that cover both the Arabic world and other countries they might include both global catastrophes like earthquakes, alongside political events
- world news
- art\_and\_culture: this news focus on artistic news such as galleries, history stories or musical pieces
- economy
- health
- IT: specifically related to IT
- science: a broader view covering both tech and other sciences like genetics, math, …
- sports
It is clear that such a list of genres is very limited and much more detailed feeds should be considered. However, the choice of these genres is motivated by 2 reasons:
- the goal of this experiment is to build an initial genre classification system, having a lower number of genres means faster development and relatively better performance
- as we will show next we have found out that our dataset is extremely imbalanced even with such a small number of classes, increasing the number of genres might create a lot of small classes on which machine learning model would struggle to train and generalize
The following figure illustrates the distribution of the whole data based on this new mapping, we can easily see the imbalance between the classes with news covering nearly half of the data.

Furthermore, on initial glence we can see that the classes are not 100% pure following are some word clouds for the classes these are simply the most frequent words apart from stop words



### Data Preparation
To simplify evaluation we created a holdout test set by randomly sampling 100 articles from each class, the following graph illustrates the distribution of the classes in the new training set.
As we already mentioned the data is extremely imbalanced, and thus training a full multi-class classifier might cause the model to favour the most frequent classes, to circumvent this for each of the classes we created a balanced one vs all training set where for a given genre X the positive samples are those belonging to the genre and the negative sample are randomly selected from the other genres with each genre contributing nearly the same amount of articles. This resulted in a nearly balanced negative set for each of the genres.
## ML Experiments
Regarding our models choice we settled with Facebook's fasttext is pretty fast and pretty simple that is why we started with it, if you are not familiar with the algorithm check [this](https://research.fb.com/downloads/fasttext/), following are the experiments we did, we will report only f1 to be brief, all the results are reported on a per-class 10% held out development set
| Experiment No | Details |
| --- | --- |
| 1 | training directly on each of the genres datasets separately |
| 2 | in the previous experiment, we noticed that the classes with a lower number of articles had the worst results, this might be linked to the fact that the model is not being able to get a good words representation. To rectify this we pre-trained a fasttext model on all the training texts in an unsupervised manner to get better words representations and then re-adapted the model for each of the classes data |
Next, we tried a fast hyperparameters search on the development set following are the details of each experiment, for simplicity we didn’t do a full grid search (No time) and thus we select the best value for each of the hyperparameters. In all of the following experiments, we start from experiment No2
- Initial learning rate: the default value in exp2 is 0.05 we tested the values 0.1 and 0.25 in experiments 3 and 4 the latter gives a small increase in overall f1 of nearly 0.5% absolute
- Number of epochs: again we started from 2 the default value is 5 we tried 10, 25, 50 and 100 these are experiments 5,6,7 and 8 we found that the value of 25 gave a tangible increase in overall f1 of 1% absolute while greater values had no real effect.
- Final layer: the default in 2 is the Softmax we tried 2 additional types Negative Sampling exp 9 and Hierarchical Softmax exp 10 we found that Negative sampling loss had a small increase of 0.5% absolute
- Best parameters model: we selected the best values for every parameter that is lr=0.25, epochs=25 and loss=Negative Sampling the model gives a sizable increase of 1.5% absolute over the model in 2
The following table shows the overall results on the development set, while they seem high we are afraid of over fitting especially in sports class since it’s results were extremely high.
| **name** | **tech** | **politics** | **health** | **sports** | **economy** | **science** | **world\_news** | **art\_and\_culture** | **mean** |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | 0.706989 | 0.90438 | 0.75697 | 0.98760 | 0.93455 | 0.741667 | 0.901678 | 0.59687 | 0.81634 |
| 2 | 0.91129 | 0.88100 | 0.95219 | 0.98829 | 0.94764 | 0.8875 | 0.887242 | 0.925 | 0.92252 |
| 3 | 0.905914 | 0.88340 | 0.96015 | 0.99035 | 0.95113 | 0.885417 | 0.887242 | 0.925 | 0.92357 |
| 4 | 0.903226 | 0.89622 | 0.95617 | 0.98966 | 0.95200 | 0.906250 | 0.894265 | 0.93125 | 0.92863 |
| 5 | 0.911290 | 0.89542 | 0.96015 | 0.99104 | 0.95200 | 0.893750 | 0.893874 | 0.93125 | 0.92860 |
| 6 | 0.911290 | 0.90823 | 0.95219 | 0.99173 | 0.95462 | 0.910417 | 0.903629 | 0.92500 | 0.93214 |
| 7 | 0.913978 | 0.90422 | 0.96414 | 0.99242 | 0.94764 | 0.916667 | 0.905970 | 0.93125 | 0.93453 |
| 8 | 0.913978 | 0.89958 | 0.96414 | 0.99311 | 0.93630 | 0.918750 | 0.907920 | 0.93125 | 0.93313 |
| 9 | 0.913978 | 0.89493 | 0.96015 | 0.98966 | 0.95026 | 0.895833 | 0.895045 | 0.92812 | 0.92850 |
| 10 | 0.911290 | 0.88196 | 0.96015 | 0.98898 | 0.94677 | 0.875000 | 0.884120 | 0.93125 | 0.92244 |
| 11 | 0.930108 | 0.90502 | 0.96414 | 0.99311 | 0.94589 | 0.918750 | 0.914163 | 0.93750 | 0.93858 |
## Testing on Live Data
In order to validate our results we tested the model on a fresh version of our news feed.
### The **Data**
We are using a snapshot of Almeta dataset to do this analysis, the original set contains 12935 samples.
Furthermore, the text fields in this snapshot do not include any English words as they are extracted from the results of the ESL model. However, our models were trained on articles with English words retained this might affect the results of the model on this set. However, Note that this deficiency will not be present during the deployment.
Furthermore, we removed all the articles that have less than 50 words in their text
The distribution of the articles among the set is shown below

### The **Model**
The model used in this analysis is trained using the best configuration we obtained in exp No11 with one difference, we removed the uncategorizable class from the negative samples of the training set since it was extremely impure. The model is tested in 2 settings single and multi-tag:
- Single\_tag: here we simply retrieve the tag with the highest confidence if it has confidence higher than 0.5 or uncategorized. In this case, the model is correct if the tag is suitable for the article.
- Multi\_tag: here for a given threshold T we report all the tags with confidence higher, here the model is correct if all the tags reported are suitable for the article.
### **Single** T**ag Analysis**
#### **Confidence Distributions**
First, let us see the distribution of the genres among the dataset and see if it is rational, the distribution seems ok, although we expected the news to be larger while the uncategorizable class have very small size.

Next let us see the confidence distribution across the whole data, from the distribution we see that nearly all the tags have very high confidence this might seem problematic, However, remember that for each article we are selecting the model with the highest confidence. To assess this let us see the confidence distribution of each of the individual models.

Following are the distributions of each individual model, we can see that most of these model have near binary confidence which might be problematic meaning that each individual classifier is very strict, but this does not imply that the bootstrapping of these classifiers is bad, actually it is the exact opposite.



#### **Confidence Analysis**
We split the data into 3 ranges only based on model confidence, these ranges are based on the extremely binary distribution of confidence alternatively:
- strong confidence [100-90]
- normal confidence [90-50]
- uncategorizable [<50]
Note that these ranges apply only to the most probable tag
**Strong Confidence**
This constitutes most of the dataset 11813 articles, as we have seen the distribution is very binary, we analyzed only 2% of them basically 236 articles, furthermore, to avoid the bias towards the news class due to it’s size the selected sample was uniformly distributed between the genres to facilitate per genre error evaluation. The following table shows the analysis of the errors on this set. Where:
- #false\_positives mean actual false positives where the model totally failed to select the correct tag and where it is possible to give the model a tag from the classes we have considered,
- #incidents represents one class that we have seen a lot in the false positives namely the incidents like killing, theft, or accidents. Example
- شاهد.. إخلاء جوي للمتواجدين في محطة قطار الحرمين خلال الحريق
- شرطة سلطنة عمان تصدر بيانا بعد القبض على أفراد عصابة من جنسية عربية
- فيديو.. حريق في منشأة نفطية إيرانية
- # unknown represents articles that won’t fit any of the classes used examples
- صراع تحت الماء من أجل البقاء
- #invalid articles with invalid text from the collection process
Now let us analyze the patterns of the false positives:
- Incident class or local news, this is a relatively common class it’s a percentage in this small sample is around 3% but it was not accounted for.
- Weapons industry: a lot of articles specifically from RT talks about weapon technology and industry, this class was not considered during training and in test time it represents the bulk of the science and tech classes examples:
- الولايات المتحدة تعلن عن اختبار ناجح لصاروخ مضاد للسفن قرب جزيرة غوام
- انطلاق مبيعات النماذج المدنية لـ"كلاشينكوف آكا -12" بالتجزئة
- شاهد.. إنزال فوج من مظليي البحرية الروسية
| Class | #true | #false positive | #incidents | #unknown | #invalid | accuracy |
| --- | --- | --- | --- | --- | --- | --- |
| art\_and\_culture | 29 | 3 | 1 | 0 | 1 | 0.872941 |
| economy | 30 | 1 | 3 | 0 | 0 | 0.882353 |
| health | 26 | 3 | 3 | 1 | 1 | 0.785294 |
| news | 31 | 2 | 0 | 0 | 0 | 0.935765 |
| science | 32 | 2 | 0 | 0 | 0 | 0.941176 |
| sports | 34 | 0 | 0 | 0 | 0 | 1.000000 |
| tech | 28 | 6 | 0 | 0 | 0 | 0.823529 |
| Sum | 209 | 19 | 7 | 1 | 2 | 0.891579 |
**Normal Confidence:**
This sample amount to 1115 article from the set, we select 10% of it amounting to 111 articles uni-formally among the genres. And manually check the results.
| Class | #true | #false positive | #incidents | #unknown | #invalid | accuracy |
| --- | --- | --- | --- | --- | --- | --- |
| art\_and\_culture | 9 | 1 | 0 | 0 | 2 | 0.900000 |
| economy | 8 | 2 | 0 | 0 | 2 | 0.800000 |
| health | 11 | 1 | 0 | 0 | 0 | 0.916667 |
| news | 12 | 0 | 0 | 0 | 0 | 1.000000 |
| science | 11 | 0 | 0 | 0 | 1 | 1.000000 |
| sports | 12 | 0 | 0 | 0 | 0 | 1.000000 |
| tech | 6 | 5 | 0 | 0 | 1 | 0.545455 |
| art\_and\_culture | 9 | 1 | 0 | 0 | 2 | 0.900000 |
The main problem in this range is the prevalence of invalid articles following is an example (sorry it is normalized so a bit hard to read), Note How the article is basically word salad
الوسائط مكتبه التقارير P D P D D P D P P D P D P P P P P P D D D D D D D D D P P P D P D D P D P P D P D يجمع خبراء المال والاعمال علي ان هناك عده اسباب تقف وراء فشل الكثير من المشاريع التي تكون واعده في بداياتها P ومن الاسباب التقليد الاعمي للمشاريع الاخري P او عدم الاستعانه بالخبرات المؤهله P او عدم القدره علي ضبط المصاريف P تقرير P محمد فاوري قراءه P محمد رمال تاريخ البث P D بحسب وسائل اعلام تركيه فان المنطقه الامنه تمتد من نهر الفرات غربا الي مدينه المالكيه بمثلث الحدود التركيه العراقيه بعمق يتراوح بين D كيلومترا و D كيلومترا P ويسيطر علي معظمها وحدات حمايه الشعب الكرديه P تاريخ البث P D اطلقت تركيا عمليه عسكريه بشمال شرق سوريا اسمتها P نبع السلام P وقالت ان هدفها محاربه التنظيمات الارهابيه P الا ان ردود الفعل الدوليه تباينت بشانها P فقد ندد بها الاوروبيون ورفضت اميركا المشاركه فيها P في حين اظهرت روسيا وايران تفهمها للعمليه P تقرير P ناصر ايت طاهر تاريخ البث P D اكدت رئاسه تركيا علي لسان رئيس دائره الاتصال فيها ان الجيش التركي جاهز لعبور الحدود السوريه وتنفيذ عمليه عسكريه P في حين توقع مسؤولون اميركيون ان يبدا الاتراك العمليه التي يشارك فيها الجيش السوري الحر خلال D ساعه P تقرير P فاضل ابو الحسن تاريخ البث P D تمور فلسطينحسب تاكيد البنك المركزي الاردني ارتفاع حجم مديونيه الافراد الي D مليار دولار P وهناك عوامل عده قللت من القدره الشرائيه للاردنيين واجبرتهم علي الاقتراض P في مقدمتها ارتفاع مطرد لاسعار السلع والخدمات وزياده معدل الضرائب P تقرير P رائد عواد تاريخ البث P D
The second important problem is the prevalence of the weapons articles in this range which particularly hurts the tech class
**Uncategorizable:**
This set has very low confidence across all the models and is comprised only of uncategorizable articles the size of this set is only 205 articles and we manually annotated 25 articles out of it only 11 were actually uncategorizable, the best option for this range is to simply be ignored since it holds very little articles.
#### Text Length Impact
To measure the impact of text length on the classifier results let us see the distribution of the errors based on the length of the article in words, the following graph (left) shows the histogram of this over the three manually annotated sets. We can clearly see that the number of errors especially concentrates in the area of short articles and that the errors nearly diminishes as the articles grow longer, to validate this, even more, let us inspect the distribution of the High confidence range alone since it represents the bulk of the results, the following graph (right) illustrates the results, we can see even clearer correlation between the article length and the errors rate (basically most of the errors happens in articles shorter than 150 words), most importantly for articles shorter than 50 words the accuracy falls below 50%


### Multi Tag Analysis
Again, we split the confidence range into the following sub-ranges:
- Strong confidence [100-90]
- Normal confidence [90-50]
Here we test based on 2 settings:
- Strict: all the tags must be reported correctly for the example to be assigned as correct
- Mild: the majority of the tags reported should be correct for the example to be assigned as correct
**Strong Confidence:**
The range covers 1122 articles that can be annotated by at least 2 tags. The following chart illustrates the distribution of the number of tags, most of the articles can be assigned 2 tags, some 3 and hardly any 4.

We sample 10% of these articles this results in 111 articles and manually checked them in the 2 aforementioned settings.
- Strict fashion: 86 articles had correct mapping, i.e. an accuracy of 77% out of these errors 28% were caused by confusion in the health class, 24% by tech, 12% by culture
- Mild fashion: 108 articles i.e. accuracy of 97% this is mainly motivated by the fact that our classes are really wide.
**Error Analysis:**
The main patterns found here are extremely similar to the single tag case specifically weapons and military training in case of tech and incidents in case of health
N**ormal confidence:**
This part of the results contains only 140 articles (there are 140 articles that have 2 or more classes with confidence between 0.7 and 0.9), The following chart illustrates the distribution of number of tags, most of the articles can be assigned 2 tags, some 3 and hardly any 4. we manually annotate 50% of this set to find any weird error patterns.

## Conclusion:
In this article we detailed the process we took to develop our initial genre classification mode, here are the notes we found:
- The plain accuracy of the models in this case is not awful but does need to be improved.
- The pre-normalized text in the test set have reduced the accuracy and we expect better results on deployment
- Short articles represents a problem for this model
- Most of the models have a nearly binary confidence, this is good since a bootstrapped or a voting model will have less confusion, but on the other hand this means that much less of the results will have 2 or more tags which seems not accurate as many of the articles we have reviewed can have more than a single tag
- The main cause for errors is the under-representation of some classes most notably incidents, weapons, tourism,… the best way to handle this issue is to create a fairly detailed taxonomy based on our choices of important classes rather than the tags given by the authors and scrapping specialized sites for it, this can enable us to do something like Hashtags and hashtags following in a manner similar to Tumblr.
- Multi-tag classification can and should be implemented since the accuracy of it is rather OK
- The binary nature of the confidence nearly wipes out the other class this might be caused by the selected model (fasttext) we can resolve this in 2 ways:
- Add much more models to cover as many other classes
- Use different models that have less edgy confidence
### News Stream Clustering - Sequential Clustering in Action
- URL: https://mohammadshaker.com/en/blog/news-stream-clustering-sequential-clustering-in-action
- Date: 2020-01-29T00:00:00.000Z
- Tags: ml, nlp, research, Human-Written
“Sequential clustering for news streams groups incoming articles into event clusters in real time, without a fixed cluster count. Each document is compared to existing cluster centroids and assigned to the best match above a similarity threshold — or starts a new cluster. We show this algorithm applied to an Arabic news stream with real results.”
#### Content
In a previous post, we talked about [“How to” Event Detection in Media using NLP and AI](/blog/event-detection-in-media-using-nlp-and-ai). In another post, we presented the [Sequential Clustering](/blog/an-implementation-of-a-news-stream-sequence-clustering-algorithm). Today we're introducing an online (sequential) clustering algorithm specialized in aggregating news articles into fine-grained story clusters.
## Problem Formulation
We focus on the clustering of a stream of documents, where the number of clusters is not fixed and learned automatically. We denote by ***D*** (potentially infinite) space of documents. We are interested in associating each document with a cluster via the function ***C***(***d***)∈ ***N***, which returns the cluster label given a document.
For each cluster, we maintained a centroid that needs to be incrementally updated to reflect new information that exists in a new incoming document.
## The Clustering Algorithm
The online clustering process works as follows:
With a new incoming document ***d***, represented as a vector, we compute a similarity metric between the document vector and each of the existing centroids. If the largest similarity exceeds a threshold **τ** for cluster index ***j***, then we set ***C***(***d***) = ***j***. In that case, we also update the ***j***th cluster centroid to include new information from document ***d***. If none of the similarity values exceeds a threshold **τ**, we find the first cluster's *id* which is still unassigned ***i*** and set ***C***(**d**) = ***i***, therefore creating a new cluster.
We maintain a centroid for each cluster that is created as the average of all the representations of documents that belong to that cluster. We recalculate this average for each document insertion in the cluster.
## Document Representation
In the proposed clustering algorithm each news document can be represented by multiple subvectors, where each subvector denotes a feature type e.g. several TF-IDF subvectors with words, word lemmas and named entities. Besides these textual vectors, we use the documents' timestamp.
### Features Overview
In our previous post: [“How to” Event Detection in Media using NLP and AI](/blog/event-detection-in-media-using-nlp-and-ai), we discussed a bunch of representative features for events in news documents. We're going to briefly give an overview of these features in the rest of this section.
The textual features related to an event usually try to capture the answers of the 5W1H questions in the news article to gather information about a story. Questions like *who* and where can be answered by extracting the named entities in the article. While *what* and *how* are more complicated, some representations like the TF-IDF of the whole content or the LDA representation of the article can be used to capture them. Finally, the when question is answered using the timestamp feature.
## Similarity Metrics
The used similarity metric computes weighted cosine similarity on the different subvectors. Formally, the similarity is given by a function defined as:
Here, ***dj*** is the ***j***th document in the stream and ***cl*** is a cluster.
### Textual Similarity
The textual similarity is computed on the features subvectors where ***K*** is the number of subvectors for the relevant document representation.
The function ***φi***(***dj***, ***cl***) returns the cosine similarity between the document representation of the ***j***th document and the centroid for cluster ***cl***. The vector ***q0*** denotes the weights through which each of the cosine similarity values for each subvector is weighted.
### Time Similarity
The function ***γ***(***d***, ***c***) that maps a pair of a document and a cluster is defined as follows:
for a given ***μ*** and ***σ*** > 0. For each document ***d*** and cluster ***c***, we generate the following three-dimensional vector ***γ***(***d***, ***c***) = (***s1*** , ***s2***, ***s3***):
- ***s1*** = ***f***(***t***(***d***) − ***n1***(***c***)) where ***t***(***dj***) is the timestamp for document ***d*** and ***n1***(***c)*** is the timestamp for the newest document in cluster ***c***.
- ***s2*** = ***f***(***t***(***d***)***−******n2***(***c***)) where ***n2***(***c***) is the average timestamp for all documents in cluster ***c***.
- ***s3*** = ***f***(***t***(***d***) − ***n3***(***c***)) where ***n3***(***c***) is the timestamp for the oldest document in cluster ***c***.
The vector ***q1*** denotes the weights for the timestamp features
These three timestamps features model the time aspect of the online stream of news data and help disambiguate clustering decisions since time is a valuable indicator that a news story has changed, even if a cluster representation has a reasonable match in the textual features with the incoming document.
Regarding the hyper-parameters related to the timestamp features, we can fix ***μ*** = 0 and tune ***σ*** on the development set.
## Learning To Weight The Similarities
Perhaps the similarity value related to a specific feature type matters more than others in taking the decision towards choosing the best cluster for a document. ***q0*** and ***q1*** encode the contributions of each of these feature types to this decision. In order to determine the weights in ***q0*** and ***q1*** we can consider training a ranking model to learn them, where the ranking problem can be formulated as follows:
The problem is about choosing the most relevant cluster to an incoming document. This can be treated as a ranking problem where the most relevant cluster has the highest rank, thus the considered features in this problem would be our similarity metrics, and a ranking model can be trained to learn how to weight these features.
For this goal, we need to generate a collection of ranking examples, where for each incoming document in the dataset we have a ranked list of the best clusters that match it.
**How to generate the training data for the clusters ranking problem?**
In a previous post, we discussed [Building a Test Collection for Event Detection Systems Evaluation](/blog/building-a-test-collection-for-event-detection-systems-evaluation). We can generate the training data using a partition of our test collection (that we should not get evaluated on). For each document in our training partition, we create a list of clusters ranked according to their similarity to the document, this ranking can be done using only a simple vector representation like the TF-IDF of the content. Hence, we should simulate the execution of the clustering algorithm on the training partition to produce the clusters.
After all, a ranking algorithm is trained on the generated data using the document-cluster similarities as features, thus the ranking algorithm will weigh the similarities according to their contribution to the decision of choosing the best matching cluster, the resulted weights are ***q0*** and ***q1***.
## New Event Detection
New event detection refers to the process of deciding when an incoming document should be merged to the best cluster or create a new one. This decision can be made by two different common ways:
- Defining a parameter ***τ*** which is a similarity threshold. If the largest similarity exceeds this threshold***τ***the new document should be merged to the best cluster, otherwise, a new cluster should be created. This parameter can be tuned on the development set using a grid search.
- Training a binary classifier to learn this decision: simply by passing the max of the similarities between the incoming document and the current clustering pool as the input feature vector to the model. This way, the classifier learns when the current clusters, as a whole, are of different news stories than the incoming document.
## Conclusion
In this post, we described a state of the art method for clustering of an incoming stream of documents. The method works by maintaining centroids for the clusters, where a cluster groups a set of documents. We also discussed how to leverage different training procedures for ranking and classification to improve clustering decisions.
## Further Reading
[1]. Miranda, Sebastião, et al. "Multilingual Clustering of Streaming News." arXiv preprint arXiv:1809.00540 (2018).
[2]. Staykovski, Todor, et al. "Dense vs. Sparse Representations for News Stream Clustering." Text2Story@ ECIR. 2019.
[3]. Aggarwal, Charu C., and Philip S. Yu. "A framework for clustering massive text and categorical data streams." Proceedings of the 2006 SIAM International Conference on Data Mining. Society for Industrial and Applied Mathematics, 2006.
### Towards Contrary-View Detection in News
- URL: https://mohammadshaker.com/en/blog/towards-contrary-view-detection-in-news
- Date: 2020-01-27T00:00:00.000Z
- Tags: arabic, ml, nlp, research, Human-Written
Contrary-view detection finds news articles that cover the same topic but from opposing viewpoints. The problem splits into two steps: topic identification and viewpoint divergence ranking. We formalize the task, review approaches from stance detection and topic modeling, and present a document-similarity-based pipeline for surfacing contrasting perspectives in Arabic news.
#### Content
***Difference* is a fine and beautiful phenomenon. Difference should always be accepted, expected and respected.**
Difference adds richness to the topics we discuss and opens to everyone new perspectives that they never thought of.
Starting from our belief that a different viewpoint from ours is the other side of the truth that we could not reach, but was reached by others. And as a part of proceeding in our mission in Almeta to provide our users with the whole picture. In this post, we are analyzing the task of finding articles with different viewpoints than a one under consideration.
## **Problem Formulation**
Suppose we have a corpus of articles discussing different topics. Given an article with a specific viewpoint about a topic, we want to determine the articles discussing the same topic but from the contrary viewpoint.
We can split this problem into two phases, given an article *A*, discussing a topic *T*, from a viewpoint *V*:
1. Finding other articles discussing *T*.
2. Ranking these articles according to their disagreement with *V*.
## **Phase 1: Topic Identification**
In general, topics vary in their level such that higher-level topics are broader. For instance:
topic level**1**: "Sport" **>** topic level**2**: "Football" **>** topic level**3**: "The world cup" **>** topic level**4**: "France won the world cup 2018".
Let's consider the question *Which level of topics do we need to capture for this problem?*
Obviously, it doesn't make sense trying to capture contrary viewpoints in broad topics, imagine trying to measure the disagreement between two articles, one reporting a story related to basketball, while the other reporting a story related to football!
Hence, we should try to identify as finer topics as possible.
As a result, nor broad-genres text classification methods or topic modeling methods are applicable here. Simply because the intended topic-level is finer than the output of these systems.
From the example stated above, we can see that the finest topic-level is just an event reported in an article, which leads us to think about using event detection methods.
Event detection aimed to aggregate news articles into fine-grained story clusters. This task usually deals with news streams in an online manner, using sequential clustering algorithms. We usually rely on features related to named entities and timestamps to detect events. For more information about this task, you can refer to our previous post [Event Detection in Media using NLP and AI](/blog/event-detection-in-media-using-nlp-and-ai).
## Phase 2: Disagreement Based Ranking
Now we have all the articles for one event grouped together, for each article *A* in event *E* we want to rank the other articles belonging to *E* according to their disagreement with *A*. Following are our suggestions to perform this task.
### Political Orientation
When we are considering political news articles, one idea that may jump to our mind for detecting the disagreement is identifying the political orientation of the article. Then simply, articles from different orientations usually disagreed.
The emerged questions *What are these orientations? and how to detect them?*
Well, orientations vary according to the region. For instance, in western politics, the political spectrum is uniform with a clear dichotomy between the left and right. However, in Arabic politics, the vision is blurred with various orientations, that may be organizational or religious related.
Once these orientations are determined, the problem can be treated as a normal text classification problem.
For more details about the political orientation identification task, you can refer to our previous post, [Political Orientation Detection – AI and NLP Approach](/blog/political-orientation-detection-ai-and-nlp-approach).
**Limitations**:
- Genre-specific: Solves the problem only for politics articles.
- Region-specific: Orientations vary between regions.
- Not ranking: The problem can't be treated as ranking anymore.
- Vague base: Identifying political orientation tends to be vague. The orientations are indistinct, and the classification features are unclear.
- Weak assumption: Different orientations doesn't always mean disagreement in viewpoints. Perhaps two different orientations disagree on multiple concepts while agreeing on others.
### Entity-Level Sentiment Analysis
Recalling that an event is characterized by a set of entities, a viewpoint about an event can be interpreted as sentiment polarities against its entities. The entity-level sentiment analysis is a separate NLP task by itself. You can find more details about it and its solutions in our post: [Aspect-level Vs Entity-level Sentiment Analysis](/blog/aspect-level-vs-entity-level-sentiment-analysis).
In this suggested method, we rank the articles according to the agreement of their views about the named entities.
The suggested steps to rank an article is as follows:
1. Extract the named entities from the articles, using named entity extraction system.
2. Find the polarity that corresponds to each entity, using an entity-level sentiment analysis system.
3. Segment the *clitics* from the named entities for easier matching. This step is pretty important for languages like Arabic. However, the segmentation error risk is high as we're working on the entity-level e.g. "الميتا" is likely to be segmented as "الـ+ميتا". Thus, the quality of this step depends a lot on the segmentation system performance.
4. Unify the entities, one entity may be mentioned multiple times in an article and with varying polarities. Aggregate all the polarities of an entity in one value by averaging them.
5. Represent each article with a vector, where each element corresponds to the polarity of a named entity in the article. The dimension of this vector is equal to the number of named entities that appeared in all the articles belonging to the same event.
6. After all, we calculate the similarity between two articles by finding the distance between their vectors.
7. An additional step could be weighting the entities according to their significance in the event. To get these weights, for each article we generate the same vector as 5, but with the TF-IDF values of the entities as its elements. Then, all the articles' vectors are aggregated together by averaging them. The resulted vector contains the entities weights. These weights can be given to the entities by multiplying the weights vector in an element-wise manner with the vector generated in 5.
### Multi-Features
*What features can we consider other than the ones related to the sentiment polarities?*
The intuitive answer is the words used for expressing an opinion, these words should differ in articles with different viewpoints, and are in two types:
- Words that usually have polarities, and then they can be captured by sentiment lexicons. For instance:
| **Positive** | **Negative** |
| --- | --- |
| جيد | سيء |
| ذكي | غبي |
- Words that express different viewpoints, but may not be captured by sentiment lexicons. For instance:
| **Positive** | **Negative** |
| --- | --- |
| حزب | ميليشيا |
| شهيد | قتيل |
In general, features related to the article content can be represented using TF-IDF, or maybe using distributed embeddings like doc2vec. While this representation includes the whole text, the first-type words can be separated in its own features vector too after extracting them using the sentiment lexicons.
Related feature to the first-type words are:
- The overall article polarity, which is the averaged polarities of the polar article's words, extracted using a sentiment lexicon.
- The count of positive and negative words in the article. Again, using a sentiment lexicon.
Another answer to this question is considering metrics that indicate an opinion type, which should, in theory, differ in articles with different viewpoints, such as:
- Abusive: Demeaning and abusive language.
- Obscenity: Obscene or profane language.
- Racism: Demeaning and abusive language targeted towards a particular ethnicity.
- Sexism: Demeaning and abusive language targeted towards a particular gender.
- Insults: Scornful remarks directed towards an individual.
- Threats: Expressing a wish or intention for pain, injury, or violence against an individual or group.
All these features along with the entity-level sentiment analysis features, discussed in the previous section, should be aggregated in one weighted sum which is the ranking formula.
**Limitation**s:
The weights in the ranking formula are determined manually.
### Pairwise Classification
To overcome the limitation of the previously suggested method, we suggest training a pairwise classifier. The input of the classifier is the different similarity values between two articles, all concatenated together in one vector, while the output is a value indicating whether the two articles agreed on one viewpoint or not. The classifier will automatically determine the weights of the ranking formula by learning to link the input to the corresponding output.
*How to get training data?*
As we suppose having an event detection system for the first phase of the problem, we can collect an event detection test collection, where we label articles reporting the same event, you can refer to [Building a Test Collection for Event Detection Systems Evaluation](/blog/building-a-test-collection-for-event-detection-systems-evaluation) for the details about the collection process. Now, for each event in the test collection, we generate all the possible articles pairs, then for each articles pair, we decide whether they disagreed or not.
**Pros**:
- Treating the problem classification rather than a ranking task has pros on the performance. Basically, for an event of size *N*, we have to rank the whole *N* articles against a new incoming article to get a ranked list of contrary view articles, instead, we can compare against the previously recognized *P* viewpoints whose number *P* << *N*, where each viewpoint can be represented for example using the average of its articles vectors.
- Using the similarities as the features of the classifier should ideally isolate it from variations in genre, regions, etc. thus we can train a single global classifier.
### Entity Linking
Until now we depend heavily on the named entities in our suggestions, Let's consider the named entity linking task and see how it can improve the algorithm's understanding to the named entities.
Named entity recognition (NER) systems identify and classify named entities occurrences in text into pre-defined categories. On the other hand, Named entity linking (NEL) will assign a unique identity to entities mentioned in the text. In other words, NEL is the task of linking entities mentioned in a text with their corresponding entities in a knowledge base.
*But how does it help us in our mission?*
Consider the following example:
| In **English** | In **Arabic** |
| --- | --- |
| Ajoubair said during a press conference in the British capital London “we are convinced based the evidence we have that the Iranian Army had a role to play in Aramko attacks” | وقال الجبير في مؤتمر صحفي بالعاصمة البريطانية لندن “مقتنعون من خلال الأدلة الموجودة لدينا بتورط الجيش الإيراني في هجمات أرامكو |
One of the entities extracted from this text would be “الجيش الإيراني” (the Iranian Army) and we can see that the sentence convenes a negative sentiment towards this entity.
Now consider the following excerpt describing the same event just in a different way:
| In **English** | In **Arabic** |
| --- | --- |
| Aljoubair repeated his accusations of the Iranian armed forces in being behind the Aramco incident | و جدد الجبير اتهامه للقوات العسكرية الإيرانية بالوقوف وراء حادثة أرامكو |
In order to have a better understanding of the articles, it is important for our contrary view detection system to be able to say that the entities “الجيش الإيراني” (the Iranian Army) and “للقوات المسلحة الإيرانية” (the Iranian armed Forces) represents the same physical thing.
For further information about NEL and how to implement it, refer to our post [Aspect Detection and Named Entity Linking (NEL): Using SPARQL and DBpedia](/blog/aspect-detection-and-named-entity-linking-nel-using-sparql-and-dbpedia).
## Conclusion
In this post, we introduced a new task namely: Contrary View Detection in its two phases: 1. articles grouping and 2. disagreement based ranking. We presented our suggestions to solve this problem including methods to be used and features to represent the problem.
### An Implementation of a News Stream Sequence Clustering Algorithm
- URL: https://mohammadshaker.com/en/blog/an-implementation-of-a-news-stream-sequence-clustering-algorithm
- Date: 2020-01-26T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Sequential clustering for news streams assigns each incoming article to an existing event cluster or creates a new one, without knowing the number of clusters in advance. We implement this online algorithm with incremental centroid updates and a similarity threshold, then evaluate it against our custom Arabic news test collection.
#### Content
In a previous post, we discussed a [News Stream Sequential Clustering Algorithm](/blog/news-stream-clustering-sequential-clustering-in-action). In this post, we're discussing the details of implementing this algorithm with minimal tuning, and showing the results produced by this implementation.
Along with this post, we're evaluating our results and tuning the algorithm's hyper-parameters using a test collection built in a previous effort according to a method described in [Building a Test Collection for Event Detection Systems Evaluation](/blog/building-a-test-collection-for-event-detection-systems-evaluation).
## Applying The Algorithm
In this section, we're discussing the implementation details of the previously proposed news stream clustering algorithm.
### Features Extraction
In this algorithm, each news article is represented by a set of subvectors indicating textual features, and three values indicating the time feature.
To capture an event in text we relied on two textual features: the TF-IDF of the main content of the news article, and the named entities extracted from that content. We also considered the time feature which is pretty important for event detection. For more information about the event detection problem features you can refer to our article [Event Detection in Media using NLP and AI](/blog/event-detection-in-media-using-nlp-and-ai).
### Evaluation Metric
In order to tune the hyper-parameters, we need an evaluation metric. The chosen evaluation metric based on the existence of a labeled test dataset (news articles assigned to the events they're reporting).
We evaluated the clustering in the following manner: let **Tp** be the number of correctly clustered-together document pairs, let **Fp** be the number of incorrectly clustered-together document pairs and let **Fn** be the number of incorrectly not-clustered-together document pairs. Then we report precision as **Tp**/(**Tp**+**Fp**), recall as **Tp**/(**Tp**+**Fn**) and F1 as the harmonic mean of the precision and recall measures.
### Hyper-parameters Tuning
The algorithm has many hyper-parameters:
- The weights of the features, that denote which feature matters more in measuring the similarity between an article and a cluster. We Fixed all these weights to "1" in this experiment, which means all our features were equally contributing to the similarity measure.
- ***σ*** is a parameter denotes a standard deviation of a Gaussian function related to the time features, and it refers to the difference \_in days\_ between two dates. Was fixed to "3 days".
- **μ** is the mean of a Gaussian function related to the time features. Was fixed to "0".
- We only tuned **τ** which is the similarity threshold, if the similarity between a document and cluster exceeds it, the document would be added to that cluster. We tuned this parameter on our test collection using the F1 measure metric introduced above: the best value for **τ** was "4.5", where F1 = 60%.
## Results
Let's show you some of our results after applying the algorithm using the previously discussed configurations.
### Our Results on The Test Collection
We got ~160 clusters for ~2000 news articles in the test collection. Examples of the good clusters produced by the model are shown below:
Dhamar attack
Hurricane Dorian
The death of Robert Mugabe
Here is another example where the model got confused:
Peace negotiations with Taliban + Donald Trump sacks John Bolton
The model merged between two distinct events, which is a common error probably caused by the following reasons altogether, that appeared in the previous example:
- Many of the named entities are shared between the two events e.g. ترامب، أمريكا
- Many of the non-named entities are shared between the two events too e.g. قرار
- They are too close to each other in terms of time.
Follow up on our suggestions to improve the model performance in the last section of this post.
### Representing The Event
In this section, we're trying to summarize the main event using expressive words or phrases, in simple manners.
#### Word-Cloud
Let's try to represent the event as a word-cloud.
Let's take one event as an example:
Dhamar attack
The following word-clouds were generated by aggregating the TF-IDF vectors of the titles/contents of all the articles in the event.
Following are word-clouds generated using the content:
1-Gram
2-Grams
3-Grams
Following are word-clouds generated using the title:
1-Gram
2-Grams
3-Grams
#### The Best Word/Phrase
Let's try to represent the event in one word/phrase, which should be the most expressive one.
First, we'll try to take the word/phrase with the highest TF-IDF score, after aggregating the TF-IDF vectors over all the articles title/content in the event.
For the same cluster shown in the previous section:
Choosing the word/phrase from the content:
| **1-Gram** | **2-Grams** | **3-Grams** |
| --- | --- | --- |
| ذمار | التحالف السعودي | اللجنة الدولية للصليب |
Choosing the word/phrase from the title:
| **1-Gram** | **2-Grams** | **3-Grams** |
| --- | --- | --- |
| ذمار | المبعوث الأممي | عشرات القتلى والجرحى |
One example was enough to prove that these words/phrases are generic and incomplete to represent the whole event.
Another way we tried, is searching for the longest words sequence that is occurred in more than half of the event articles.
For the same previous example we got:
اللجنة الدولية للصليب الأحمر
And for the following cluster:
Hurricane Dorian
We got:
الإعصار دوريان
So, sometimes we're getting good results using this method but other times we're still getting generic phrases. Moreover, in some clusters where there are huge shared text parts among the articles e.g. when they are telling that someone said something, we got very huge phrases.
#### **The Best Named Entities:**
For each article in the event, we extracted three types of named entities: people's names, locations, and organizations' names. We expressed each event by the named entities from each type that were mentioned in more than half of the articles in the event.
To show the results let's consider the same clusters above:
Dhamar attack:
| **Person** | **Location** | **Organization** |
| --- | --- | --- |
| يوسف، عبد | ذمار، اليمن، صنعاء، السعودية | الحوثي |
Hurricane Dorian
| **Person** | **Location** | **Organization** |
| --- | --- | --- |
| دوريان، دونالد، ترامب | دوريان، الولايات، المتحدة، فلوريدا | - |
A new example cluster:
Amazon Fires
| **Person** | **Location** | **Organization** |
| --- | --- | --- |
| - | الأمازون | - |
Where the used named entity extractor is FARASA.
#### Expressing The Events with Concepts
This time we're trying to express each event by the concepts presented in its articles. For this task, we're using [Wikifier](http://wikifier.org/) which is a web service that annotates a text with links to relevant Wikipedia concepts.
For each article related to the event, we extracted its concepts. We used the page rank value associated with each concept as its weight, as the page rank values of concepts can be used to disambiguate them, you can refer to [1] for more information about the Wikifier algorithm. We aggregated the weights of the presented concepts over all the articles. After all, we chose the concepts that their weights exceeded a threshold to express the event. For the following results, we chose the threshold to be 0.004.
Again to show the results, let's consider the same example clusters above using Wikifier:
| **Event** | **Expression** |
| --- | --- |
| **Dhamar attack** | حوثيون، اللجنة الدولية للصليب الأحمر، السعودية، اليمن، صنعاء، التدخل العسكري في اليمن، الأمم المتحدة، ذمار، الجزيرة (قناة)، |
| **Hurricane Dorian** | إعصار، باهامس، فلوريدا، الولايات المتحدة، المحيط الأطلسي، الشرق الأوسط، جزر أباكو، دونالد ترامب، النمسا، دوريان، |
| **Amazon Fires** | جايير بولسونارو، غابة الأمازون، رسام، البرازيل، منظمة الصحة العالمية، غابة، الشرق الأوسط، نظام بيئي، العالم، طفل |
We can see huge improvement over all the previous techniques.
Pros:
- Concepts are unified. In other words, in the previous techniques, we faced the following problem which does not exist for this technique:
إعصار <> وإعصار
- Concepts are always complete. In the previous technique some extracted phrases were not complete e.g. :
الولايات، المتحدة، اللجنة الدولية للصليب
Errors:
- The presence of concepts that are mentioned in the articles but are not really key concepts e.g. الجزيرة (قناة)
### The Clusters' Count Effect on The Clustering Time
As the number of the clusters increases over time with the developing news stream, the number of the similarity measurements increases too, which is reflected in the clustering time of an incoming article.
The following graph, which was generated using the test collection, shows this effect visually:
Obviously the number of clusters affects the clustering time badly. One way to reduce this effect is by eliminating the outdated clusters which we're discussing in the next section.
#### Eliminating The Outdated, Stale Clusters
We defined an outdated cluster as a cluster that has not been updated for a while. The question emerges here: how much is that "*while*"?
Let's say we have a cluster with one article that has not been updated for three days and another cluster with 100 articles that has not been updated for three days too. Clearly, these can't be treated in the same way, while the small cluster seems to be outdated, the larger one does not seem so, as large events usually tend to expand over a long period of time and may not be updated daily.
Hence, for eliminating the outdated clusters we should consider both the cluster last update date and its size.
For this goal let's propose the following two normalized Gaussian functions:
F1 = exp(- (x - ***μ1***)^2 / (2\****σ1***^2) )
Where the input is the difference between today's date and the publish date of the newest article that was added to the cluster, in terms of days.
***μ1***: is the mean, was set to "0"
***σ1***: is the standard deviation, denotes a number of days, was set to "3".
The longer the cluster has not been updated, the higher the value of this function is.
F2 = 1 - exp(- (x - ***μ2***)^2 / (2\****σ2***^2) )
Where the input is the number of articles in the cluster.
***μ2***: is the mean, was set to "0"
***σ2***: is the standard deviation, denotes a number of articles, was set to "5".
The larger the cluster is, the higher the value of this function is.
We combined both functions in one formula as the following:
**outdated**\_**indicator**(***c***) = **F1**(***c***.news\_cluster\_date) + 2 × **F2**(***c***.articles\_num)
The maximum value of this function is "3".
Finally, we defined a threshold, if the outdated indicator value doesn't exceed it, the cluster would be eliminated, otherwise, it would be kept. And we set this threshold value to "2", by manually tuning it.
As an additional constraint, any cluster that was updated in the last three days no matter what its size is, shouldn't be eliminated.
To prove the effectiveness of the proposed method we generated inputs for F1 in the range [3, 50], and for F2 in the range [1, 100], then we took their combinations to generate inputs for the outdated indicator formula.
We randomly sampled these combinations and plotted them in the following graph where each point denotes a combination, while it's color denotes whether the cluster was kept or removed to its outdated indicator value.
You can see that large clusters can be preserved for about a month maximum without any update, while small clusters are eliminated after a few days from being frozen.
After applying this cutoff method while clustering our test collection, the following graph was generated:
The maximum number of the stored clusters we reached is ~40, while it was the whole 170 clusters before the cutoff. And the maximum clustering time for an incoming document is ~0.2s, while it was ~0.7s before the cutoff.
In comparison with the first graph without the cutoff, there is definitely a huge improvement in the clustering time. The cutoff tries to keep the clustering time stable. A major point to add is that applying this cutoff did not affect the clustering performance in any way.
### Our Results on The Real-World Data
In this section, we're validating the algorithm and the chosen hyper-parameters against the real-world data.
We tried to detect the events in Almeta's news feed data over three months. Here are some of the produced clusters:
Some clusters make sense, like the following:
Dhaka attack
Paris protests
Russia Maneuvers Center
While others tend to be totally disastrous, like the following one:
Another one:
In general, our evaluation on the test collection was better than the actual results we got on real-world data. This can be interpreted because of the noise in the real-world data (e.g. articles from different domains other than news) while the test collection is clean and determined.
Another common error type that we noticed is articles that their similarity with their cluster exceeded the threshold just because of the time features, although their textual similarity with the actual articles in the cluster is too low. *Thus, we applied a textual similarity threshold such that if the textual similarity between the incoming article and the cluster doesn't exceed it the article won't be merged into the cluster.* In the following example is one cluster before and after applying the textual threshold:
Before applying the textual threshold:
Tunisia Elections
Where red titles are unrelated.
After applying the textual threshold:
Tunisia Elections
We got rid of all the unrelated documents.
The textual threshold was tuned manually on the real-world data and was set to **1.5**.
As a try to improve the results on the real-world example, to be more specific in our clustering we tried a 1 more final thing:
Instead of taking just the TF-IDF of the whole article content we took both the title and the content TF-IDF and concatenate them with each other. This resulted in higher F1 on the test collection 62% taking the same hyper-parameters as above. However, the results on the real-world examples improved for just a little bit with this step.
## Further Improvements
For further experimenting and improvements, we're listing some aspects we would like to explore more in the future.
**Better clustering decisions**:
- Improving our decision about the best matching cluster for an incoming document, by learning to weight our similarity features. By treating the problem of making this decision as a ranking algorithm, where the ranked items are the clusters, which are ranked with reference to an incoming document. A ranking model can be trained to perform this task, and it would implicitly learn to weight our similarity features. For further information, you can refer to [News Stream Clustering – Sequential Clustering in Action](/blog/news-stream-clustering-sequential-clustering-in-action).
- Improving our decision about when to make a new cluster for an incoming article and when not to. By replacing the threshold parameter ***σ*** with a binary classifier that has the complete ownership about making this decision. Also, for further information, you can refer to [News Stream Clustering – Sequential Clustering in Action](/blog/news-stream-clustering-sequential-clustering-in-action).
**What about further experiments with the features?**
- Different combinations of the different parts of the document e.g title, body, first paragraph.
- What's the effect of some preprocessing steps like lemmatization and stemming?
- Does using a dense representation like doc2vec improve the performance?
## Conclusion
In this post, we discussed the details of our implementation of a news stream clustering algorithm for event detection with minimal tuning. we also presented the produced results by this implementation on both test and real-world data.
## References
[1] Janez Brank, Gregor Leban, Marko Grobelnik. Annotating Documents with Relevant Wikipedia Concepts. Proceedings of the Slovenian Conference on Data Mining and Data Warehouses (SiKDD 2017), Ljubljana, Slovenia, 9 October 2017.
### Automatic Sentence Paraphrasing
- URL: https://mohammadshaker.com/en/blog/automatic-sentence-paraphrasing
- Date: 2020-01-26T00:00:00.000Z
- Tags: arabic, nlp, Human-Written
Automatic sentence paraphrasing rewrites text to preserve meaning while changing surface form — useful for data augmentation, summarization, and plagiarism detection. Rule-based systems use hand-crafted transformation templates; statistical approaches learn from parallel corpora; neural models like seq2seq generate more fluent paraphrases but require large training sets.
#### Content
Most of us are not good writers, and if you are like me you may have sometimes struggled in communicating your ideas in a written form.
The new field of automated paraphrasing can provide a solution to this issue. In the optimal version of a software paraphrasing system, you would be able to write your essay on your own, not-very-great style and then give it to the paraphrasing system. This system will modify your text to make it shorter, more formal, or less biased.
In this post, we will explore in a very shallow manner how you can build your own paraphrasing system.
# What is a Paraphrase
The concept of paraphrasing is most generally defined as semantic equivalence (having the same meaning): A paraphrase is an alternative entity in the same language expressing the same meaning as the original form. And these paraphrases may occur at several levels.
- Words having the same meaning are usually referred to as **lexical paraphrases** or, more commonly: synonyms. For example: (hot, warm) and (eat, consume). However, this concept can also include hyperonymy, where one of the words is either more general or more specific than the other, for example, (reply, say) and (landlady, hostess).
- **Phrasal paraphrase** refers to phrases sharing the same meaning. Although these phrases usually form full syntactic phrases like (work on, soften up) and (take over, assume control of ) they may also be patterns with linked variables. For example, Y was built by X, X is the creator of Y, etc.
- Two sentences that represent the same semantic content are termed **Sentential paraphrases**, for example, (I finished my work, I completed my assignment). Although it is possible to generate very simple sentential paraphrases by simply substituting words and phrases in the original sentence with their respective semantic equivalents, it is significantly more difficult to generate more interesting ones such as (e.g. He needed to make a quick decision in that situation, the scenario required him to make a split-second judgment)
# Automatic Paraphrasing
## What is It
The task of automated paraphrasing is basically an umbrella that covers multiple more specialized tasks. Most notable are:
- Paraphrases extraction: when given a large corpus of text the goal is to extract paraphrases i.e sentences, phrases, or words that have the same meaning, this is important for data generation and thus supporting the rest of these tasks list.
- Sentence compression: here, given a sentence, the goal is to reduce the size of the sentence either through deleting unimportant parts or merging words, see [this list of examples](http://homepages.inf.ed.ac.uk/mlap/cgi-exp/examples1.html) of how sentence compression works in English from a linguistic point of view.
- Style transfer: following the success of style transfer in the domain of computer vision and image processing, the goal of style transfer in NLP is to altering a piece of text to follow a certain style of writing (imitating a famous writer, a more formal/ less formal way)
- Text simplification: this can be considered a subtask of style transfer where the goal is making the text more readable and simpler. Here is an example of this task
- Original: Owls are the order Strigiformes, comprising 200 bird of prey species.
- Simplified: An owl is a bird. There are about 200 kinds of owls.
## Why Should I Invest in It
Some of the main applications of an automated paraphrasing system include:
- Query and Pattern Expansion: the automatic generation of
query variants for searching in an IR system (e.g. search engine)
*Original: circuit details*
*Variant 1: details about the circuit*
*Variant 2: the details of circuits*
- Expanding Sparse Human Reference Data: some tasks need human-generated output for either training or evaluation, most famous of which is machine translation and abstractive summarization, and due to the great cost of human labor the ability to use paraphrasing to generate more synthetic data for either evaluation or training is greatly important.
- Improving the responsiveness of chatbots.
- Summarization: some of the paraphrasing approaches employ compression to generate the paraphrases this can be helpful as a level on top of an extractive summarizer.
## How Can We Evaluate It
Similar to many NLG tasks, given a generated paraphrase and a reference paraphrase metrics we want to measure how accurate is the generated paraphrase relative to the reference. BLEU, ROUGE and METOR are usually used here. If you are not familiar with these metrics, please follow [this gentle introduction](https://medium.com/explorations-in-language-and-learning/metrics-for-nlg-evaluation-c89b6a781054).
# Get Me Some Data
In order to build any machine learning system, the main concern is having a large-enough and diverse enough dataset for this model to learn from. In the case of paraphrasing, each data point is composed of a set of 2 sentences/paragraphs that hold the same semantic content but are written in a relatively different manner.
## Available Datasets
In this section, we include dataset specifically developed for the task of paraphrasing or those that can be utilized for this task with minimal effort.
### For Multiple Languages
- PPDB [1] is an automatically extracted database containing millions of paraphrases in 16 different languages. The goal of PPBD is to improve language processing by making systems more robust to language variability and unseen words. The entire PPDB resource is [freely available here](http://paraphrase.org/#/download) but since the dataset is automatically generated the quality might be low, **and it supports Arabic.** This dataset includes phrasal level paraphrases that are relatively short, here are some examples of pairs from this dataset:
البرنامج
الجارى لتامين ||| البرنامج
الراهن لتامين
الجناة والضحايا : المساءلة ||| المجرمون والضحايا: المساءلة
الرسائل
المقدمة بموجب المادة 22 ||| الرسائل
الواردة بموجب المادة 22
- In [2], the authors build the first dataset for Arabic sentence compression, which is a very small dataset of fewer than 100 documents, and it is not publicly available yet.
### Specifically for English
- Wikipedia offers a simpler version for English learners [Zu et al. (2010)](http://aclweb.org/anthology/C10-1152) a compiled a parallel corpus with more than 108K sentence pairs from 65,133 Wikipedia articles, allowing **1-to-1 and 1-to-N alignments**. The latter type of alignments represents instances of sentence splitting. The original full corpus can be found [here](https://www.informatik.tu-darmstadt.de/ukp/research_6/data/sentence_simplification/simple_complex_sentence_pairs/index.en.jsp). This simple Wikipedia is only available for the English language.
- [The Grammarly’s Yahoo Answers Formality Corpus (GYAFC)](https://github.com/raosudha89/GYAFC-corpus) was constructed by identifying 110 000 informal responses containing between 5 and 25 words on Yahoo Answers. Each of these was then rewritten to use more formal language by Amazon Mechanical Turk workers. However, this corpora covers only the English language.
- Another dataset of Shakespeare plays and their modern translations are also available. This corpus contains 17 plays and their modernizations from . And versions of eight of these plays from . While the alignments appear to mostly be of high quality, they were still produced using automatic sentence alignment which may not perform the task as proficiently as a human. This dataset includes around 21k sentences in total.
### Utilizing Datasets from Other Tasks
- The data usually used in machine translation takes the form of pairs of sentences, where one sentence is from the source language e.g. English and the other is from the target language, say Arabic. In this case, for every source sentence there is only one target reference. However, some translation datasets like [this](https://catalog.ldc.upenn.edu/LDC2003T17) or [this](https://catalog.ldc.upenn.edu/LDC2006T04) include multiple references for the same source sentences, i.e. every example in the dataset looks like this (E, F1, F2, F3), where F1-3 are different translations of the same sentence E. It is accurate to assume that different translations (references) constitute paraphrases of each other. The same logic applies to the summarization task were multiple references are usually used to measure the performance of the model.
- The same idea can be extended to other tasks e.g. MSCOCO was originally an image captions dataset, containing over 120K images, with five different captions from five different annotators per image. All the annotations toward one image describe the most prominent object or action in this image, which is suitable for the paraphrase generation task (However, we know of no image captioning datasets in Arabic).
## How Can We Collect More Data
In this section, we explore various methods to build a dataset from paraphrasing models.
### Bi-lingual Pivoting
Assuming 2 languages: E and F, the basic assumption is that any two source strings e1 and e2 that both translate to a reference string f1 have a similar meaning. In [3], the authors apply this approach to the EUROPARL machine translation dataset, basically the English-Arabic translation pair. They first align the Arabic and English sentences, to get a sentence to sentence translation pairs, then they find Arabic sentences that translate to the same English sentence (or to relatively close sentences).
### Using Automatic Machine Translation
The authors in [3] as well suggest a second method also using a machine translation dataset. In very simple terms, given a translation data between, say French and English, for every sentence pair (Fi, Ei), use machine translation to translate both sentences to Arabic, this will generate two new sentences in Arabic (A1, A2), that can be considered paraphrases.
The approach is depicted in the following figure:
The authors evaluated 200 pairs generated through this approach (which is very low considering that the generated dataset had over 2M pairs). They claim that these sentences were found to be grammatical or near grammatical with minor mistakes, but were found to be highly similar in meaning. Also, they were sufficiently different in surface form and thus usable for training data for the a seq-to-seq paraphrasing model.
The main drawbacks of such method are:
- The possibility that the en→ar and fr→ar translation tools we used were trained on the same Arabic data with similar model architecture and parameters, in which case the output would not be diverse enough to be useful as input to the paraphrasing system.
- The possibility that the French source was translated to English and then to Arabic.
- The low quality of the machine-translated sentences.
One way to improve this approach is to start with a translation dataset that already includes Arabic sentences. For example, if we have an English-Arabic translation pair, only the English sentences need to be machine-translated to Arabic. This way we can gain two improvements:
- Lower effort in machine translation.
- Higher quality: basically by training the paraphrasing model (seq2seq) to take the machine translated sentences as input and generate their human labeled paraphrases, we can ensure the fluency of the model output.
### Duplicate Questions/Posts on Online Community Sites
Many online community sites like stack exchange and quora label duplicated questions in order to control their content. These duplicated questions represent very close paraphrases, for example, in Quora alone there are over 400K lines of potential question pairs that are duplicated to each other. However, we know of no sites that practices this behavior in Arabic.
### Using Parallel Translations
In [Evaluating prose style transfer with the Bible], the authors utilize the various versions of the bible as a source for paraphrases. The bible is split into chapters and verses, where each verse is between 1 and 6 sentences. These verses can be considered parallel across the various version. Overall, there is 31,102 verses in the bible. Each version of the bible is understood as embodying a unique writing style. The versions in this corpus were created with a wide range of intentions. Versions such as the Bible in Basic English were written to be simple enough to be understood by people with a very limited vocabulary. Other versions, like the King James Version, were written centuries ago and use very distinctive archaic language. The authors didn’t publish the [whole dataset](https://github.com/keithecarlson/StyleTransferBibleData) mainly because of distributional licenses, However, they claim that re-scraping the dataset is relatively simple.
There are 34 stylistically different versions of the English bible (new and old testament). However, in the case of Arabic, there are 8 versions, with only 6 of them covering both the new and old testament, two of these versions are available through [bible Gateway](https://www.biblegateway.com/versions/). The following figure table illustrates the variations between the Arabic bible versions:
| Translation | [Genesis](https://en.wikipedia.org/wiki/Book_of_Genesis) 1:1–3 (التكوين) | [John](https://en.wikipedia.org/wiki/Gospel_of_John) 3:16 (يوحنا) |
| --- | --- | --- |
| *Van Dyck* (or *Van Dyke*) فانديك | في البدء خلق الله السماوات والأرض. وكانت الأرض خربة وخالية وعلى وجه الغمر ظلمة وروح الله يرف على وجه المياه. وقال الله: «ليكن نور» فكان نور. | لأنه هكذا أحب الله العالم حتى بذل ابنه الوحيد لكي لا يهلك كل من يؤمن به بل تكون له الحياة الأبدية. |
| *Book of Life* كتاب الحياة | في البدء خلق الله السماوات والأرض، وإذ كانت الأرض مشوشة ومقفرة وتكتنف الظلمة وجه المياه، وإذ كان روح الله يرفرف على سطح المياه، أمر الله: «ليكن نور» فصار نور. | لأنه هكذا أحب الله العالم حتى بذل ابنه الوحيد لكي لا يهلك كل من يؤمن به بل تكون له الحياة الأبدية. |
| *Revised Catholic* الترجمة الكاثوليكية المجددة | في البدء خلق الله السموات والأرض وكانت الأرض خاوية خالية وعلى وجه الغمر ظلام وروح الله يرف على وجه المياه. وقال الله: «ليكن نور»، فكان نور. | فإن الله أحب العالم حتى إنه جاد بابنه الوحيد لكي لا يهلك كل من يؤمن به بل تكون له الحياة الأبدية. |
| *Good News* الترجمة المشتركة | في البدء خلق الله السماوات والأرض، وكانت الأرض خاوية خالية، وعلى وجه الغمر ظلام، وروح الله يرف على وجه المياه. وقال الله: «ليكن نور» فكان نور. | هكذا أحب الله العالم حتى وهب ابنه الأوحد، فلا يهلك كل من يؤمن به، بل تكون له الحياة الأبدية. |
| *Sharif Bible* الكتاب الشريف | في البداية خلق الله السماوات والأرض. وكانت الأرض بلا شكل وخالية، والظلام يغطي المياه العميقة، وروح الله يرفرف على سطح المياه. وقال الله: «ليكن نور.» فصار نور. | أحب الله كل الناس لدرجة أنه بذل ابنه الوحيد لكي لا يهلك كل من يؤمن به، بل ينال حياة الخلود. |
| *Easy-to-read Version* النسخة سهل للقراءة | في البدء خلق الله السماوات والأرض. كانت اﻷرض قاحلة وفارغة. وكان الظلام يلفّ المحيط، وروح الله تحوّم فوق المياه. في ذلك الوقت، قال الله: «ليكن نور.» فصار النور. | فقد أحبّ الله العالم كثيرا، حتى إنه قدّم ابنه الوحيد، لكي لا يهلك كل من يؤمن به، بل تكون له الحياة الأبدية. |
Finally, there are some dialectical translations of the bible in the public domain, like [this](http://worldbibles.org/language_detail/eng/arz/Normal+Egyptian+Arabic). However, we are not sure if they are in an electronically accessible format.
# How Can We Implement This Task
The approaches to this task can be easily categorized into either supervised, unsupervised and ad-hoc.
## Supervised
In this type, the task boils down to a machine translation task, where given a relatively large parallel dataset of sentences, where the target sentences represent a paraphrase of the source sentence, for example, training a model to translate between van-dyke and the “book of life” versions of the bible.
The two predominant approaches to this task are the same two types of approaches used in machine translation, namely statistical machine translation using frameworks like [Moses](http://www.statmt.org/moses/) or the neural machine translation approach following the encoder-decoder structure, in this type framework like [OpenNMT](http://opennmt.net/) can be of help.
The earliest work on paraphrases generation used statistical machine translation techniques to generate novel paraphrases [4]. More recently, phrase-based statistical machine translation software was used to create paraphrases, in [5], they report a relatively high BLEU score of 50.4.
The Shakespeare dataset was used with a Seq2Seq model [6]. Their
results are impressive, showing improvement over statistical machine
translation methods as measured by automatic metrics. They experiment
with many settings, but in order to overcome the small amount of
training data, their best models all require the integration of a
human-produced dictionary which translates approximately 1500
Shakespearean words to their modern equivalent.
In [7], the authors use both Moses phrase-based statistical translation and a neural seq2seq model, in order to overcome the limited size of corpora when training the Seq2Seq model they utilized the tagging trick from [8], where they build a single seq2seq model to translate between the different version in a “multi-lingual” fashion, this is accomplished by adding a small token to the start of the source sentence to represent the target version, for example, if the target style for a verse pair is that of the American Standard Version, we start off the source sentence with an ‘ASV’ token. They report BLEU results as high as 71 for the statistical machine translation variant and 52 for the Seq2Seq model. The following figure shows some of the results
## Unsupervised
The main bottleneck in many of the paraphrasing tasks is the lack of parallel datasets to train a supervised model, especially for tasks like style transfer, where the target of the training can be specialized (following a certain style like Shakespeare, or following a certain sentiment positive vs negative).
Most of the available methods to tasks like style transfer are unsupervised, meaning that it is possible to train such models without any parallel data. [This is an amazing list](https://github.com/fuzhenxin/Style-Transfer-in-Text) of papers and code-bases in this field.
## Other
- One of the main methods used to generate more synthetic data and a baseline for comparison with the other methods is lexical synonym replacement. Basically, for some words in the article (mostly adjectives and nouns), the system will automatically change the word using one of its synonyms from lexicons like WordNet, datasets like the aforementioned PPDB can also be helpful since it contains paraphrases on the lexical, phrasal, and syntactic levels. To get an idea of the performance of such an approach you can try some of the online tools that use them like [this](https://paraphrasing-tool.com/), [this](https://www.prepostseo.com/paraphrasing-tool), or [this](https://www.rewritertools.com/paraphrasing-tool).
- The literature regarding the field of sentence compression includes several rule-based methods to delete or replace parts of the sentence in order to reduce its size, some of the older approaches to the task apply modifications on the parse tree of a sentence, these approaches employ rules in the form of context-free grammar CFG that is either written by hand or learned in a stochastic manner, see the following figure for an example. However, we couldn’t cover these approaches in detail within the scope of this post but further details can be found in this amazing summary [9].
- Another line of work is sentence fusion [10], [11], this task is heavily used in multi-document summarization. This approach is used in [Quill-bot](https://quillbot.com/). Check it out to get an idea of the performance of these systems. In very simple terms the main approach is clustering sentences into semantically similar clusters and then for each cluster a single word is generated. For example in [10], the main approach used is to generate a word graph from the cluster sentences and then finding the shortest path in this graph to create the fused sentence. For example, for the following two sentences the graph is shown in the following figure, again note that covering this task in detail is beyond the limits of this post and the reader is left to explore the references.
- In Asia Japan Nikkei lost 9.6% while Hong Kongs Hang Seng index fell 8.3%.
- Elsewhere in Asia Hong Kongs Hang Seng index fell 8.3% to 12,618
## Conclusion
In this article, we have explored in a very broad manner the field of paraphrase generation. This is an extremely wide field with multiple sub-tasks in it. However, if you decided to build a paraphrasing system yourself then the following tips might help you:
- There is a decent size of parallel datasets. And due to the nature of this task, it is also possible to reuse datasets built for different tasks like machine translation or image captioning.
- There are several ways to scrap datasets for this task and some of them like machine translation or bi-lingual pivoting can generate large amounts of datasets.
- Furthermore, we believe that further research can reveal other ways to collect data for this task.
- Methods based on supervised machine translation reports high results in the literature. One weird phenomenon is the fact that statistical machine translation models constantly outperform their neural counter-parts throughout the reviewed literature. This can be attributed to the small size of training data.
- Unsupervised NMT models like the approaches used in style transfer are very complicated. Yet the reported results in the literature are rather decent in comparison with the supervised approach. Furthermore most of them include publicly available code-bases.
- Some of the ad-hoc methods like phrase substitution have their own pros and cons:
- Very simple and can be easy to build and debug.
- Their quality depends on the quality of lexicons used like word2vec or the paraphrases dataset like PPDB, the main fear is that if synonyms are not suitable for the context.
- They only perform phrasal or lexical changes to the sentences.
The following table compares the aforementioned approaches:
| Approach | Cons | Pros |
| --- | --- | --- |
| Neural machine translation using OpenNMT | Require parallel data this data can be generated in various ways or we can use some ready-made datasets like PDBB. These models often depends of the domain of the training data mainly because their vocabulary is limited and can often overfit horribly when working with new domains. The quality of the model depends of the size and quality of the training dataset | Relatively easy to train models and optimize them using the machine translation data generation approach it is possible to build datasets for a large number of domains can achieve high performance when large data is available |
| Statistical Machine translation | Also requires parallel data. Moses is not very simple to learn | Reported results in the literature is higher |
| Unsupervised methods | Relatively complicated and hard to understand. Their performance is lower (yet comparable) to their supervised counterparts | Several code-bases are available, and can be relatively used in an off-the-shelf manner does not require parallel data, but does require un-parallel corpora of the 2 styles |
| Phrase substituting | Relatively simple to implement | Their results are often low, the generated sentences are templatic and can generate paraphrases not suitable for the context |
##
## References
[1] J. Ganitkevitch, B. Van Durme, and C. Callison-Burch, “PPDB:
The paraphrase database,” in *Proceedings of the 2013 Conference
of the North American Chapter of the Association for Computational
Linguistics: Human Language Technologies*, 2013, pp. 758–764.
[2] R. Belkebir and A. Guessoum, “TALAA-ASC: A sentence
compression corpus for Arabic,” in *2015 IEEE/ACS 12th
International Conference of Computer Systems and Applications
(AICCSA)*, 2015, pp. 1–8.
[3] F. Al-Raisi, A. Bourai, and W. Lin, “NEURAL SYMBOLIC ARABIC
PARAPHRASING WITH AUTOMATIC EVALUATION,” *Comput. Sci. Inf.
Technol.*, vol. 1, 2018.
[4] C. Quirk, C. Brockett, and W. Dolan, “Monolingual machine
translation for paraphrase generation,” in *Proceedings of the
2004 conference on empirical methods in natural language processing*,
2004, pp. 142–149.
[5] S. Wubben, A. Van Den Bosch, and E. Krahmer, “Paraphrase
generation as monolingual translation: Data and evaluation,” in
*Proceedings of the 6th International Natural Language Generation
Conference*, 2010, pp. 203–207.
[6] H. Jhamtani, V. Gangal, E. Hovy, and E. Nyberg, “Shakespearizing
modern language using copy-enriched sequence-to-sequence models,”
*ArXiv Prepr. ArXiv170701161*, 2017.
[7] K. Carlson, A. Riddell, and D. Rockmore, “Evaluating prose
style transfer with the Bible,” *R. Soc. Open Sci.*, vol. 5,
no. 10, p. 171920, 2018.
[8] M. Johnson *et al.*, “Google’s multilingual neural
machine translation system: Enabling zero-shot translation,”
*Trans. Assoc. Comput. Linguist.*, vol. 5, pp. 339–351, 2017.
[9] E. Pitler, “Methods for sentence compression,” 2010.
[10] M. T. Nayeem, T. A. Fuad, and Y. Chali, “Abstractive
unsupervised multi-document summarization using paraphrastic
sentence fusion,” in *Proceedings of the 27th International
Conference on Computational Linguistics*, 2018, pp. 1191–1204.
[11] K. Thadani and K. McKeown, “Supervised sentence fusion with
single-stage inference,” in *Proceedings of the Sixth
International Joint Conference on Natural Language Processing*,
2013, pp. 1410–1418.
### Contrary View Detection Based on Document Similarity
- URL: https://mohammadshaker.com/en/blog/contrary-view-detection-based-on-document-similarity
- Date: 2020-01-26T00:00:00.000Z
- Tags: almeta.io, arabic, nlp, Human-Written
Contrary view detection using document similarity works in two stages: first find articles on the same topic via high similarity, then re-rank by dissimilarity of viewpoint signals — sentiment polarity, entity framing, and opinion markers. We implement and evaluate this pipeline on Arabic news, measuring how well similarity metrics alone can approximate ideological opposition.
#### Content
After reading a news article on your favourite news aggregator or your news site of choice, Most of the current news aggregators allows you to read other articles that are related to the one you have already read, these suggestions are usually articles describing the same event.
Here at Almeta, one feature we believe that can be helpful for our readers is the ability to suggest not only related articles but also articles that discuss the same event but suggest different viewpoints or opinions.
We believe that reading different opinions given by different anchors is the best way to find the whole truth.
This article is the third part of our series on detecting contrary view articles you are strongly advised to review the previous articles to learn:
- [What is contrary view-point detection and what are the theoretical ways we can implement such a system?](/blog/towards-contrary-view-detection-in-news)
- [Can we use multidimensional topic modeling to automatically find different view-points among articles discussing the same event.](/blog/contrary-view-detection-based-on-vodum)
In a very intuitive sense given a set of articles that discuss the same event for example “Palestinian-Israeli conflect” the goal is to cluster the articles accourding to the side they support.
In this article, we will follow a different path to finding articles that express varying viewpoints. Given a certain numerical representation of each article, it is possible to measure the distance between every 2 articles. For a given article A and a specific “optimal” representation and distance function, we can assume that the articles closest to article share the same viewpoint with it while articles that are further-away have different view-points.
The goal of this article is to investigate these “optimal” representation and distance functions.
# **Experiment steps**
Similar to the previous article all of our experiments are carried out on a small in-house developed dataset of news articles classified into the event they represent, you can review [our previous article](/blog/viewpoint-topic-and-opinion-discovery-in-an-opinionated-document) to learn about the details of building this dataset.
Each of the experiments reported here follow the following steps:
1. Text is fully normalized (Alefs, English, nums and punctuation).
2. Articles are converted into a numerical representation. We will test different types of representations e.g. TFIDF.
3. Select a specific distance metric based on the aforementioned representation.
4. Calculate distance matrix (using all the articles of the event) using the selected representation using the aforementioned distance function.
5. Filter the articles to eliminate anomalies using the distance matrix
6. For each article (row in the distance matrix) order the other articles based on their similarity to this article.
7. Return the ordered articles.
# **Experimentation and Analysis**
In this section we will discuss the different findings we got:
## **Controversial Events**
What we found is that most events (in our dataset) does not include any contrary views and that all the news outlets simply report the same events.
- After manual inspection. Out of 81 events present in the manual test data only 10 of them included at least 2 different viewpoints, most of the events talked about non-political stuff earthquakes, Nobel prize, … or talked about non-controversial global politics like trump impeachment and British elections (These events are not controversial **for Arabic news anchors), so please note that all the following experiment is evaluated using 10 events only which is not optimal**
- Even in the case of controversial events. Not all the articles picked a side and some do maintain neutrality. For example in the case of نبع السلام event, some articles talked about the cease-fire agreement brokered by Trump. This sub-event has no contrary view since all the outlets report it in the same manner. See the following example
أعلن الرئيس الأميركي دونالد ترامب أنّ نائبه مايك بنس ووزير الخارجية مايك بومبيو سيتوجّهان إلى تركيا للتفاوض على وقف لإطلاق النار في شمال سوريا، وذلك غداة اتصاله بنظيره التركي رجب طيب أردوغان.
ووفق بيان البيت الأبيض فان بنس سيدعو خلال المفاوضات في تركيا لوقف إطلاق النار فوراً في سوريا وبدء المحادثات.
كما أن بنس سيلتقي مع الرئيس التركي أردوغان يوم الخميس، وسيؤكد التزام ترامب بإبقاء العقوبات الاقتصادية على تركيا لحين التوصل إلى حل، بحسب البيان.
- Furthermore, in some event the opposing news outlets simply ignore the whole event and does not describe it at all like for example when describing the Houthi attack in southern Saudi Arabia in the operation نصر من الله. The Saudi side simply ignore the whole event.
**These issues raises the question of usability of this system across all the articles, since most of them won’t have a contrary view and furthermore we have no way of saying if an article do or don’t have a contrary view.**
## **Article representation and distance metric**
### **Simple Baseline using TFIDF:**
We started by trying a very simple representation using only TFIDF and with cosine distance However, this simple baseline gave very reasonable results see the following example (original article top, contrary view bottom),
أعلنت الوكالة الدولية للطاقة الذرية اليوم الاثنين عن مضاعفة ايران عدد أجهزة الطرد المركزي المتقدمة لتخصيب اليورانيوم.
وقال المتحدث باسم الوكالة فريدريك دال إن مفتشي الوكالة تحققوا يوم أمس من تركيب 22 جهاز طرد مركزي من طراز "آي آر – 4" في منشأة تخصيب اليورانيوم في نطنز، مقابل 11 خلال الأشهر القليلة الماضية.
وظلت أعداد طرازات أجهزة الطرد الأخرى عند نفس مستوياتها تقريباً.
ودعا كورنيل فيروتا، القائم بأعمال مدير الوكالة الدولية للطاقة الذرية إيران إلى "الرد فوراً" على أسئلة الوكالة المتعلقة ببرنامجها النووي.
وزار فيروتا إيران أمس والتقى وزير الخارجية محمد جواد ظريف وعدد من كبرار المسؤولين الإيرانيين للاطلاع على أحدث التدابير التي أعلنتها إيران بشأن تعزيز الأبحاث والتطوير المرتبط بتخصيب اليورانيوم.
ويأتي هذا الإعلان بعد اتخاذ ايران الخطوة الـثالثة لخفض التزاماتها النووية، رداً على انتهاكات واشنطن للاتفاق النووي.
++++++++++++++++++++
0.7658784302734492
وصفت لندن، اليوم السبت، قرار إيران تشغيل أجهزة طرد مركزي متطورة لزيادة مخزونها من اليورانيوم المخصب بأنه "مخيب للغاية". وذكرت الخارجية البريطانية في بيان أن هذا التطور "الذي يخالف التعهدات في الاتفاق المبرم مخيب للغاية، في الوقت الذي نسعى فيه مع شركائنا الأوروبيين والدوليين لنزع فتيل الأزمة مع إيران".
وكان مسؤول إيراني رفيع، أكد في مؤتمر صحافي عقد السبت أن بلاده بدأت، الجمعة، خطوة جديدة في تخفيض التزاماتها في الاتفاق النووي الإيراني.
كما لوح المتحدث باسم منظمة الطاقة الذرية الإيرانية، بهروز كمالوندي، برفع نسبة تخصيب اليورانيوم، قائلاً "إيران لديها القدرة على تخصيب اليورانيوم بما يتجاوز 20%".
وفرض الاتفاق قيوداً على برنامج إيران النووي المثير للجدل مقابل رفع العقوبات عنها، لكنه بدأ يتفكك منذ انسحاب الولايات المتحدة منه العام الماضي وتحركها لتضييق الخناق على تجارة النفط الإيرانية، لإجبارها على تقديم تنازلات أمنية أوسع نطاقاً.وبدأت إيران منذ مايو/ أيار الماضي في تقليص التزاماتها ببنود الاتفاق رداً على حملة الرئيس الأميركي، دونالد ترمب، التي ينتهج فيها وضع الحد الأقصى من الضغط على طهران منذ الانسحاب من الاتفاق، والتي شملت إعادة فرض العقوبات لإجبار طهران على العودة للمفاوضات.
وكان الاتحاد الأوروبي، حض الخميس، إيران على "التراجع" عن التخلي عن التزاماتها بموجب الاتفاق النووي مع الدول الكبرى. وقال المتحدث باسم المفوضية الأوروبية، كارلوس مارتن رويز دي غورديخويلا، للصحافيين في بروكسل "إننا نعتبر هذه الأنشطة غير متوافقة (مع الاتفاق النووي) ونحضّ إيران على التراجع عن هذه الخطوات والامتناع عن أي خطوات إضافية تقوض الاتفاق النووي".
Out of the 10 controversial events this simple method produced good results for 8 of them which is relatively good.
The 2 events where the system gave non-optimal results were relatively large and heterogeneous (Lebanon riots and Iraq riots).
Mainly this simplistic approach sometimes fails with larger clusters like Lebanon riots where different details from different periods are all clustered together. In this case this simple representation usually assumes different events as contrary since they are the furthermost.
However, this behavior is more related to purity rather than size, for example in the case of other events like Egypt riots (demonstrations against president Sisis) event nearly all the articles describe the same thing and we can see a clear distinction between the different views, see the following example:
في تظاهرات محدودة بعدد من المدن المصرية، خرج مئات المحتجين رافعين شعارات منددة بالحكومة والرئيس عبد الفتاح السيسي
في تجاوب مع دعوات أطلقت على شبكات التواصل الاجتماعي، قادها رجل الأعمال والمقاول المصري المقيم في الخارج محمد علي، ورغم تأكيد ناشطين ومنظمات حقوقية اعتقال العشرات واستخدام قنابل الغاز في تفريق التظاهرات إلا أن بعض التقارير وصفت التعامل الأمني بغير الحاد..فما هي دلالات تجاوب الشارع هذه المرة مع دعوات التظاهر؟ وهل ستكون لها تبعات في الأيام المقبلة؟ وما حقيقة ارتباطها بأطراف تقليدية وأخرى غير تقليدية؟
++++++++++++++++++++
0.969631312181884
اشتعلت منصات التواصل الاجتماعي في مصر بعبارات الاحتجاج السياسي اليوم الجمعة، تزامنا مع دعوات معارضة صدرت من خارج البلاد للتظاهر والخروج في احتجاجات ضد الرئيس عبد الفتاح السيسي.
وصبغ ناشطون مصريون موقع تويتر باللون الأحمر تعبيرا عن المشاركة في الاحتجاج ضد السيسي، في حين تحدث ناشطون عن بدء أولى مبادرات الاحتجاج في الشارع، وسط دعوات لضرورة عدم التراجع.
وبرزت وسوم عدة بث المشاركون عبرها رسائل الاحتجاج. وتصدر وسم "جمعة الغضب" قائمة التداول (الترند) المصري على تويتر بعد ظهر اليوم بعشرات آلاف التغريدات، استلهاما للجمعة الأولى التي أعقبت اندلاع شرارة ثورة 25 يناير/كانون الثاني 2011.
…
Furthermore This simple baseline can even detect less obvious contrary views. For example in the following example anti-Qatar outlets describe the failure of the World Athletics Competition in Qatar (TOP). While Aljazeera simply describes the winning athlets and so on, this is achievable due to the shift in vocabulary.
أثارت استضافة قطر لبطولة العالم لألعاب القوى، والتي تعد بمثابة اختبار عملي لمدى استعدادها لبطولة كأس العالم لكرة القدم 2022، الكثير من الأسئلة حول مدى قدرتها على استضافة أكبر حدث رياضي في العالم، بحسب ما جاء في تقرير أعدته مراسلة شبكة "بلومبيرغ" في الدوحة، سيمون فوكسمان.
وأوضح التقرير أن الزوار الأجانب، الذين حضروا المنافسات الأولى من بطولة العالم لألعاب القوى في الدوحة، انتقدوا مدرجات المتفرجين الشاغرة في استاد خليفة مكيف الهواء، فضلاً عن العثرات والأخطاء في التشغيل والتنظيم.
…
++++++++++++++++++++
0.945053891962537
شهد اليوم الختامي لمونديال ألعاب القوى الذي استضافته قطر تحقيق العدّاء الجزائري توفيق مخلوفي الميدالية الفضية في سباق 1500 م ليرفع الغلّة العربية إلى 7 ميداليات.
ويضمّ سجلّ مخلوفي ذهبية سباق 1500 م في أولمبياد لندن، وفضيتين في سباقَي 1500 م و800 م في ريو دي جانيرو.
وقال مخلوفي بعد التتويج: "أنا فخور بهذه الميدالية، فخور بهذا الانجاز بعد عودتي من الاصابة وبعد وقت طويل بعيداً عن المنافسات".
### **Using Entity Level Sentiment Analysis**
The main motive behind this approach is as follows: we can assume that each article describes a set of entities (political figures, locations, ideologies, ...) and every article expresses a certain opinion towards each of these entities.
We can intuitively state that article the have the same stance towards the same entities express the same view-point.
In this setting the representation of the article is found as follows:
1. **Find the named entities** in the article, we use an in-house name entity detector for this task.
2. **Calculate polarity**: Here we use our entity-level sentiment analysis system, given a certain named entity (target) this system can calculate the polarity of the article towards this target (does the article express positive, negative or neutral sentiment towards the target). by sending these named entities to ELSA we can find the polarity expressed by the article for each entity. You can learn more about our ELSA system by reviewing this [previous article](/blog/aspect-level-vs-entity-level-sentiment-analysis).
3. **Average the polarities across the named entities** The ELSA system returns the different polarities for each occurrence of the target in the text, for example, the name “سعد الحريري” might be mentioned multiple time in the text, in some sentences the article might praise the target while in others it might criticize it. In order to get a single polarity of the article towards the target, we need to average the different polarities expressed for different occurrences. Note that this averaging is done based on the exact matching of (Normalized) strings and thus “سعد\_الحريري” and “وسعد\_الحريري” represent 2 different entities. A better approach is to combine these different mentions into the same “Aspect” using named entity linking for example. We discuss the idea of named entity linking in detail in this [previous article.](/blog/aspect-detection-and-named-entity-linking-nel-using-sparql-and-dbpedia) However, for simplicity, we choose not to do this in this experiment. After this step is done each named entity will have a polarity score assigned to it.
4. **Find a common vocabulary:** we collect a vocabulary of the most prominent named entities selected from a large corpus of news article crawled between Feb 2018 and April 2018.
5. Using the common vocabulary we can construct a unified vector representation of the article based on its named entities polarities each row represents a named entity where the value of the row is the polarity of this named entity in the article (0 if the named entity does not occur in the article)
6. Then we can follow the same steps in the TFIDF steps to order the articles based on their distances.
However manual inspection (on the limited set) revealed worse than baseline performance. There are 2 main factors that may contribute to the bad results:
- **Propagation of error:** between NER and ELSA polarity which itself can be really noisy (due to the averaging process).
- **Vector sparsity:** The combined common vocabulary includes around 21k entry while debugging we found many articles where a lot of the named entities fall outside this common vocabulary due to segmentation issues, or simply because they are Strange, this issue can be resolved by building an independent vocabulary for every event, but we didn’t try this due to time constraints.
### **Improving TFIDF**
The main motivation is that most of the baseline errors happen when there are several topics (sub-events) in the cluster. Intuitively we can consider 2 articles to be contrary if they are talking about the same thing but represent different takes.
To approximate this for 2 articles to be contrary, we can assume that we have 2 different distances. A positive distance that we should minimize to make sure the 2 articles discuss the same topic. And a negative distance that should be large to make sure the 2 articles have different viewpoints. thus a combined distance can be defined as:
combined\_distance = normalize(lambda \* negative\_distance – (1-lambda) \* positive\_similarity)
Where:
positive\_similarity = 1- positive distance aka cosine distance
And lambda is a weighting hyper-parameter between 0 and 1, the larger value gives more emphasis to the baseline and lower value gives more emphasis to the similarity between articles.
This step is extremely similar to MMR in multi-document summarization (MDS) research if you are interested in this idea you can [review our previous article](/blog/multi-document-summarization-the-what-why-and-how) on MDS and how we used MMR to diversify the summary.
We assumed that if 2 articles have the same-named entities then they are likely to be discussing the same topic and thus the above mentioned NER vocabulary could be used for positive similarity, and the original TFIDF is the negative distance, we did the following:
1. Take the TFIDF representation as a negative representation this will give us the negative vector.
2. Induce a positive representation: we used the above-mentioned method to extract named entities and run them into a count-vectorizer this will give us a positive vector.
3. Calculate the 2 distance matrices using the positive and negative representations, we use cosine distance for both of them, this will give us the positive and negative distance matrices.
4. Calculate the positive similarity matrix Similarity\_pos = 1 – Distance\_pos.
5. Calculate the final distance and sort the articles based on it.
We tried 3 different values for Lambda 0.2, 0.5, 0.8 and found 0.8 to be the best value.
This operation was extremely difficult since it requires manually evaluating the events for each value. Nonetheless, this approach results were similar or worse than the original baseline, this might be due to the issues with the NER vectorizer representation (sparsity).
## **Cluster impurity**
The Impurity of clusters causes the system to fail regardless of the article representation mainly because an anomaly would be far from all the articles, the following example illustrates the idea
أكد الرئيس التركي رجب طيب أردوغان اليوم الأحد أن التهديدات الغربية بفرض عقوبات على بلاده وحظر تصدير الأسلحة إليها لن تدفعها لوقف عملية "نبع السلام" المستمرة ضد المقاتلين الأكراد في شمال سوريا.
وقال أردوغان في خطاب متلفز "بعدما أطلقنا عمليتنا واجهنا تهديدات مثل عقوبات اقتصادية وحظر على بيع الأسلحة.. ومن يعتقدون أن بإمكانهم دفع تركيا للتراجع عبر هذه التهديدات مخطئون كثيرا".
...
++++++++++++++++++++
0.9597737245179864
الجزيرة نت-طهران
أعلنت إيران فجر اليوم الخميس عن انتهاء القوات البرية في الجيش من مناورات مفاجئة تحت شعار "هدف واحد.. رصاصة واحدة" في المنطقة المحاذية لتركيا بهدف "تقييم جاهزية قواتها القتالية"، مؤكدة أن المناورات تبعث رسائل منفصلة إلى الشعب الإيراني وأعدائه.
وأعلنت دائرة العلاقات العامة للجيش الإيراني صباح أمس الأربعاء انطلاق مناورات مفاجئة للقوات البرية في منطقة أورومية شمال غربي البلاد بالقرب من حدود تركيا.
...
To resolve this issue for each article we first calculate the sum of its distances from all the other articles using the similarity matrix.
Intuitively an anomaly should be far from all the articles and thus this sum should be overly large, if we assume that these sums follow a gaussian distribution (which they seem to do see the following figure), then we can state that most of the data should fall between avg+2\*std and avg-2\*std.

Based on this reasoning any article whose distances sums is greater than avg+2\*std can be considered as an anomaly and discarded from consideration. After applying this mechanism on the previous example the resulting contrary article is as follows
أكد الرئيس التركي رجب طيب أردوغان اليوم الأحد أن التهديدات الغربية بفرض عقوبات على بلاده وحظر تصدير الأسلحة إليها لن تدفعها لوقف عملية "نبع السلام" المستمرة ضد المقاتلين الأكراد في شمال سوريا.
وقال أردوغان في خطاب متلفز "بعدما أطلقنا عمليتنا واجهنا تهديدات مثل عقوبات اقتصادية وحظر على بيع الأسلحة.. ومن يعتقدون أن بإمكانهم دفع تركيا للتراجع عبر هذه التهديدات مخطئون كثيرا".
...
++++++++++++++++++++
0.927225742795052
دانت الخارجية السورية بشدة عزم تركيا تنفيذ عملية عسكرية داخل أراضي سوريا، واعتبرت أن ذلك سيفقد أنقرة دور الضامن في عملية أستانا ويوجه ضربة قاصمة للعملية السياسية برمتها.
وقالت الخارجية، اليوم الأربعاء، إن سوريا تدين بأشد العبارات "التصريحات الهوجاء والنوايا العدوانية للنظام التركي والحشود العسكرية على الحدود السورية التي تشكل انتهاكا فاضحا للقانون الدولي وخرقا سافرا لقرارات مجلس الأمن الدولي التي تؤكد جميعها على احترام وحدة وسلامة وسيادة سوريا".
...
**Please note the following:**
- That this simple ad-hoc method can fail in removing all anomalies due to 2 factors:
- If the cluster is extremely noisy or impure and include many events like Lebanon riots all the articles will be far from each other and thus the std will be too big that it includes everything
- Even in less noisy clusters, anomalies can have very similar language, for example, the following article is selected a part of a cluster describing the resignation of the Saudi power minister, and since it is talking about economy, oil and power it is similar enough to fall under the limit of avg+2\*std while still being an anomaly
أصدر خادم الحرمين الشريفين الملك سلمان بن عبدالعزيز - يحفظه الله - عدداً من الأوامر الملكية يوم الجمعة الماضي، تضمنت تعيينات لقيادات حكومية جديدة وإعفاء آخرين من مناصبهم، وإعادة تشكيل لأجهزة حكومية، وترقية بعض منها لمستوى إداري أعلى من حيث الصلاحيات والمسؤوليات.
من بين ما شملت الأوامر الملكية على مستوى الأجهزة الحكومية، فصل الطاقة عن الصناعة والثروة المعدنية، وتغيير تبعاً لذلك مسمى "وزارة الطاقة والصناعة والثروة المعدنية" ليصبح "وزارة الطاقة" وإنشاء وزارة باسم "وزارة الصناعة والثروة المعدنية" تُنقل إليها الاختصاصات والمهمات والمسؤوليات المتعلقة بقطاع الصناعة ونشاط الثروة المعدنية، لتنفصل بذلك عن وزارة الطاقة. ويُعد هذا قراراً صائباً للغاية، باعتبار أن وزارة الطاقة والصناعة والثروة المعدنية قبل الفصل بين مهامها ومسؤولياتها،
…
- That this step does indeed remove some authentic articles sometimes. However, We believe it is a fair trade-off since by keeping anomalies all the other articles will have bad results.
- This step can as well get rid of problematic articles where say very small content is provided for example from the event of Tunisian elections this was the only excluded article
ضيوف الحوار
الصحافي التونسي منصف سليمي، بون - ألمانيا.
المحلل السياسي مُولى غازي. تونس.
برنامج "ستوديو الحدث"، ثمرة التعاون بين "
DW
- Some events have ambiguity due to combining articles from multiple genres. For example in the manual data of event detection. The event of Aramco attacks includes both political and economic articles. These cases were ignored mainly because this case will not occur in production since event-detection is limited by the political genre.
## **Does This Article Has A Contrary View**
As mentioned above for many events there isn’t a contrary viewpoint and even for controversial events some articles are simply neutral, in order to resolve this we tried 2 filtering schemes, both done of the TFIDF baseline since it gave the best results, the goal is to set a limit L where only if 2 articles have a distance greater than L these 2 articles can be faithfully considered as contrary, we tried to set the value of L using 2 ways:
- Direct values: e.g. if 2 articles distance is greater than 0.95 then they are contrary. This method can be used in theory to filter out non-controversial events since they should, in theory, have all their articles very similar.
- Event-based: for every event assume the values of the pairwise distances to be some sort of a population and find the 75th, 80th, … percentile. This method can’t detect non-controversial events. However, it can in theory detect non-controversial articles within a controversial event. since they would be similar to the 2 different viewpoints (because they are neutral).
However, both methods failed in filtering the non-contrary articles. Please note that all the evaluation is done manually. We could not pin a specific value of L (neither global or event-based). For example using a value of 0.8 some articles might have a contrary article whose distance is 0.7, while other articles might not be contrary but have a distance greater than 0.9. This is mainly due to the simplicity of representation of TFIDF. Basically if 2 articles are far they might be contrary but they might as well be describing different sub-events.
Theoretically this method should work if we have a better way to calculate the distance between the articles.
# **Conclusions**
- Most of the events out there are not controversial and do not include multiple viewpoints.
- The settings of this experiment is not-optimal at best since there are only 10 events with possible controversial articles and that all the evaluation is done manually. This setting greatly complicated the process of hyperparameters search and evaluation. We are a bit sceptical about the results we have now. But we believe if we have an annotated dataset for evaluation. We might be able to get something usable based on these simplistic approaches.
- Overall the best method was the plain TFIDF baseline. You can think of this method as a reversed content-based suggestion system. This method can’t be used to present contrary views. if it to be deployed we should instead market as “other viewpoints ”, “other takes on the event”, … and not exact contrary views.
- We tried filtering out the non-controversial events or articles but this is not possible using the current method.
- It is possible to reduce the impurity in clusters using a simple ad-hoc anomaly detection
- Event impurity and large events are also major factors in the quality of the results.
##
### Contrary View Detection Based on VODUM
- URL: https://mohammadshaker.com/en/blog/contrary-view-detection-based-on-vodum
- Date: 2020-01-26T00:00:00.000Z
- Tags: almeta.io, analysing-content,-not-publishers, biased-and-sensational, ml, nlp, research, values, Human-Written
VODUM (Viewpoint and Opinion Discovery Unsupervised Model) extends LDA by jointly modeling topics and viewpoints, making it a natural fit for contrary view detection. We apply VODUM to Arabic news to surface articles that cover the same event from opposing ideological angles, comparing it against document similarity baselines on viewpoint divergence metrics.
#### Content
While reading the news each one of us perceives it in a different manner. We have our own biases and we tend to search for information that confirms our previous beliefs. Thus different people might have drastically different viewpoints of the same event, and this effect extends to the news anchors which are also subject to this kind of bias.
Here at Almeta, we are in no position to decide if a certain viewpoint is right or wrong. However, we believe that reading different viewpoints of the same event, listening to arguments that are contrary to your beliefs and accepting differences, can be a step toward really uncovering the whole picture.
In this article we will explore the following question: “for a
given news article can we automatically find other articles that
describe the same event but have different view-point”. This
feature will hopefully allow our users to read different articles
describing the same event yet having different takes on it, and we
leave the judgment of these viewpoints as valid or not to the users.
# What is VODUM
VODUM stands for Viewpoint and Opinion Discovery Unified Model, we
will gloss over it for now but if you wish to have more details
please review [our
previous article](/blog/viewpoint-topic-and-opinion-discovery-in-an-opinionated-document)
This model falls into the category of multi-dimensional topic
models, where the model tries to capture multiple aspects of the text
including topic, opinion and viewpoint. In very simple terms the
model is a generative model based on 3 different probability
distributions:
- When a writer starts an article she selects a certain viewpoint from the list of possible viewpoints based on a probability distribution PI, this distribution can be described by a vector of probabilities, This selected viewpoint is the same over the whole document.
- After selecting a viewpoint the author selects for every new sentence a certain topic from a list of viewpoint-specific topics based on a second distribution over topics called theta.
- Finally, after selecting a viewpoint and a viewpoint-specific topic, the author proceeds to select the words of the sentence, at this step the author chooses between 2 types of words Topical (related to topics), or Opinion-words also using 2 different distributions phi0 and phi,
This means that the overall model can be described using 4 distributions shown in the following figure,
where:
The goal of the training step is to estimate these distributions using the training text in an unsupervised manner similar to any other topic model with the goal of selecting a set of distributions that minimizes the perplexity as much as possible.
In the inference phase, the text is fitted to the different
distributions and then the most probable viewpoint/topic
distributions are returned.
The hyper-parameters of the training process are the number of topics, the number of viewpoints and the priors of the aforementioned distributions.
To use this model for contrary view detection a model should be fitted to the articles of each event using a pre-defined number of viewpoints (for now 2) and then articles are clustered using VODUM. Afterwards, for each article that is assigned by VODUM to a certain viewpoint, we can suggest other articles from other viewpoints.
# Experiment set up
All the results and experiments are carried out on a small in house dataset initially used for event-detection task, you can read more about this dataset creation [from our previous article](/blog/building-a-test-collection-for-event-detection-systems-evaluation). We took a sample of this set that only contains controversial events (events where at least 2 different viewpoints exist) plus a random sample of non-controversial events. All the evaluation is done in a manual way.
For VODUM we used the [official
implementation.](https://github.com/tthonet/VODUM)
For each of the events we extract all the articles and then do the
following:
- Text is fully normalized (Alefs, English, nums and Punctuation)
- Stop words removal
- Lemmatization to improve topic modelling
- Word type selection: recall that the model needs to build different distributions for topical and opinion models, we followed 2 different methods to select word type:
- In the [paper](http://www.irit.fr/publis/SIG/2016_ECIR_TCBPS.pdf), the authors used POS tags to select word type, where all Nouns are considered Topical where the other POS types are Opinion-words. This is clearly an inaccurate approximation
- Another option is to use sentiment lexicons where subjective words (words that contain sentiment) are considered opinion words while non-subjective words are considered topical words. We used Arsenl lexicon for this step.
Surprisingly, the first approach gave better results than the latter one this can be attributed to the high number of OOVs.
- VODUM hyper-parameters: The problem in automating the process of VODUM is the selection of hyper-parameters for every new event. We can either use a pre-defined set of parameters for all the events or try to select them automatically based on the event information:
- for the priors hyper-parameters, we always use the values reported in the paper for all the events
- for the number of viewpoints we used 2 for now across all the events (This is obviously not accurate)
- the number of topics is selected based on the number of articles in the event using simple linear interpolation, in the original paper the authors used 50 events to represents data of nearly 600 articles, so for example for an event of 30 articles we use 3 topics.
# Results
The results we observed were not satisfactory even on the manually
annotated events, here are the issues with this approach.
## Low cluster purity
The view-points generated by this model are not very pure, for small controversial events encompassing less than 30 articles we have observed some good results. However, as the event size starts to grow the 2 viewpoints becomes increasingly noisy. There are 2 main possible causes for error:
### The simplicity of the model
VODUM is a very simple unsupervised model and thus we don’t expect its performance to be extremely high. In the original paper, the authors run the system several times for every topic they report a maximum accuracy of 75% across all topics and across all the runs and an average of just above 60%.
### Problems in selecting Hyper-parameters
As noted before the selection of Hyper-parameters is done in an automatic manner. The number of view-points and distribution priors are kept the same across all the events, and the number of topics is selected only based on the number of articles in the event.
This simple way of selecting
hyper-parameters is not optimal since it ignores the actual content
of the articles. We
have verified this experimentally by manually optimizing the
hyper-parameters of one event “نبع
السلام” and
while this step did improve the results it can’t be done on
production since it requires
manual intervention.
One way to select the best parameters of the model is by relying on the perplexity of the model after training. this is the only “confidence-like” value that this algorithm can produce. In theory, we can run a Hyper-parameters optimization process using this metric to guide the selection of hyper-parameters in an automatic manner. However we have chosen not to do this for 2 main reasons:
- Firstly, even if this approach does work the computational overhead of training the topic model (multiple times) and carrying out grid-search to find the parameters with the lowest perplexity is too high, especially if we consider the issue of event retraining (see below).
- Secondly, the perplexity while being an indicator of how pure are the view-points. It can be a noisy indicator. In the previous example of “نبع السلام,” we tried to optimize the parameters only using the perplexity as an indicator. However, for some parameters values with lower perplexities didn’t correspond to improvements in the clustering performance (as measured by manual inspection).
### Inability to classify events as controversial:
An important issue is trying to
guess if a certain event is controversial or not (have multiple
viewpoints).
The algorithm we suggested at the start of this article works only on controversial events. But when considering events where there is no real difference between the coverage of multiple anchors (natural disasters for example) i.e. when all the articles have the same view-point it is impossible to find a “contrary view” to any of the articles.
Therefore, a major first step is to decide if a certain event is controversial or not, and then if it is controversial try to find the different viewpoints in it.
As we discussed above the only indicator of the confidence that we can extract from this model is perplexity. One way to detect non-controversial events is to measure the model perplexity on the articles of this event after training the model. If the perplexity exceeds a certain limit we can say that a model with multiple viewpoints couldn’t fit the data and thus there isn’t multiple viewpoints, i.e. the events with values higher than such a limit can be considered non-controversial.
However, in practice, this approach fails. Mainly because while events that have a better separation between different view do have lower perplexity. Other factors influence the perplexity value including the number of topics in the event, events size and even different runs of the model.
### Growing the event
Assume
that we start with a small event of say 20 articles and we train a
model that can correctly fit this event. The main issue is expanding
this model when the event size increases (new
articles are published that describe this event).
While
retraining the model every say 10 new articles is not computationally
extensive. Re-selecting
the hyper-parameters properly (in
a manual manner) is
problematic.
And this requires answering the following questions: when does an article introduce a new topic? Or a new viewpoint? And can we detect this automatically?
### Major Tasks in Dialectical Arabic Processing
- URL: https://mohammadshaker.com/en/blog/major-tasks-in-dialectical-arabic-processing
- Date: 2020-01-26T00:00:00.000Z
- Tags: almeta.io, arabic, ideas, nlp, research, Human-Written
Dialectal Arabic NLP is harder than Modern Standard Arabic because dialects vary widely across 22 countries, lack standardized spelling, and have far fewer labeled datasets. This survey covers the best available systems and benchmarks for dialect identification, sentiment analysis, and machine translation across Arabic varieties.
#### Content
There have been some recent advancements in Dialectical Arabic processing across various NLP tasks, in this article the goal of this article won’t be to explore any particular task but to explore as many tasks as possible and give an overview on the best available systems, datasets, and methodologies for each of them.
## Properties of Dialectical Arabic
The Arabic language is a well-known example of diglossia, in this type of languages the formal variety of the language, which is taught in schools and used in written communication and formal speech (religion, politics, etc.) differs significantly from the informal varieties that are acquired from the home.
These
informal variants are used in all other types of communication either
in speaking or in social media. The spoken varieties of the Arabic
language (which we refer to collectively as Dialectal Arabic) differ
widely depending on the geographic location and the socio-economic
conditions of the speakers, and they can be quite different from the
formal variety known as Modern Standard Arabic (MSA) فصحى.
While NLP applications in the MSA have had some success, these applications can’t be directly applied to the Arabic dialects because there are some significant differences across the dialects themselves and between the dialects and MSA, these difference can be in:
- Phonology: the way letters are pronounced in different dialects
- Morphology: the way the words are composed from smaller parts for example in the Palestinian dialect the suffix ش is usually used to negate an action for example عملت (I did) can be negated by adding ش to its end creating عملتش (I did not), there is no similar mechanism in MSA and most of the other dialects, most of the dialects have special ways to build words
- Lexicon: the commonly used words to describe an entity, for example, the word knife in MSA is سكين but in the Syrian dialect, it is موس in Palestinian it is خوصة
- And even syntax: the way words are combined together to build a sentence, basically the MSA follows the Verb-Subject-Object structure (VSO) in the sentences for example رمى الولد التفاحة which translates word-to-word to (throw the boy the apple) however in most of dialects both VSO and SVO is permitted meaning both الولد زت التفاحة (the boy throw the apple) and زت الولد التفاحة (throw the boy the apple) are ok while the former being the most common
These differences render some dialects incomprehensible to the speakers of other dialects and make NLP systems built on MSA not usable for handling dialectical content found in everyday speech or social media.
## Textual Dialect Identification:
In
this task, the given an input text the system should try to identify
the dialect of the text.
The
applications of such a system includes:
- Geotagging of reviews/tweets: giving a coarse view of the users location and cultural background,
- As pre-processing step to enable specialized dialect-specific models for the other tasks,
- And finally it can be used as a post-processing step in systems such as ASR, autocorrect, auto-complete and to a lower extent in OCR as identifying the dialect can allow us to use specialized error correction (language models).
The task has 2 main categories based on the scope of the identification:
- **Coarse classification**: in which the goal is the identification of the main Arabic dialects (Levent, Gulf, Egypt, Maghrib, Iraqi, other [Sudan, Somalia, Yemen, …], and MSA)
- **Fine-grained classification**: here the goal is to identify more specific tags like the country or the city/region
### Datasets
This task has a relatively large number of sizable datasets mainly due to the simplicity of building them in an automatic fashion. The following table shows the freely available datasets for dialect classification
| Name | Aligned (can be used in translation) | size | Notes |
| --- | --- | --- | --- |
| [ADD](https://www.lancaster.ac.uk/staff/elhaj/corpora_files/ArabicDialectsDataset.zip) | No | 10K | covering the 5 main dialects (Egy, Gulf, Levant, North Africa and MSA) |
| [Shami](https://github.com/GU-CLASP/shami-corpus) | No | 66k | tweets automatically scraped and annotated based on geo-location only for the 4 Levant dialects |
| [PADIC](https://smart.loria.fr/corpora/) | Yes | 6.4k | PADIC includes four dialects from the Maghreb: two from Algeria, one from Tunisia, one from Morocco and two dialects from the Middle- East (Syria and Palestine). |
| [comparable Wikipedia](https://github.com/motazsaad/comparableWikiCoprus) | No | 10k | Wikipedia Articles from Arabic and Egyptian Wikipedia aligned using [wikipedia aligner](https://github.com/motazsaad/WikiDocsAligner) although the articles represent a “translation” of the same Wikipedia entry the aligned documents while discussing the same entry rarely have the exact same content and thus this can’t be used for translation |
| [DART](https://www.dropbox.com/s/jslg6fzxeu47flu/DART.zip?dl=0) | No | 27.5k | The covered dialects are the 5 main dialects the data is composed of 25k automatically annotated tweets and phrases with 2.5k human-annotated data |
| [VarDial2016](http://ttg.uni-saarland.de/vardial2016/dsl2016.html) | No | ? | Automatic Speech recognition transcripts covering the Egyptian, Gulf, Levantine, and North-African, and Modern Standard Arabic (MSA) |
| [AOC](https://github.com/sjeblee/AOC/tree/master/stuff-from-omar/annotated-aoc-data) | No | Huge | Covers the same 5 main dialects, and is created by scraping comments found on online sites This dataset contains 2 subsets: a huge 46M automatically annotated set a smaller 100k manually annotated set through Mturk Furthermore the dataset includes for each of the comments the following fields: the original article URL the subtitle of the article and thus can be used in other tasks like topic modelling and classification, keyphrase extraction, … |
| [MADAR](https://camel.abudhabi.nyu.edu/madar/) | Yes | Varying | MADAR project offers 2 datasets one is parallel that can be used for translation between dialects and another is not parallel that can be used to do dialect classification: the First data set (parallel one) is split into 2 sub-sets: 6-cities set :that have tuples of parallel sentences from 6 Arabic cities across the regions of the MENA the set contains 12k tuples 26-cities set: that have tuples of parallel sentences from 26 Arabic cities the set contains 2k tuples The second Dataset (None parallel) is also split into 2 subtasks the first task includes predicting the city tag of a sentence, this set is further split in 2 smaller tasks predicting the 6-cities tags, this set contains 54k examples predicting the 26-cities tags, this set contains 41.6K examples the second task is simpler and aims at predicting the whole country, the training set alone contains 217k tweets annotated by country |
| [LDC2012T09](https://catalog.ldc.upenn.edu/LDC2012T09) | Yes | 350M words | Parallel text for English, MSA, Levant and Egypt dialects, Note that this dataset is paid (2250$) on LDC |
Furthermore, if the goal is coarse classification, it is relatively easy to collect more data from social media or blogs by utilizing features like Geo-location, localized groups, …
### Approaches
Both the coarse and the fine-grained tasks can be tackled using basic text classification methods. For the **coarse** variant, the state of the art is around 82% in [1]. Honestly, with this much data, it is possible to train any model you want.
On the other-hand for the **find-grained** case, the task is a bit harder due to the confusion between the related dialects and the limited data size. The [MADAR](https://camel.abudhabi.nyu.edu/madar/) project system for identification which is called [ADIDA](https://adida.abudhabi.nyu.edu/#/) [2] can identify the dialect of over 26 cities around the Arabic world. The authors claim that they can identify the exact city of a speaker at an accuracy of 67.9% for sentences with an average length of 7 words, and reach more than 90% when the text is longer than 16 words.
### Conclusion
We believe that the task of dialect identification can be considered solved with several systems having relatively high performance and a wealth of data
## Cross-dialect translation
The
goal of this task is to translate text from one dialect to another or
to the MSA.
### Datasets
Some of the dialect Identification datasets also include bi-text that can be used to train machine translation systems, these datasets are [LDC2012T09](https://catalog.ldc.upenn.edu/LDC2012T09), [PADIC](https://sourceforge.net/projects/padic/), [Shami](https://github.com/GU-CLASP/shami-corpus) and [MADAR](https://camel.abudhabi.nyu.edu/madar/). See the following figure for an example of such parallel data from PADIC
However, all of the available datasets is extremely small to enable training of a machine translation system.
### Other available resources
Multilingual words embedding can be used to create a word-to-word translation system and can also serve as features in a larger translation system, [MUSE](https://github.com/facebookresearch/MUSE) by facebook AI is a tool to build these embeddings, and the people at MADAR have trained [such embeddings](https://camel.abudhabi.nyu.edu/arabic-multidialectal-embeddings/) for the case of Arabic dialects
### Ways to Collect Data
In [3] the authors suggest a method to generate synthetic data for machine translation of under-resourced languages using words-embeddings mapping, the algorithm is rather delicate and we refrain from describing it here. The authors test this method to generate data for Levantian-English translation pair, generating only 50K synthetic examples and adding them to the available 160k manual examples. And the NMT and SMT model was trained on both the manual data and the manual+synthetic data with the generated data increasing the performance of the baseline translator by around 2 BLEU points. Their best model scored 17.33 BLEU points. It is not clear how good is the generated data and if the 50k size was chosen in order not to influence the manual set.
### Approaches
The translation of dialects falls in the category of under-resourced machine translation since as we saw above. There is not a lot of data to train such models. There is a full field regarding this task and in this section we will try to explore the approaches that we believe can be applied to the dialectical Arabic translation. Please note that most of the following approaches can be applied to translating any under-resourced language.
#### Note about evaluation metrics
The main metric used in the automatic evaluation of translation systems is BLEU. If you are not familiar with the metric follow [this link](https://towardsdatascience.com/evaluating-text-output-in-nlp-bleu-at-your-own-risk-e8609665a213).
#### Rule-based
- In [4] an algorithm was proposed that normalizes Sanaani dialect to MSA based on morphological rules. Input text was tokenized and segmented. The stem and the affixes can be either dialect-specific, MSA-specific, or both. A rule-based system is built on top of the segmenter output.
- In [5], a rule-based approach for machine translation from Arabic dialects to MSA was presented. The approach relies on morphological analysis, morphological transfer rules and dictionaries, in addition to language models to produce MSA paraphrases of dialectal sentences. The treated dialects are Levantine, Egyptian, Iraqi, and Gulf Arabic
However,
the performance of these systems is questionable at best
#### Supervised SMT
Most of the research on dialectical translation and any automatic translation in Arabic have relied mostly on statistical machine translation (SMT) using tool kits like [Moses](http://www.statmt.org/moses/).
- In [6] the authors manually annotate a small dataset of approximately 150k pairs covering MSA, Levantine and Egyptian dialects. Their reported BLEU scores are 16.7 and 18.5 for the Levant and Egyptian dialects respectively.
- In [7] the first real cross-dialect translation experiments are done using SMT methodology, the reported results are not optimal see the table below. However, some pairs have a surprisingly high performance this can be partially due to the limited size of test data
| Source | ALG KN | ALG WB | ANB KN | ANB WB | TUN KN | TUN WB | SYR KN | SYR WB | PAL KN | PAL WB | MSA KN | MSA WB |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| ALG | — | — | 61.06 | 60.81 | 9.67 | 9.36 | 7.29 | 7.95 | 10.61 | 10.14 | 15.10 | 14.64 |
| ANB | 67.31 | 65.55 | — | — | 9.08 | 8.64 | 7.52 | 7.95 | 10.12 | 9.84 | 14.44 | 13.95 |
| TUN | 9.89 | 9.48 | 9.34 | 9.01 | — | — | 13.05 | 12.93 | 22.55 | 22.21 | 25.99 | 25.21 |
| SYR | 7.57 | 7.50 | 7.50 | 7.64 | 13.67 | 13.23 | — | — | 26.60 | 25.74 | 24.14 | 22.96 |
| PAL | 11.28 | 10.67 | 9.53 | 9.15 | 17.93 | 16.64 | 23.29 | 23.07 | — | — | 40.48 | 39.76 |
| MSA | 13.55 | 13.05 | 12.54 | 11.72 | 20.03 | 20.44 | 21.38 | 20.32 | 42.46 | 41.37 | — | — |
`KN` is Kneser–Ney smoothing and `WB` is Witten–Bell smoothing. Each cell is the BLEU score reported in Table 5 of [7].
However, the results of SMT in Dialectical Arabic is very low.
#### Supervised NMT
Neural Machine translation (NMT) dominates the current state of the art in machine translation in many languages. However the main obstacle against the widespread of this mode is its reliance on large quantities of data. The only result we found is reported in [8]. The authors used the transformer architecture and trained a model on the LDC2012T09 dataset mentioned above to translate from dialects to English. The authors tested whether using a pipelined model (given an input text the model first detects the dialect and then direct the text to a model tuned specifically for that dialect) is better than using a multi-lingual model (where a single model is used for all the inputs). The results were **22.79 and 23.78 BLEU** score for the Egyptian and Levant dialects respectively.
#### Unsupervised NMT
Unsupervised neural machine translation trains seq2seq models to translate from a source language to a target language without having any parallel dataset (sentences in the source language translated to the target language) and using only 2 monolingual independent texts.
There has been some recent work on full-scale unsupervised neural machine translation [9]–[11]. The best SOTA for translation between English and French [12] (using the WMT dataset) is 35 BLEU which is considerably less than the supervised SOTA of 46 on the same translation pair and test set.
**This same technique is language-agnostic** and can be applied to the dialects of Arabic.
#### Model adaptation in NMT
Since the major success in NMT was seen in western languages were large quantities of parallel data is available. Many people [13]–[15] considered the task of adapting the models trained on western languages to work on under-resourced languages.
The steps of the process of using the parameters of a parent model (trained on the large dataset) in initializing a child model (to be trained on a small dataset) usually go like this [14]:
1. Learn monolingual embedding of the child language in a monolingual fashion using word2vec
2. Extract source embedding from a pre-trained parent NMT model.
3. Learn a cross-lingual linear mapping between 1 and 2
4. replace the source embedding with the linearly mapped ones
See the following diagram for an illustration
This approach can increase the performance of under-resourced NMT models substantially. The following table is from [14]. **Although the absolute results may look low, these pairs are extremely under-resourced: Slovenian-to-English has 17k parallel sentence pairs, while the other child-language datasets have only 5k–10k pairs.**
| System | Basque–English | Slovenian–English | Belarusian–English | Azerbaijani–English | Turkish–English |
| --- | ---: | ---: | ---: | ---: | ---: |
| Transformer baseline (child only) | 1.7 | 10.1 | 3.2 | 3.1 | 0.8 |
| Multilingual (parent + child) | 5.1 | 16.7 | 4.2 | 4.5 | 8.7 |
| Transfer | 4.9 | 19.2 | 8.9 | 5.3 | 7.4 |
| + cross-lingual word embedding | 7.4 | 20.6 | 12.2 | 7.4 | 9.4 |
| + artificial noise | 8.2 | 21.3 | 12.8 | 8.1 | 10.1 |
| + synthetic parent data | 9.7 | 22.1 | 14.0 | 9.0 | 11.3 |
We have not come across any work tackling the issue of NMT model adaptation from Arabic MSA to Arabic dialects.
#### Cross-lingual embedding mapping
This method depends on the closeness between the dialects of Arabic and the MSA. **As mentioned above it is *possible* to build cross-lingual words embedding where the representations of words in one language are very close to the representation of the translation of that word in another language, see the figure below from Facebook’s MUSE.**
In [16] the author describes a similar approach to carry out word-by-word translation of dialects. They basically train 2 separate embedding models for the source language (Egy - Epyptian) and the target language (MSA - Arabic Fusha) and then they use a bi-text dictionary to train a linear mapping that takes the representation of the word in the source language and return its representation in the target language, their model can predict possible translations for each word see the figure below.
This approach stops at this point. However, it is possible to extend this model to translate full sentences by treating the suggested words as a lattice and disambiguate it in a similar manner to speech recognition using a language model. Tools like [SRILM](http://www.speech.sri.com/projects/srilm/) is usually used to solve this task. See the following graph for an example lattice from speech recognition
While this method is relatively simple and language agnostic, such a simplistic approach have several issues including: named entities handling, out of vocabulary words, and the fact that although such a system considers the context in the target language it does not consider it in the source language.
#### Other Approaches
- In [17] the authors describe a dialectical translation system using a hybrid rule based and SMT approach. The paper does not describe any evaluation results on test data but **the system is available for researchers upon contact**. This system was used in [18] to improve the translation of Arabic to english. However the improvement to plain MSA baseline is negligible.
- In [19] the authors follow a similar pattern of 1- translating Egyptian dialectical speech first to MSA and 2- then translating the MSA results to english, the results are not impressive with **16.8 BLEU score**.
### **Conclusion**
- First I would take these conclusions with a grain of salt because we need more research before taking any decision regarding the best path we should follow. The best source to start with is [20].
- Supervised SMT is rather worrisome since we have so little data
- Rule-based and ad-hoc approaches have been used previously. I am not sure of their potential but we should try to acquire the ELISSA system just to measure the performance of such system.
- We need to do more research on the venues of cross-lingual words embeddings as well as a way to collect more data.
- **What we should start with is:**
- Supervised NMT is a possible solution; if approaches like domain adaption and synthetic data generation were followed.
- The unsupervised NMT is worth a shot since codebases are available, and it can serve as an initialization for a supervised NMT.
## Arabizi Transliteration and Vowelization
The Arabizi reefers to text written in Arabic language but using roman script e.g. “Sho Hal7ki”. Such type of text is used a lot in the social media. The text varies due to the variations of the Arabic dialects. The goal here is to convert the aforementioned text to the Arabic script.
This task is important for other applications related to dialectical Arabic (basically anything related to social media analysis)
The
task can be classified into several levels:
- **Phonological**: by basically doing a char-to-char translation in which [a simple letter based mapping](https://en.wikipedia.org/wiki/Arabic_chat_alphabet) between Arabizi symbols and Arabic letters. Such an approach would be suitable in creating a very simple editor.
- **Lexical**: this a bit higher level in which the translation is done on a word-to-word basis. This task can be declared solved mainly because there is a lot of open software that provides this service like Google’s [Taarib](https://www.google.com/intl/ar/inputtools/try/), [Yamli](https://www.yamli.com/editor/ar/).
- **Syntactic**: basically doing a fully-fledged machine translation between Arabizi and Arabic which is harder since the context has a role to play in this variation.
### Datasets
There is a decent corpora size for this task the following table list the ones we found
| Name | Free or paid | Size in words | Notes |
| --- | --- | --- | --- |
| [ILPS](https://ilps.science.uva.nl/resources/arabizi/) | Free | 10K | From [21] but the details of data collection in their paper are rather vague and the quality of the dataset needs to be checked |
| [ELRA-W0126](http://catalog.elra.info/en-us/repository/browse/ELRA-W0126/) | Paid | | Contains 3,452 Arabizi tokens manually transliterated into Arabic, and a set of 127 Arabizi tweets containing 1,385 words also manually transliterated into Arabic. And while the commercial license costs 650 euros there is a free license for research purposes. |
| [Camel](https://camel.abudhabi.nyu.edu/arabscribe/) | Free | 10k | Arabic words transcribed using roman code (Romanization) and in local Arabizi |
| [BOLT](https://catalog.ldc.upenn.edu/LDC2017T07) | Paid | ? | Might be outdated |
| QCRI | Free | ? | I remember they had a dataset for this task I just can’t find it anywhere |
### Methods to collect more data
In One paper by the QCRI (I really couldn’t find it) the authors generated a pronunciation table using first names from students records in several schools. These records included the names in Arabic and a romanized version of it. The goal was to improve the transliteration performance on OOV words.
### Approaches
- The authors in [21] collect the ILPS dataset and then build a simple translation system using a method similar to the lattice disambiguation approach from the machine translation section. **The software is open.**
- In [22] the QCRI ppl achieve an F1 measure of 0.93 on the task of normalizing people names mentions in multiple documents.
- [23] tackle both the task of code-switching detection (if the writer is writing in English or Arabizi) as well as the transliteration. They report accuracy of 93% on the code-switching task and an 88.7% conversion accuracy, with roughly a third of errors being spelling and morphological variants of the forms in ground truth. Their method used the same idea of lattice generation and disambiguation using a language model.
### Conclusion
- We believe that this task can be successfully implemented at least on the phonological and morphological lexical levels by for example training a character level decoder since there is some data for this task.
- For full syntactical transliteration, we should try at least the code from [ILPS](https://ilps.science.uva.nl/resources/arabizi/) as an off the shelf component.
## Sentiment analysis
In this task given a piece of text, the system should extract the overall polarity in it (positive, negative, neutral).
This is usually done in 2 steps. First, the subjectivity of the text is determined aka (subjective vs objective) and if it is subjective the second step determines the polarity.
There is a wealth of datasets, lexicons, and pre-trained embeddings and quite a lot of work in the literature. The best line of work known to me is by [Sief Mohamad](http://saifmohammad.com/WebPages/ResearchInterests.html) from the CNRC. The dataset is all in the [corpora list document](https://docs.google.com/spreadsheets/d/1HnUJH7N-43AOTdishzHFtNqFta0jgXils_LWTTPqu0M/edit#gid=0).
However, in total, the performance of these models is relatively low (when the results are reported on shared challenges and not on homemade corpora ) and highly depends on the domain specificity as over-fitting is a real threat [22]. This is due to issues like the excessive use of sarcasm and contextual language, code-switching, and the dependence of the system on the dialect or the genre. The best approach to tackle this task is building a universal model and then adapting that model using data from a specific domain, genre, and dialect to ensure the highest performance. However, in a real industrial application, the size and type of data needed for adaptation become crucial and further investigation is needed.
There are some already established social media Analysis services for Arabic including [crowdAnalyzer](https://www.crowdanalyzer.com/) and [trend25](https://25trends.me/).
## Summarization
This task is also listed under dialectical language processing mainly because of applications like social media summarizations and customers reviews summarization which would deal closely with dialectical content. Such applications can be helpful in tasks like social media analysis or stuff like google alert. In these applications the customer would like to see a gist of the user's reviews in, say the last month, related to a specific brand or product and hows sentiment is negative.
### Datasets
We have not come across any summarization datasets for dialectical Arabic. Furthermore, none of the available summarization datasets in Arabic includes dialectical content (they are taken from Wikipedia, news outlets, …)
### Approaches
In [24], a microblog summarization technique based on machine learning for Egyptian dialect was presented. The results achieved were compared to several well-known algorithms such as SumBasic, TF-IDF, TextRank, and human summaries.
### Conclusion
- Simple language-agnostic summarization methods like TextRank can work easily on dialectical content.
- There is no work on abstractive summarization in dialectical Arabic due to the lack of dataset (there is a lack of data in MSA itself let alone dialects), this task poses a challenge.
- Most of the aforementioned applications of dialectical summarization focus on the multi-document variation (finding the gist of multiple users reviews/ tweets)
## Information Retrieval
This addresses the issue of building search services that can deal with dialectical language. The main complexity in applying usual IR methods comes from the variations in orthography that is found in dialectical content due to the usage of Arabizi for example. This issue can be resolved by either transliteration or dialect specific normalization.
### Approaches
- In [25] the author addressed the issue of linguistic differences in IR. The presented tool automatically generates dialect search terms with relevant morphological variations from English or Standard Arabic query terms.
- [26] is a very good summary on the IR in Arabic (couldn’t go through it cause it is too long.)
## Preprocessing Tasks
In this section we will explore the various tasks related to text pre-processing:
- Normalization: different normalization schemes need to be employed to handle dialectical text mainly because, in contrast to MSA, dialectical Arabic has no orthographic standard. The same word can be written in different forms. This poses difficulties in NLP tools. This task is extremely related to the task of Arabizi transliteration.
- Segmentation: splitting the words into their constituents e.g. playing → play+ing
- Part of Speech (POS) tagging: basically identifying the pos tag for each of the words in the sentence e.g. “boy plays the violin” → [(boy, Noun), (plays, Verb), (the, Identifier), (violin, Noun)]
- Named entity recognition (NER): recognizing named entities like persons names, organizations or places.
### Multi-task Datasets
- The people at QCRI have built [a sizeable dataset](http://alt.qcri.org/resources/da_resources/) for dialectical Arabic segmentation and POS tagging for dialectical content.
- [Curras dataset](https://portal.sina.birzeit.edu/curras/download.html) is a similar dataset that contains the lemma and fine-grain pos for text in the Palestinian dialect.
### Normalization
- In [27] the first steps towards normalizing Arabic dialects orthography for Levantine and Egyptian were made. For that, different similarity measures were employed that exploit string similarity and contextual semantic similarity, to unify different writings of the same word.
- Some people [28] suggested creating a specialized orthography (way of writing) that is comparable between dialectical and MSA. However, this never gained traction.
### Named Entity Recognition (NER)
In [29], the authors address the issue of named entity recognition in microblogs, the describe the various complexities of NER in tweets. **They utilize a weekly supervised language-agnostic method and report a rather low result of 0.65 f1 scores on dialectical tweets.**
Following approaches, we didn’t have time to investigate are [30], [31].
### Morphological Analysis
- [32] rule-based Morphological analyzer for dialectical Arabic, didn’t gain tract and very old. Not sure of the importance.
- In [33], two morphological analyzers for Gulf, Levantine, Egyptian, North African, Sudani, and Iraqi dialects were presented. The first one relies on an MSA morphological analyzer. The second one applies word segmentation and uses web data as a corpus to produce statistical information about the frequency of different segment combinations.
- [34] morphological analyzer for Egyptian dialect.
- [35] training a supervised pos tagger on a manually annotated set of Egyptian dialect text.
- [36] a full morphological analyzer for Egyptian that supports part-of-speech tagging, diacritization, lemmatization, and tokenization.
## Speech Recognition
In speech recognition, the system takes as input the speech waveform and is tasked with transcribing the content into words.
### Datasets
- A series of challenges by DARPA created a set of paid datasets most famous of which is the [call home](https://catalog.ldc.upenn.edu/LDC97S45) data set that contained Egyptian Arabic speech.
- However, the largest free dataset for MSA 1200 hours was provided by the QCRI through their [MGB-2](http://www.mgb-challenge.org/arabic_download.html) challenge that focused on speech recognition in the wild (i.e. under various variations in the channel, noise, and speakers conditions) by using transcribed TV broadcasts from Aljazeera.
- [MGB-3](http://www.mgb-challenge.org/MGB-3.html) introduced a small dataset of 16 hours of Egyptian dialect youtube videos for adaptation of systems developed for MSA using the MGB-2 to the Egyptian dialect.
- Similarly, MGB-5 added another small Moroccan dialectical dataset of 18 hours also for adaptation of MGB-2 systems. All the MGB datasets are available on contact.
- [This collection](https://github.com/qcri/dialectID/tree/master/data) of datasets by the QCRI includes several useful data for speech recognition.
### Methods to collect more data
In some cases speech and transcripts exist independently. For example, all the articles of the *[Syrian Researchers](http://syrian-researchers.com)* are read aloud by some of the team members and saved on SoundCloud. In such cases it is possible to use forced aligners like [aeanes](https://github.com/readbeyond/aeneas/) to align the articles with the spoken version, and thus create large dataset for ASR and TTS. The only pain in this process is the normalization of text to align with speech. However, I don’t know any resources like this in dialects. Other examples include the audio version of the bible and the Quran (which are *not* dialectical.)
### Approaches
- The dissertation of Ahmad Ali [37] is by far the most important work in the field currently and it sums the effort of 5 years collaboration between Edinburgh University, QCRI, and Aljazeera on Arabic speech. They even provide a recipe for an ASR system in Arabic using Kaldi.
- The major approach many of the applicants have used in MGB-3 challenge is based on training an ASR on MSA using te MGB-2 data and then adapting it to the dialect[38]. The best result on the Egyptian dialect is 29% WER (word error rate). Please note that the results of MGB-5 are not public yet.
- Another approach is proposed for English by Google AI team [39], [40]. In these, they are investigating the applicability of multi-dialect speech recognition using a single end-to-end model in a similar fashion to google machine translation system. This new method has been applied successfully to multi-dialect English ASR and the results seems to be promising. Especially in performing zero-shot ASR (recognizing speech from dialects on which the system was never trained). There have even been some attempts at developing multi-lingual ASR [41].
### Conclusion
- The task of speech recognition is one of the most important tasks in dialectical Arabic mainly because dialects are mostly expressed through speech and to a lower degree through text (social media and texting.)
- While it is possible to build ASR for MSA, the dialectical ASR is limited to dialects that have data for adaptation. MGB challenge provide data for Egyptian and Moroccan dialects only. We need to do more research to explore if more dialectical data is available or if there are more ways to collect data.
## Speech Dialect Identification
The Arabic dialect Identification task is a special case of the more general language Identification, it is a vital component in most of the multi-dialect speech recognition systems, and is usually used as well in some speaker recognition systems.
In very simple terms the input of such a model is a speech wave and the output is the dialect of the speech. However, in contrast to the language detection, dialect detection is harder due to the similarity between the dialects.
### Datasets
- [MGB-3](http://www.mgb-challenge.org/arabic_download.html) data set has been utilized for dialectical speech recognition (by adapting systems developed for MSA using the larger MGB-2 data to the Egyptian dialect) as well as dialect identification. This dataset is rather small it contains only 16 hours including adaptation, development and evaluation set, the dataset only includes Egyptian YouTube programs and is centred on podcasts.
- [MGB-5](http://www.mgb-challenge.org/MGB-5.html) ADI data: this dataset contains 3000 hours of YouTube videos covering 17 different dialects and many genres.
- VarDial conference 2018 and 2017 have a harder shared task on Arabic speech dialect identification.
- Finally, [this](http://alt.qcri.org/resources/aljazeeraSpeechCorpus/) is also a pretty large dataset from Aljazeera programs of over 1200 hours of speech annotated for the 5 main dialects.
### Approaches
- [42]–[45] have utilized the MGB-3 dataset to train their system with results reaching up to 80% f1 and [here](https://dialectid.qcri.org/) is a live demo the best model by QCRI and MIT utilized I-vectors in the detection, a method commonly used in user identification.
- On VarDial the results are much worse mainly because the data is much smaller with much lower results of around 60% see [46]–[48].
### Conclusion
We believe that also this task can be implemented since there is a wealth of data and a sizable literature on the field
## Text to Speech
While TTS is available in MSA I have come across no work related to TTS in the dialects, and I hardly see any point in developing such a system. Mo thinks that maybe I'm wrong.
##
## References
[1] M.
Elaraby and M. Abdul-Mageed, “Deep Models for Arabic Dialect
Identification on Benchmarked Data,” in *Proceedings
of the Fifth Workshop on NLP for Similar Languages, Varieties and
Dialects (VarDial 2018)*,
2018, pp. 263–274.
[2] M.
Salameh and H. Bouamor, “Fine-grained arabic dialect
identification,” in *Proceedings
of the 27th International Conference on Computational Linguistics*,
2018, pp. 1332–1344.
[3] H.
Hassan, M. Elaraby, and A. Tawfik, “Synthetic data for neural
machine translation of spoken-dialects,” *ArXiv
Prepr. ArXiv170700079*,
2017.
[4] G.
H. Al-Gaphari and M. Al-Yadoumi, “A method to convert Sana’ani
accent to Modern Standard Arabic,” *Int.
J. Inf. Sci. Manag.*,
vol. 8, no. 1, 2010.
[5] W.
Salloum and N. Habash, “Dialectal to standard Arabic paraphrasing
to improve Arabic-English statistical machine translation,” in
*Proceedings
of the first workshop on algorithms and resources for modelling of
dialects and language varieties*,
2011, pp. 10–21.
[6] R.
Zbib *et
al.*,
“Machine translation of Arabic dialects,” in *Proceedings
of the 2012 conference of the north american chapter of the
association for computational linguistics: Human language
technologies*,
2012, pp. 49–59.
[7] K.
Meftouh, S. Harrat, S. Jamoussi, M. Abbas, and K. Smaili, “Machine
translation experiments on padic: A parallel arabic dialect corpus,”
in *The
29th Pacific Asia conference on language, information and
computation*,
2015.
[8] P.
Shapiro and K. Duh, “Comparing Pipelined and Integrated Approaches
to Dialectal Arabic Neural Machine Translation,” in *Proceedings
of the Sixth Workshop on NLP for Similar Languages, Varieties and
Dialects*,
2019, pp. 214–222.
[9] M.
Artetxe, G. Labaka, E. Agirre, and K. Cho, “Unsupervised neural
machine translation,” *ArXiv
Prepr. ArXiv171011041*,
2017.
[10] G.
Lample, L. Denoyer, and M. Ranzato, “Unsupervised Machine
Translation Using Monolingual Corpora Only,” *ArXiv
Prepr. ArXiv171100043*,
2017.
[11] G.
Lample, M. Ott, A. Conneau, L. Denoyer, and M. Ranzato, “Phrase-Based
& Neural Unsupervised Machine Translation,” *ArXiv
Prepr. ArXiv180407755*,
2018.
[12] M.
Artetxe, G. Labaka, and E. Agirre, “An effective approach to
unsupervised machine translation,” *ArXiv
Prepr. ArXiv190201313*,
2019.
[13] M.
Gheini and J. May, “A Universal Parent Model for Low-Resource
Neural Machine Translation Transfer,” *ArXiv
Prepr. ArXiv190906516*,
2019.
[14] Y.
Kim, Y. Gao, and H. Ney, “Effective Cross-lingual Transfer of
Neural Machine Translation Models without Shared Vocabularies,”
*ArXiv
Prepr. ArXiv190505475*,
2019.
[15] T.
Kocmi and O. Bojar, “Trivial transfer learning for low-resource
neural machine translation,” *ArXiv
Prepr. ArXiv180900357*,
2018.
[16] E.
H. Almansor, “Translating Arabic as low resource language using
distribution representation and neural machine translation models,”
2018.
[17] W.
Salloum and N. Habash, “Elissa: A dialectal to standard Arabic
machine translation system,” in *Proceedings
of COLING 2012: Demonstration Papers*,
2012, pp. 385–392.
[18] W.
Salloum and N. Habash, “Dialectal arabic to english machine
translation: Pivoting through modern standard arabic,” in
*Proceedings
of the 2013 Conference of the North American Chapter of the
Association for Computational Linguistics: Human Language
Technologies*,
2013, pp. 348–358.
[19] H.
Sajjad, K. Darwish, and Y. Belinkov, “Translating dialectal arabic
to english,” in *Proceedings
of the 51st Annual Meeting of the Association for Computational
Linguistics (Volume 2: Short Papers)*,
2013, pp. 1–6.
[20] S.
Harrat, K. Meftouh, and K. Smaili, “Machine translation for Arabic
dialects (survey),” *Inf.
Process. Manag.*,
2017.
[21] M.
van der Wees, A. Bisazza, and C. Monz, “A simple but effective
approach to improve arabizi-to-english statistical machine
translation,” in *Proceedings
of the 2nd Workshop on Noisy User-generated Text (WNUT)*,
2016, pp. 43–50.
[22] W.
Magdy, K. Darwish, O. Emam, and H. Hassan, “Arabic cross-document
person name normalization,” in *Proceedings
of the 2007 Workshop on Computational Approaches to Semitic
Languages: Common Issues and Resources*,
2007, pp. 25–32.
[23] K.
Darwish, “Arabizi detection and conversion to Arabic,” *ArXiv
Prepr. ArXiv13066755*,
2013.
[24] N.
El-Fishawy, A. Hamouda, G. M. Attiya, and M. Atef, “Arabic
summarization in twitter social network,” *Ain
Shams Eng. J.*,
vol. 5, no. 2, pp. 411–420, 2014.
[25] A.
Pasha *et
al.*,
“Dira: Dialectal arabic information retrieval assistant,” in *The
Companion Volume of the Proceedings of IJCNLP 2013: System
Demonstrations*,
2013, pp. 13–16.
[26] K.
Darwish and W. Magdy, “Arabic information retrieval,” *Found.
Trends® Inf. Retr.*,
vol. 7, no. 4, pp. 239–342, 2014.
[27] P.
Dasigi and M. Diab, “Codact: Towards identifying orthographic
variants in dialectal arabic,” 2011.
[28] N.
Habash, M. T. Diab, and O. Rambow, “Conventional Orthography for
Dialectal Arabic.,” in *LREC*,
2012, pp. 711–718.
[29] K.
Darwish and W. Gao, “Simple Effective Microblog Named Entity
Recognition: Arabic as an Example.,” in *LREC*,
2014, pp. 2513–2517.
[30] A.
Zirikly and M. Diab, “Named entity recognition for dialectal
arabic,” *ANLP
2014*,
p. 78, 2014.
[31] A.
Zirikly and M. Diab, “Named entity recognition for arabic social
media,” in *Proceedings
of the 1st Workshop on Vector Space Modeling for Natural Language
Processing*,
2015, pp. 176–185.
[32] N.
Habash and O. Rambow, “Morphophonemic and orthographic rules in a
multi-dialectal morphological analyzer and generator for arabic
verbs,” in *International
symposium on computer and arabic language (iscal), riyadh, saudi
arabia*,
2007, vol. 2006.
[33] K.
Almeman and M. Lee, “Towards developing a multi-dialect
morphological analyser for arabic,” in *4th
international conference on arabic language processing, rabat,
morocco*,
2012.
[34] W.
Salloum and N. Habash, “ADAM: Analyzer for dialectal Arabic
morphology,” *J.
King Saud Univ.-Comput. Inf. Sci.*,
vol. 26, no. 4, pp. 372–378, 2014.
[35] R.
Al-Sabbagh and R. Girju, “A supervised POS tagger for written
Arabic social networking corpora.,” in *KONVENS*,
2012, pp. 39–52.
[36] N.
Habash, R. Roth, O. Rambow, R. Eskander, and N. Tomeh, “Morphological
analysis and disambiguation for dialectal Arabic,” in *Proceedings
of the 2013 Conference of the North American Chapter of the
Association for Computational Linguistics: Human Language
Technologies*,
2013, pp. 426–432.
[37] A.
M. A. M. Ali, “Multi-dialect Arabic broadcast speech recognition,”
2018.
[38] S.
Khurana, A. Ali, and J. Glass, “DARTS: Dialectal Arabic
Transcription System,” *ArXiv
Prepr. ArXiv190912163*,
2019.
[39] M.
Johnson *et
al.*,
“Google’s multilingual neural machine translation system:
Enabling zero-shot translation,” *Trans.
Assoc. Comput. Linguist.*,
vol. 5, pp. 339–351, 2017.
[40] B.
Li *et
al.*,
“Multi-dialect speech recognition with a single
sequence-to-sequence model,” in *2018
IEEE International Conference on Acoustics, Speech and Signal
Processing (ICASSP)*,
2018, pp. 4749–4753.
[41] S.
Toshniwal *et
al.*,
“Multilingual speech recognition with a single end-to-end model,”
in *2018
IEEE International Conference on Acoustics, Speech and Signal
Processing (ICASSP)*,
2018, pp. 4904–4908.
[42] C.
Zhang, Q. Zhang, and J. H. Hansen, “Semi-supervised Learning with
Generative Adversarial Networks for Arabic Dialect Identification,”
in *ICASSP
2019-2019 IEEE International Conference on Acoustics, Speech and
Signal Processing (ICASSP)*,
2019, pp. 5986–5990.
[43] S.
Khurana, M. Najafian, A. M. Ali, T. Al Hanai, Y. Belinkov, and J. R.
Glass, “QMDIS: QCRI-MIT Advanced Dialect Identification System.,”
in *Interspeech*,
2017, pp. 2591–2595.
[44] A.
E. Bulut, Q. Zhang, C. Zhang, F. Bahmaninezhad, and J. H. Hansen,
“UTD-CRSS submission for MGB-3 Arabic dialect identification:
Front-end and back-end advancements on broadcast speech,” in *2017
IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)*,
2017, pp. 360–367.
[45] S.
Shon, A. Ali, and J. Glass, “Convolutional neural networks and
language embeddings for end-to-end dialect recognition,” *ArXiv
Prepr. ArXiv180304567*,
2018.
[46] A.
M. Butnaru and R. T. Ionescu, “UnibucKernel Reloaded: First place
in Arabic dialect identification for the second year in a row,”
*ArXiv
Prepr. ArXiv180504876*,
2018.
[47] M.
Zampieri *et
al.*,
“Language Identification and Morphosyntactic Tagging: The Second
VarDial Evaluation Campaign,” in *Proceedings
of the Fifth Workshop on NLP for Similar Languages, Varieties and
Dialects*,
2018.
[48] P.
Nakov, M. Zampieri, N. Ljubešić, J. Tiedemann, S. Malmasi, and A.
Ali, “Proceedings of the Fourth Workshop on NLP for Similar
Languages, Varieties and Dialects (VarDial),” in *Proceedings
of the Fourth Workshop on NLP for Similar Languages, Varieties and
Dialects (VarDial)*,
2017.
### Multi-Document Summarization. The What, Why and How
- URL: https://mohammadshaker.com/en/blog/multi-document-summarization-the-what-why-and-how
- Date: 2020-01-26T00:00:00.000Z
- Tags: almeta.io, arabic, ml, nlp, research, Human-Written
Multi-document summarization merges information from several articles covering the same event into one coherent summary. The main challenges are redundancy elimination, cross-document coreference, and information ordering. Extractive methods copy sentences directly; abstractive methods generate new text. Both need to handle temporal inconsistencies across articles.
#### Content
While reading the news you are most likely to encounter several articles that describe the same event or incident and each of these articles comes from a different news anchor and provides a different viewpoint to the event. However most of us do not have the time to read all of these articles in order to get a fully unbiased view of the current events, and we usually select the anchor that best suites our ideology or agrees with our preconceived opinions, leading to a biased view of the events.
Here at Almeta, we believe in the importance of discussing different viewpoints to any argument. Imagine having a service that can identify the various articles the describe the same event and then combine them into a short condensed and informative summary that covers all details given by the different anchors.
We have already described in a different series the details of our event detection system. However, in this article, we look closer at the second part namely combining various documents in one informative summary.
# What is Multi-document Summarization?
Summarization methods can be categorized into single or multi-document approaches. In contrast to the single document, the multi-document setup can utilize the fact that in some domains like news articles there are different sources describing the same event and thus these articles have a lot of similarities. In such systems, the input would be two or more pieces of text and the output should be a summary that covers the details of all the inputs.
Furthermore, the summarization approaches can be classified into two other classes based on the way the summary is constructed from the input text. Extractive methods solely rely on the words of the input and e.g. extract whole sentences from it and does not produce any new sentences, while abstractive approaches, on the other hand, the output of the system is rarely bound to any constraints and these systems can generate any text that constitutes a valid summary of the input, these systems have gained a lot of traction recently due to current advances in Deep Learning.
# Data Sources
In order to build any machine learning system, the main concern is having a large-enough and diverse enough dataset for this model to learn from. In the case of multi-document summarization, each data point is composed of a set of two or more articles alongside a summarization of them.
## Available Datasets
[In our previous article](/blog/can-you-measure-a-text-informativeness-using-its-summary), we explored the available datasets for summarization in both Arabic and English languages. All of these datasets are built for single-document summarization with the exception of DUC2004, this dataset includes only 100 samples and is designed mainly for evaluation, other datasets are introduced through the [MultiLing](http://multiling.iit.demokritos.gr/pages/view/1571/datasets) conference but they as well suffer from the issue of data sparsity.
On the contrary, there are several English multi-document summarization datasets including DUC(2001-2007), TAC(2008-2011) and Multiling Nonetheless even in the case of English all of these datasets are extremely small with the largest being the DUC2004 dataset. Last year, the researchers at google brain produced the first large scale Multi-document summarization dataset for English by basically using Wikipedia [1] and then on Wiki-news [2] (more on this in the next section). However, the size of the publicly available data is restricted by the authors to around 10k from over a million datapoint collected in the original dataset. And earlier this year some researchers at Yale [3] created another large scale dataset mostly by relying on the site [newser.com](http://newser.com/) this 50K dataset is available for [download from GitHub.](https://github.com/Alex-Fabbri/Multi-News)
## How Can We Collect More Data
Although there is a limited number of publicly available datasets to train a multi-document summarizer, there are several methods reported in the literature to build a dataset for multi-document summarization in an automatic or semi-automatic manner. While most of these methods are applied to English, we focus on this part on methods that can be extended into relatively under-resourced languages especially Arabic.
### Summarizing Wikipedia
Several people have used Wikipedia in one way or another to generate data for summarization. In [5], the Google Brain team uses the whole English Wikipedia as a multi-document summarization data set. This application is extremely complicated.
However, we believe that the following approach is applicable by a relatively small team with limited resources. The authors of [6] **used Wikinews to generate a large multi-document summarization dataset. The idea is that Wikinews articles can be considered as the summaries, while the sources of the article can represent the sources of the articles.**
**Pros**
- While treating all of Wikipedia as a summarization dataset is hard in under-resourced languages, mainly because many of the sources are usually written in English or other western languages, in the case of Wikinews most of the sources come for news outlets speaking the native language.
- The same approach can be applied to many languages.
**Cons**
- The complexity of cleaning data from multiple outlets would usually need language detection modulos.
- The size of data is rather limited to 52k articles.
## Approaches
In this section, we will highlight the main methods to implement a multi-document summarizer
### Extractive Methods
Most of these algorithms are unsupervised and thus does not require any training data, in these approaches the sentences of all the document are extracted and scored using certain metrics, then the highest-scoring sentences are selected. The difference between these algorithms mainly lies in the difference between scoring methods.
#### Graph-based
One of the oldest of these algorithms is LexRank. LexRank is very similar to other graph-based summarizers like TextRank and DivRank, in that they “recommend” other similar sentences to the reader. **Thus, if one sentence is very similar to many others, it will likely be a sentence of great importance. The importance of this sentence also stems from the importance of the sentences “recommending” it.** Thus, to get ranked highly and placed in a summary, a sentence must be similar to many sentences that are in turn also similar to many other sentences. **This makes intuitive sense and allows the algorithms to be applied to any arbitrary new text.**
To implement this, a graph of all the sentences in all the documents is generated where the distances between these sentences represent the lexical/semantic similarity and then the LexRank algorithm is used to rank the sentences.
For a very good overview of the differences between DivRank, TextRank, and LexRank we recommend [this awesome review](https://blog.peiyingchi.com/2019/10/29/TextRank-LexRank-DivRank/).
Finally, there were also a few attempts at using deep learning for performing sentence ranking for MDS. The system in [4] employs a Graph Convolutional Network (GCN) on the sentence relation graphs, with sentence embeddings obtained from Recurrent Neural Networks as input node features. Through multiple layer-wise propagations, the GCN generates high-level hidden sentence features for saliency estimation. A greedy heuristic is then applied to extract salient sentences while avoiding redundancy.
#### Reconstruction Based
Sparse reconstruction methods [5]–[7] employ the idea of data reconstruction in the summarization task. In this approach, the task of summarization is treated as a data reconstruction problem. Intuitively, a good summary should recover the whole documents, or in other words, reconstruct the whole documents. And thus the summary should meet three key requirements: Coverage, Sparsity, and Diversity
- Coverage means the extracted summary can conclude every aspect of all documents. This is estimated by trying to reconstruct every document as a weighted average of the summary sentences.
- Sparsity meaning that the summary sentences should describe different topics.
- Diversity: meaning that the summary sentences should not be correlated and be as distinctive as possible.
The overall structure is shown in the following graph:
To represent this in a mathematically feasible manner, the main
outline of these approaches goes as follows:
- represent every sentence as a vector e.g. TF-IDF or Doc2Vec
- assuming a summary set of sentences S define the following metrics:
**Coverage metric**: how many articles representations (vectors) can be calculated as a non-negative weighted average of the summary set sentence vectors, this is basically the MSE error as follows, where Si is a sentence from the original documents, s∗j is a sentence from summary and k is the number of sentences in the summary and aij is the weight of the jth summary sentence in constructing the ith original sentence
**Sparsity metric**: the previous step would result in a matrix of weights aij, we want every original sentence to be represented by a small number of the summary sentence, and thus sparsity can be achieved by imposing a sparsity constraint on the columns of the matrix, for example, using L1 norm
**Diversity metric:** the final metric tries to minimize the correlation between the sentence in the summary, the correlation between two sentences Si and Sj in the summary can be calculated as follows, where the Sbar is the average vector of the summary sentences:
based on these steps the overall error function looks like this:
The last step is to find the summary set of sentences that minimizes this error, using a stochastic search algorithm like simulated annealing or genetic algorithms.
The different implementations of this approach differ in the choice of:
- Sentence representations TF-IDF, Doc2Vec, neural representation using auto-encoders, …
- Metrics definition
- Optimization algorithms
#### Topic-based
This is not to be confused with query-based summarization (also sometimes called topic-based because linguists are great at naming. In query-based methods, the system accepts a set of documents and a query and should construct a summary most relevant to that query.)
The algorithms used in these approaches are similar. The general structure of these algorithms goes like this:
1. Find the topics/events described in the documents. The number of topics to be considered is usually based on the desired summary length. The topic modeling is usually carried out using LDA or similar algorithms in an unsupervised manner.
2. Rank each of the sentences based on their similarity to the topics, the usual way of doing this is by constructing a query from each topic by concatenating prominent words, and then each sentence in the document is ranked using IR metrics like IDF or TF/IDF, although using other similarity metrics based on manifolds or documents embeddings was also suggested.
3. After selecting a specific number of sentences for each of the topics, the final summary is usually constructed using the Maximum Marginal Relevance (MMR) algorithm [8]. MMR relies on a greedy approach to select sentences basically by choosing sentences that are similar enough to the topic (query) and in the same time dissimilar enough from the already selected sentences. The goal is to select the sentences with the highest relevant to the topics and the most novel at the same time.
**Pros**
- Relatively simple to understand and implement, and have several ready-made implementations.
- Is language-agnostic although language-dependent text pre-processing can improve the results.
- Unsupervised and thus require no training data.
- Can give relatively decent results in practice based on the domain.
- Relatively very fast (with the exception of reconstruction based method).
**Cons**
- These algorithms work on the sentence level and thus have no information about the sentence context, this can lead to discontinuities in the resulting summary that will look more like a bullet point list than as a coherent paragraph.
- The possibility of eliminating some documents that are too short or there are too many documents.
- No abstraction or entailment of ideas.
### Intermediate Methods
These methods serve as a natural extension and usually as a second processing step to an extractive summarizer. In this step the system applies further processing on the selected sentences.
#### **Sentence Compression**
Compression is achieved by removing unimportant or redundant words/phrases from the selected sentences, making them more concise and general. this task is also called sentence compression, and you can learn more about it through [our previous post on automatic paraphrasing in NLP.](/blog/automatic-sentence-paraphrasing)
#### **Summary Revision**
This step is also deployed as a post-processing step after an extractive summarizer. Basically, it tries to improve the quality of the summary by rewriting noun phrases and resolving co-references, see for example [9]. Furthermore [our previous discussion of text paraphrasing](/blog/automatic-sentence-paraphrasing) can also be of help here.
**Pros**
- There are several possible ways to implement this task in a supervised or unsupervised manner and in a language-dependent or language-agnostic way.
- Relatively simple to understand and implement, and have several ready made implementations.
**Cons**
- These algorithms also work on the sentence level and have no information about the sentence context.
- Possibility of propagation error since these systems are usually implemented as a second step to extractive summarizers. Also abstraction or entailment of ideas.
- The relative improvement over the simple extractive methods is limited when measured using automatic metrics like ROUGE however the authors claim improvement in human evaluations.
### Abstractive Methods
Unlike the previous methods, abstraction-based methods can
generate new sentences whose fragments come from different source
sentences, using sentence aggregation and fusion. Here are some of
the main approaches to the task.
#### **Sentence Fusion**
Several publications have addressed this issue, as mentioned above the goal here is to find similar sentences in a piece of text and then merge them to create a shorter version of the text. This approach is used in [Quill-bot](https://quillbot.com/) check it out to get an idea of the performance of these systems. The main structure of the algorithm is shown in the following figure
The main structure
[10], [11] follows the following approach:
1. First, split the documents into sentences and cluster them, one approach followed by [11] identify the most important document in the multi-document set. Then the sentences in the most important document are aligned to sentences in other documents to generate clusters of similar sentences.
2. Second, generate K-shortest paths from the sentences in each cluster using a word-graph structure to generate the summary sentences. See the following figure for illustration
3. Finally, Rank these sentences based on informativeness, … and combine them to generate the summary. For example, integer linear programming (ILP) is used to select the summary sentences.
**Pros**
- Unsupervised and language agnostic.
- Can generate abstractions of the original sentence.
- Relatively simple to implement.
**Cons**
- The generated abstractions are not extremely novel since they are generated from within the original text words.
- The shortest path approach has issues when dealing with multi-word expressions.
- The quality of the model depends on the clustering quality.
#### **Deep Learning Methods**
Basically, by treating the task as a Many-to-One machine translation task where the input is the original documents and the target is the summary, the models are trained in a supervised manner and thus require a sizable amount of data to generalize, using approaches like Wikipedia summarization as outlined above.
**Pros**
- Truly abstractive and can yield
better summaries
**Cons**
- Need of large parallel training dataset.
- Overfitting issues when working with different genres or domains, this can be extremely problematic since the language model component of these systems can start generating word salad.
##
## **References:**
[1] P. J. Liu *et al.*, “Generating wikipedia by summarizing
long sequences,” *ArXiv Prepr. ArXiv180110198*, 2018.
[2] J. Zhang and X. Wan, “Towards Automatic Construction of News
Overview Articles by News Synthesis,” in *Proceedings of the
2017 Conference on Empirical Methods in Natural Language Processing*,
2017, pp. 2111–2116.
[3] A. R. Fabbri, I. Li, T. She, S. Li, and D. R. Radev,
“Multi-News: a Large-Scale Multi-Document Summarization Dataset
and Abstractive Hierarchical Model,” *ArXiv Prepr.
ArXiv190601749*, 2019.
[4] M. Yasunaga, R. Zhang, K. Meelu, A. Pareek, K. Srinivasan, and
D. Radev, “Graph-based neural multi-document summarization,”
*ArXiv Prepr. ArXiv170606681*, 2017.
[5] H. Liu, H. Yu, and Z.-H. Deng, “Multi-document summarization
based on two-level sparse representation model,” in *Twenty-ninth
AAAI conference on artificial intelligence*, 2015.
[6] J. Yao, X. Wan, and J. Xiao, “Compressive document
summarization via sparse optimization,” in *Twenty-Fourth
International Joint Conference on Artificial Intelligence*, 2015.
[7] S. Ma, Z.-H. Deng, and Y. Yang, “An unsupervised
multi-document summarization framework based on neural document
model,” in *Proceedings of COLING 2016, the 26th International
Conference on Computational Linguistics: Technical Papers*, 2016,
pp. 1514–1523.
[8] J. G. Carbonell and J. Goldstein, “The use of MMR,
diversity-based reranking for reordering documents and producing
summaries.,” in *SIGIR*, 1998, vol. 98, pp. 335–336.
[9] A. Nenkova, “Entity-driven rewrite for multi-document
summarization,” 2008.
[10] M. T. Nayeem, T. A. Fuad, and Y. Chali, “Abstractive
unsupervised multi-document summarization using paraphrastic
sentence fusion,” in *Proceedings of the 27th International
Conference on Computational Linguistics*, 2018, pp. 1191–1204.
[11] S. Banerjee, P. Mitra, and K. Sugiyama, “Multi-document
abstractive summarization using ilp based multi-sentence
compression,” in *Twenty-Fourth International Joint Conference
on Artificial Intelligence*, 2015.
## Frequently Asked Questions
### What is multi-document summarization?
Multi-document summarization (MDS) is the task of creating a single coherent summary from multiple input documents about the same topic. Unlike single-document summarization, MDS must handle redundancy, contradictions, and temporal ordering across sources.
### What is the difference between extractive and abstractive summarization?
Extractive summarization selects and concatenates the most important sentences from source documents. Abstractive summarization generates new sentences that paraphrase and condense the source content. Abstractive methods produce more natural summaries but are harder to implement and evaluate.
### How is multi-document summarization evaluated?
MDS is typically evaluated using ROUGE metrics (ROUGE-1, ROUGE-2, ROUGE-L) that measure n-gram overlap between generated and reference summaries. Human evaluation assesses coherence, informativeness, and fluency. BERTScore provides semantic similarity beyond surface-level matching.
### Smart Services for Social Media Marketing
- URL: https://mohammadshaker.com/en/blog/smart-services-for-social-media-marketing
- Date: 2020-01-26T00:00:00.000Z
- Tags: ideas, ml, nlp, social-media-marketing, Human-Written
NLP-powered smart services for social media marketing go beyond scheduling tools — they handle content selection, consumer intent analysis, trend detection, and automated generation. We survey the key service categories and the NLP techniques behind them: sentiment analysis, topic modeling, entity recognition, and text generation pipelines.
#### Content
Social media presents today a massive source of information for marketers and decision-makers to both better understand users trends and influence these users decision.
With the field of AI and NLP conquering various aspects of our day
to day life, social media and marketing using it is also an open
field for NLP applications.
In this post, we will explore the various different type of smart services that can help your social media marketing campaign.
### Content Selection
[Twizoo](https://www.linkedin.com/company/twizoo/about/) now acquired by Skyscanner was a startup that reads tweets feed on twitter filter it to include tweets most similar to your brand and then returns the tweets that will hopefully attract more customers to your site, this simple service would be a plugin your website where Twizoo will display the tweets about your brand.
The platform also includes some customers analysis and targeting features, the system in it’s simplest form can be built using 3 steps:
- Feed aggregator that collects recent tweets
- Feed filter that will tag the tweets using the entities that appear in them
- Sentiment analysis of other engagement metrics (like the number of likes found in the tweet) this is aimed at selecting only the positive feedback to be displayed on your home page
### **Consumer Intelligence**
This can be seen as the next step after content selection. In this step the goal is extracting insights from the collected data, [converseon](https://converseon.com/) is an example of such services where their services include:
- Pre-defined classification
- User customized classification
- Industry Trend Analysis
- New Product Development discovery
- Campaign Analysis
- New Market Analysis
- Data Visualization Workshop
While their site does not show what exactly do these features entail or how are they implemented, it is possible to schedule a demo.
### **Customer Service**
There has been a lot of advancements in the field of chatbots, mainly in limited chatbots specialized for a specific list of use cases.
Some of the leading companies in this field include ([conversocial](https://www.conversocial.com/) , [getJenny](https://www.getjenny.com/customer-service-chatbots-explained?utm_source=Google&utm_medium=CPC&utm_campaign=Sweden&utm_term=customer-service-chatbot&gclid=Cj0KCQiA0NfvBRCVARIsAO4930n-V46FddQeOOeAKgtMXYkWmLqfsrwah6SXZiO4TYP6EUaLhk58dzcaAhwrEALw_wcB), and [livePerson](https://www.liveperson.com/products/ai-chatbots/?utm_source=google&utm_medium=cpc&utm_campaign=BotsAI&utm_term=customer%20service%20chatbot&matchtype=e&gclid=Cj0KCQiA0NfvBRCVARIsAO4930m1Edhg7ntd6XGN6pmbsnBJwtPxWoZEx6neXD7WX0KtcSVZyD3IW2saAjHSEALw_wcB)) but the performance is not that good at the moment and these systems are usually as an assistant to a human employee.
### Influencer Marketing
The goal of this technology is to connect companies with influential people on social media, these influencers can range from celebrities to small content creators, in order to pitch the company message, One example is ( [InsightPool,](http://s.bl-1.com/h/J7QJzGV?url=https://insightpool.com/) currently named trendkit) a platform that searches through more than 600 million influencers across the social media spectrum to find the influencers who fit a brand’s unique characteristics, personality and goals, there is a whole field of NLP centred around public relations intelligence to facilitate connecting to new customers, suppliers, or marketing outlets. And there are a plethora of companies working in this area.
### **Content Optimization**
*The New York Times* internally built an app that they call Blossom that works within Slack. It's an intelligent bot that uses story engagement data to help them decide which of the 300 odd stories they have in a given day deserves a promotion or a higher amount of focus.
### **Competitive Intelligence**
Basically data mining to find competitors and gaps in the market, an example of such service is [unmetric](https://unmetric.com/).
##
### Auto-Tagging Content with NLP
- URL: https://mohammadshaker.com/en/blog/auto-tagging-content-with-nlp
- Date: 2020-01-21T00:00:00.000Z
- Tags: almeta.io, ideas, ml, nlp, research, Human-Written
Auto-tagging articles with NLP can save writers significant time and improve content discoverability. The main approaches are NER-based candidate extraction, graph-based keyword ranking like TextRank, statistical methods like TF-IDF, and deep learning keyphrase generation. Each trades precision for coverage in different ways.
#### Content
Many sites on the internet allow their users to specify tags for their content. The most famous example of such sites is Tumblr where each post on this social network can hold a manually selected set of tags. These tags can be useful to group the posts into related sets based on their topic and to facilitate searching.
Such tags can also be seen in news outlets or blogs where the author often add them as meta-data to improve their articles ranking in search engines like Google. If you are a writer you will recognize how tedious it can be.
In this article, we will explore the various ways this process can be automated with the help of NLP. Such an auto-tagging system can be used to generate possible tags for your posts or articles and allow you to select the most sensible for your article.
We will also delve into the details of what resources you will need to implement such a system and what approach is more favourable for your case.
# NER and NEL
Named Entity Recognition (NER) is the task of extracting Named Entities out of the article text, on the other hand, the goal of Named Entity Linking (NEL) is linking these named entities to a taxonomy like Wikipedia. If you are not familiar with NER and NEL you can review our previous article on this task.
One possible way to generate candidates for tags is to extract all the Named entities or the Aspects in the text as represented by say Wikipedia entries of the named entities in the article. Using a tool like [wikifier](http://wikifier.org/).
While this method can generate adequate candidates for other approaches like key-phrase extraction. It faces 2 issues:
- Coverage: well not all the tags in your article have to be named entities they might as well be any phrase.
- Redundancy: Not all the named entities mentioned in a text document are necessarily important for the article. For example in the following sentence “في بيان أصدرته مساء اليوم الأربعاء وتسلم مراسلنا ناصر حاتم نسخة منه، إن قطاع الأمن الوطني للوزارة رصد، في إطار جهوده "لكشف مخططات جماعة الإخوان الإرهابية والدول الداعمة لها” the named entity ناصر حاتم is irrelevant to the whole article purpose. Nonetheless, this would be suggested as a tag, which is not desired.
# Key Phrase Extraction
In keyphrase extraction the goal is to extract major tokens in the text. There are several methods. This can be done, and they generally fall in 2 main categories:
## 1. Unsupervised Methods
These are simple methods that basically rank the words in the article based on several metrics and retrieves the highest ranking words. These methods can be further classified into statistical and graph-based:
### Graph-Based Methods
In these methods, the system represents the document in a graph form and then ranks the phrases based on their centrality score which is commonly calculated using PageRank or a variant of it. The main difference between these methods lies in the way they construct the graph and how are the vertex weights calculated. The algorithms in this category include (TextRank, SingleRank, TopicRank, TopicalPageRank, PositionRank, MultipartiteRank)
### Statistical Methods
In this type the candidates are ranked using their occurrence
statistics mostly using TFIDF, some of the methods in this category
are:
- TFIDF: this is the simplest possible method. Basically we calculate the TFIDF score of every N-gram in the text and then select those with the highest TFIDF score.
- KPMiner: [1] the main drawback of using TFIDF is that it inherently has a bias for shorter n-gram since they would have larger scores. In KPMiner the system modifies the candidate selection process to reduce erroneous candidates and then adds a boosting factor to modify the weights of the TFIDF.
- YAKE: [2] introduces a method that relies on local statistical features from every term and then generates the scores by combining the consecutive N-words into keyphrases.
- EmbedRank: [3] This simple method uses the following steps:
- Candidates are phrases that consist of zero or more adjectives followed by one or multiple nouns
- These candidates and the whole document are then represented using Doc2Vec or Sent2Vec
- Afterwards, each of the candidates is then ranked based on their cosine similarity to the document vector
## 2. Supervised Methods
- [KEA](http://community.nzdl.org/kea/description.html) is a very famous algorithm for key phrase extraction. Basically it extract candidates from documents using TFIDF and then the trained model is then used to restrict the candidates set.
- Deep methods were also suggested to tackle this task, basically the task is converted to a sequence tagging problem where the input is the article text while the output is the BOI annotation.
## Other Notes
- A major distinction between key phrase extraction is whether the method uses a closed or open vocabulary. In the closed case, the extractor only selects candidates from a pre-specified set of key phrases this often improve the quality of the generated words but requires building the set as well it can reduce the number of key words extracted and can restrict them to the size of the close-set.
- Most of the aforementioned algorithms are already implemented in packages like [pke](https://github.com/boudinfl/pke).
- Some articles suggest several post-processing steps to improve the quality of the extracted phrases:
- In [3] the authors suggest using maximal marginal relevance(MMR) to improve the semantic diversity of the selected key-phrases. **They ran a manual experiment with 200 human participants and found that although reducing the phrases’ semantic overlap leads to no gains in F-score, the increased diversity selection is preferred by humans.** if you are not familiar with MMR you can learn more about it from our previous article on multi-document summarization.
- Several other approaches follow the same pattern to diversify their key phrases including [1, 2]
- Several cloud services including AWS comprehend and Azur Cognitive does support keyphrase extraction for paid fees. However, their performance in Arabic is not always good.
## Data
As mentioned above most of these methods are unsupervised and thus
require no training data. However, if you wish to use supervised
methods then you will need training data for your models.
In the case of Arabic, no large scale corpora are available, the largest we know of is AKEC which is still too small to be used by deep seq2seq models. However many sites add keyphrases to their articles especially news anchors making it fairly simple to scrap corpora of article, key phrases pairs. We already have the data for that in Almeta.
## Pros And Cons
- These methods are generally very simple and have very high performance.
- Most of these algorithms like YAKE for example are multi-lingual and usually only require a list of stop words to operate
- The unsupervised methods can generalize easily to any domain and requires no training data, even most of the supervised methods requires very small amount of training data.
- Being extractive these algorithms can only generate phrases from within the original text. This means that the generated keyphrases can’t abstract the content and the generated keyphrases might not be suitable for grouping documents
- The quality of the key phrases depends on the domain and algorithm used.
# Key Phrase Generation
A major draw back of using extractive methods is the fact that in
most datasets (in Arabic and other languages) a significant portion
of the keyphrases are not explicitly included within the text [4].
Key Phrase Generation treats the problem instead as a machine
translation task where the source language is the articles main text
while the target is usually the list of key phrases. Neural
architectures specifically designed for machine translation like
seq2seq models are the prominent method in tackling this task.
Furthermore the same tricks used to improve translation including
transforms, copy decoders and encoding text using pair bit encoding
are commonly used.
While the supervised method usually yield better key phrases than
it’s extractive counter-part there are some problems of using this
approach:
- These methods are usually language and domain-specific: a model trained on news article would generalize miserably on Wikipedia entries. This increases the cost of incorporating other languages.
- The deep models often require more computation for both the training and inference phases
- These methods require large quantities of training data to generalize. However as we mentioned above, for some domain such as news articles it is simple to scrap such data.
# Text Tagging
Another approach to tackle this issue is to treat it as a
fine-grained classification task. Where the input of the system is
the article and the system needs to select one or more tags from a
pre-defined set of classes that best represents this article. There
are 2 main challenges for this approach: choosing a model that can
predict an often very large set of classes, and obtaining enough data
to train it.
The first task is not simple several challenges have tackled this task especially the [LSHTC](http://lshtc.iit.demokritos.gr/) challenges series. The models often used for such tasks include boosting a large number of generative models [5] or by using large neural models like thous developed for object detection task in computer vision.
The second task is rather simpler, it is possible to reuse the data of the key-phrase generation task for this approach. Another large source of categorized articles is public taxonomies like Wikipedia and [DMOZ](https://dmoz-odp.org/).
One interesting case of this task is when the tags have a
hierarchical structure, one example of this is the tags commonly used
in a news outlet or the categories of Wikipedia pages. In this case
the model should consider the hierarchical structure of the tags in
order to better generalize. Several deep models have been suggested
for this task including HDLTex [6] and Capsul Networks [7,8]
The drawbacks of this approach is similar to that of key-phrase
generation namely, the inability to generalize across other domains
or languages and the increased computational costs.
# Ad-hoc solutions
In [9] a very interesting method was suggested. The authors basically indexed the English Wikipedia using Lucene search engine. Then for every new article to generate the tags they used the following steps:
- Use the new article (or a set of its sentences like summary or titles) as a query to the search engine
- Sort the results based on their cosine similarity to the article and select the top N Wikipedia articles that are similar to the input
- Extract the tags from the categories of resulted in Wikipedia articles and score them based on their co-occurrence
- filter the unneeded tags especially the administrative tags like (born in 1990, died in 1990, ...) then return the top N tags
This is a fairly simple approach. However, it might even be unnecessary to index the Wikipedia articles since Wikimedia already have an open free API that can support both querying the Wikipedia entries and extracting their categories. However, this service is somewhat limited in terms of the supported end-points and their results.
# Customizable Text Classification by Tagging
Several commercial APIs like [TextRazor](https://www.textrazor.com/topic_tagging) provide one very useful service which is customizable text classification. Basically, the user can define her own classes in a similar manner to defining your own interests on sites like quora. Next, the model can classify the new articles to the pre-defined classes.
Regardless of the method, you choose to build your tagger one very cool application to the tagging system arises when the categories come for a specific hierarchy. This case can happen either in hierarchical taggers or even in key-phrase generation and extraction by restricting the extracted key-phrases to a specific lexicon, for example, using DMOZ or Wikipedia categories.
The customizable classification system can be implemented by making the user define their own classes as a set of tags for example from Wikipedia, for example, we can define the class football players like the following set {Messi, Ronaldo, … }.
In the test case, the tagging system is used to generate the tags and then the generated tags are grouped using the classes sets.
If the original categories come from a pre-defined taxonomy like in the case of Wikipedia or DMOZ it is much easier to define special classes or use the pre-defined taxonomies.
# Conclusion
- There are several approaches to implement an automatic tagging system, they can be broadly categorized into key-phrase based, classification-based and ad-hoc methods
- For simple use cases, the unsupervised key-phrase extraction methods provide a simple multi-lingual solution to the tagging task but their results might not be satisfactory for all cases and they can’t generate abstract concepts that summarize the whole meaning of the article.
- More advanced supervised approaches like key-phrase generation and supervised tagging provides better and more abstractive results at the expense of reduced generalization and increased computation. They also require a longer time to implement due to the time spent on data collection and training the models. However, it is fairly simple to build large-enough datasets for this task automatically.
- The approach presented in [9] is a fairly general and simple, and it is possible to leverage c APIs to implement it in a fairly simple manner.
- **The simplest way to build a tagging system I out opinion is** to combine shallow key-phrase extraction with tags from WikiMedia to generate adequate tags. If the quality of the generated tags is not satisfactory to your application or if you want to support a limited set of tags then you may want to consider stronger options like a key-phrase generation or supervised tagging.
- One fascinating application of an auto-tagger is the ability to build a user-customizable text classification system. Such a system can be more useful if the tags come from an already established taxonomy.
##
# References
[1] El-Beltagy, Samhaa R., and Ahmed Rafea. "Kp-miner:
Participation in semeval-2." *Proceedings of the 5th
international workshop on semantic evaluation*. 2010.
[2] Campos, Ricardo, et al. "YAKE! Keyword extraction from
single documents using multiple local features." *Information
Sciences* 509 (2020): 257-289.
[3] Bennani-Smires, Kamil, et al. "Simple Unsupervised
Keyphrase Extraction using Sentence Embeddings." *arXiv
preprint arXiv:1801.04470* (2018).
[4] Meng, Rui, et al. "Deep keyphrase generation." *arXiv
preprint arXiv:1704.06879* (2017).
[5] Puurula, Antti, Jesse Read, and Albert Bifet. "Kaggle
LSHTC4 winning solution." *arXiv preprint arXiv:1405.0546*
(2014).
[6] Kowsari, Kamran, et al. "Hdltex: Hierarchical deep
learning for text classification." *2017 16th IEEE
International Conference on Machine Learning and Applications
(ICMLA)*. IEEE, 2017.
[7] Sinha, Koustuv, et al. "A hierarchical neural
attention-based text classifier." *Proceedings of the 2018
Conference on Empirical Methods in Natural Language Processing*.
2018.
[8] Aly, Rami, Steffen Remus, and Chris Biemann. "Hierarchical
multi-label classification of text with capsule networks."
*Proceedings of the 57th Annual Meeting of the Association for
Computational Linguistics: Student Research Workshop*. 2019.
[9] Syed, Zareen, Tim Finin, and Anupam Joshi. "Wikipedia as
an ontology for describing documents." *UMBC Student
Collection* (2008).
### Aspect-Level vs Entity-Level Sentiment Analysis
- URL: https://mohammadshaker.com/en/blog/aspect-level-vs-entity-level-sentiment-analysis
- Date: 2020-01-19T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Document-level sentiment analysis misses critical nuance: a review can hate a phone's RAM but love its price. Aspect-level sentiment analysis (ALSA) and entity-level sentiment analysis (ELSA) solve this by pinpointing the sentiment target. This post explains the difference and why it matters for political bias detection.
#### Content
First, a motivational example:
Many products on the internet allow the user to leave some feedback. This feedback is usually reviewed manually to figure out what are the users likes or dislikes in the product, what are the features they desire, and what are the problems they are facing. However wouldn’t it be amazing if we can find what specific aspect of the product does the review like or dislike? for example, in a review of a smartphone like the following:
> A: The RAM is really small, the price is low though
Would it be possible to figure out that the author hates the small RAM but is happy about the cheap price? OK, another question: what if another comment was like this:
> B: I totally hate this phone the memory is very low
Would it possible to figure out that the 2 comments are actually criticizing the same thing which is the memory capacity?
Well this is what we are talking about today. But first, let us understand the different levels of opinion mining.
This article is a part of our series on political bias detection we will hopefully introduce you to the various aspects of our political bias detection system, and you can learn about:
- [How can we predict the political orientation behind a piece of the news?](/blog/political-orientation-detection-ai-and-nlp-approach)
- [What is Stance detection? and what are the different types of it?](/blog/stance-detection-state-of-the-art)
- [What is subjective stance detection? what is distance supervision? and why they make a cute couple?](/blog/subjective-stance-detection-what-is-it-and-how-to-do-it)
- [What is ALSA, ELSA and how can opinion mining save us from political bias?](/blog/aspect-level-vs-entity-level-sentiment-analysis)
- [How to implement an initial political bias detector just from sentiment analysis and some probabilistic distribution? (warning cool visualizations)](/blog/from-sentiment-to-political-bias-in-the-arab-world-and-the-arabic-content)
- How to Visualize a Political Bias Metric
## Intro to opinion mining levels
In general, opinion mining aims to extract a quintuple from texts, where:
- e: is the entity or the target in our previous example the word "RAM" or the word "memory",
- a: is the aspect of the entity e in our example we can combine the 2 words under a single concept like "memory capacity"
- next h which is the opinion holder, here the guys A and B
- t: is the time when the opinion holder expresses her opinion on the entity e, which would be the time of the comments
- and finally s is the opinion or the sentiment expressed which is negative.
For example, opinion mining processes the review text “I bought a new iPhoneX today, the screen is great, but the voice quality is poor” and outputs two quintuples and .
The main difference between the levels of opinion mining is the
amount of information needed from this quintuple with higher level
tasks requiring lower amount of information, the following table
summarizes the various levels for the sentence “I bought a new
iPhoneX today, the screen is great, but the voice quality is poor”,
and we will revisit them in more depth later on.
| Task | Information needed | Description | Example |
| --- | --- | --- | --- |
| Sentiment analysis | | Given a piece of text extract the overall sentiment polarity | |
| Stance Analysis | | Given a piece of text and a target extract the sentiment polarity towards that target | |
| Aspect level sentiment Analysis | | Given a piece of text, a target and a particular aspect of that target extract the sentiment polarity towards that aspect | |
| Entity level sentiment analysis | | This is a lower level in which we extract the sentiment towards all the key-words (possible targets) within a piece of text regardless of whether they refer the same aspect or target | |
## What is ALSA and How is it Done
Different from stance detection, aspect level sentiment analysis aims at detecting relevant aspects and opinions. Following the general opinion mining framework, aspect mining can be formalized as the task of extracting triple ( e means target, a and s represent aspect and opinion respectively). An aspect can encompass multiple entities for example the entities “speed, latency, throughput, … ” can be grouped together in a single aspect namely responsivness. This task can be divided into 4 main subtasks, illustrated for the following sentence from [1]
“أعجب من كون الكتاب لم يصلنا إلا توًا كتاب رائع لموضوع مهم مسطرًا بلغة جميلة من كاتبة مبدعة”
- T1 Aspect term extraction: Given a piece of text the task
is to extract all the terms that can constitute targets, from the
previous example the terms would be (“كتاب”,
“لموضوع”,
“لغة”,
“كاتبة”)
Note that the terms are not necessarily a single word as they can be
any span of nominal phrases.
- T2 Aspect term polarity detection: from the output of the
previous task this task tries to assign to each of the terms a
polarity tag (positive, negative, neutral), the result of the
previous example would be
| Term | Polarity |
| --- | --- |
| كتاب | positive |
| موضوع | positive |
| لغة | positive |
| كاتبة | positive |
- T3 Aspect category identification: this is greatly similar
to target identification in [subtask
C of semeval 2019 task 6](http://alt.qcri.org/semeval2019/index.php?id=tasks) and [subtask
A of semeval 2016 task 5](http://alt.qcri.org/semeval2016/task5/) , In this task given a closed list of
predefined aspects and a piece of text the goal is to extract the
aspects within the text. here an aspect encompass more information
than a simple term, the mapping between terms and aspect of the
previous example is as follows:
| Term | aspect |
| --- | --- |
| كتاب | اصل |
| موضوع | محتوى |
| لغة | اسلوب |
| كاتبة | اصل |
Note that the
identification of the aspect category can be done either directly
from the text and aspect list, or it can be done using the output of
task T1 by using topic models like LDA [2]
- T4 Aspect
category polarity identification: finally in this task given the
identified aspects from T3 this task assigns for each of these
aspects a polarity class from (positive, negative, neutral), again
Note that this can be done using the output of T3 or by grouping the
results from T4. The results of this task on the aforementioned
example is as follows:
| Aspect | Polarity |
| --- | --- |
| اصل | Positive |
| محتوى | Positive |
| اسلوب | positive |
### The Data
There is a lot of data sets for the task of aspect level sentiment
analysis in English see the following sem-eval tasks: [2014
task 4](http://alt.qcri.org/semeval2014/task4/), [2015
task 12](http://alt.qcri.org/semeval2015/task12/), and [2016
task 5](http://alt.qcri.org/semeval2016/task5/) , in Arabic there is 3 available datasets [HAAD](https://github.com/msmadi/HAAD)
based on books reviews from good reads, [2016
task 5](http://alt.qcri.org/semeval2016/task5/) based on hotel reviews from TripAdvisor and [ABS](https://github.com/malayyoub/ABSA-for-Affective-News-Analysis-)[A](https://github.com/malayyoub/ABSA-for-Affective-News-Analysis-)
using news reviews from the 2014 war between Gaza and Israel.
The task of ALSA is well defined for the domain of product reviews, therefore the first 2 datasets have a relatively high quality although their domain is rather limited, on the other hand, while the latter data set is the most important for our task it has 2 main problems: firstly, the domain of the dataset is very limited in scope, and secondly and most importantly the annotation of the data with regards to the polarization and the aspects categories. the foll shows an example.
example of annotation errors from [ABSA](https://github.com/malayyoub/ABSA-for-Affective-News-Analysis-) dataset
### Implementation details
- The most naive manner in which this task can be tackled is
by solving T2 using lexicons. In this scheme a lexicon of sentiment
is used to find sentiment words in in the text and then linking the
sentiment with the closest target word. In both Arabic [3] and
English [4] the method usually relay on sequence tagging models
mostly shallow models mainly conditional random fields trained using
a plethora of syntactic features (POS, Lemma, dependency tree),
lexicon features, and semantic features like words embeddings, there
is a lot of emphases on the text preprocessing task in Arabic see
[5]. This is motivated by the low amount of available data.
- The SOTA in
this task depends on 2 subsystems (aspect identification and
sentiment assignment) and while the latter have usually a high
performance [6] and [4] report an f-measure of nearly 81% for that
task, the former task of target identification have really low
performance 66% in [4] and 69% in [6]
- Nonetheless This task is not directly applicable to our
task, mainly because of the absence of proper training data with the
correct aspects ([ABS](https://github.com/malayyoub/ABSA-for-Affective-News-Analysis-)[A](https://github.com/malayyoub/ABSA-for-Affective-News-Analysis-)
is the only Arabic data set related to news and its polarity and
aspects choice is very noisy). Furthermore, building such a data set
is extremely hard at least in comparison with the simpler tasks of
stance detection and direct political bias detection.
## Now What is ELSA?
This task is a lower level of the Aspect level sentiment Analysis, it can be seen as the application of both T1 and T2 from the ALSA. Namely the input to the system consists of only text, which can be comprised of one or multiple sentences, contain multiple entities with a varying sentiment, and have different domains. Our goal is to identify the important entities towards which opinions are expressed in the text; these can include any nominal or noun phrase, including events, or concepts, and they are not restricted to named entities. The only constraint is that the entities need to be explicitly mentioned in the text. See [this](https://developer.aylien.com/text-api-demo) demo to understand the task.
### Corpora and Data Sources
For the full task of ELSA the dataset of [AOT](http://www.cs.columbia.edu/~noura/Resources.html)
is an arabic dataset specifically developed for this task ,
furthermore since this task can be divided into 2subtasks as we shall
see next the datasets of [HAAD](https://github.com/msmadi/HAAD),
[ABS](https://github.com/malayyoub/ABSA-for-Affective-News-Analysis-)[A](https://github.com/malayyoub/ABSA-for-Affective-News-Analysis-),
and [2016 task 5](http://alt.qcri.org/semeval2016/task5/)
can also be used either for the whole task or for the first part
(Target detection) only
### Implementation details
As mentioned above this task incorporates the tasks T1 and T2 from
the Aspect level sentiment analysis task namely entity identification
and polarity assignment, In English the mainstream method is the
utilization of deep learning end-to-end models to both identify the
targets and assign the sentiment tag toward them [7], this is
motivated by the fact that jointly training on the 2 tasks can boost
the performance of them and that a cascaded system in which T2 is
tackled after T1 can cause a propagation of errors. However the use
of deep learning is facilitated by the abundance of data on both the
Entity level and Aspect level sentiment analysis in English.
In Arabic However the lack of data dictates the use of feature
engineering with shallow sequence tagging models such as CRFs.
Furthermore, these models address the 2 tasks separately with one
model to find the targets and another model to assign sentiment to
the found target, while this causes misses in target identification
to affect the sentient analysis model and fails to share the
information between the 2 models. This scheme have an added benefit
which is the ability to do sentiment analysis on any supplied list of
targets by substituting the closed list in place of the target
detector. Furthermore, such models are extremely fast in both
training and prediction phases.
The authors in [5] handle this task in Arabic using the aforementioned scheme, they relay on various features on the lexical and syntactic levels such as (POS, NER, dependency trees paths, Lemmas, and words segments), on the semantic level (using KNN clustering of the words embeddings) and by using Arabic and English semantic lexicons. For the tasks of segmentation, lemmatization and morphological features extraction the authors relay on the closed source [MADAMIRA](https://camel.abudhabi.nyu.edu/madamira/) , however, a lot of the features of that system can be extracted using [FARASA](http://qatsdemo.cloudapp.net/farasa/demo.html) which is open sourced (with the exception of detailed words segmentation “D3” and detailed dependency parsing and POS tags)
Based on this the final model to achieve this can utilize the method described in [5] combined with a Topic model to move from the entity to Target level.
## Conclusion
At this point hopefully you have a basic idea of the entity and aspect level sentiment analysis, so while designing the next Amazon you should be able to analyze how your customers are reacting to your product and hopefully, you will be able to get better insights on how to please them.
##
## Further reading
[1] M. Al-Smadi, O. Qawasmeh, B. Talafha, and M. Quwaider, “Human
annotated arabic dataset of book reviews for aspect based sentiment
analysis,” in *2015 3rd International Conference on Future
Internet of Things and Cloud*, 2015, pp. 726–730.
[2] D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent dirichlet
allocation,” *J. Mach. Learn. Res.*, vol. 3, no. Jan, pp.
993–1022, 2003.
[3] M. Pontiki *et al.*, “Semeval-2016 task 5: Aspect based
sentiment analysis,” in *Proceedings of the 10th international
workshop on semantic evaluation (SemEval-2016)*, 2016, pp. 19–30.
[4] T. Hercig, T. Brychcín, L. Svoboda, and M. Konkol, “Uwb at
semeval-2016 task 5: Aspect based sentiment analysis,” in
*Proceedings of the 10th international workshop on semantic
evaluation (SemEval-2016)*, 2016, pp. 342–349.
[5] N. Farra and K. McKeown, “Smarties: Sentiment models for
arabic target entities,” *ArXiv Prepr. ArXiv170103434*, 2017.
[6] A.-S. Mohammad, M. Al-Ayyoub, H. N. Al-Sarhan, and Y. Jararweh,
“An aspect-based sentiment analysis approach to evaluating arabic
news affect on readers,” *J. Univers. Comput. Sci.*, vol. 22,
no. 5, pp. 630–649, 2016.
### From Sentiment to Political Bias in the Arab World and the Arabic Content
- URL: https://mohammadshaker.com/en/blog/from-sentiment-to-political-bias-in-the-arab-world-and-the-arabic-content
- Date: 2020-01-19T00:00:00.000Z
- Tags: arabic, nlp, Human-Written
Political bias detection in Arabic news requires going beyond sentiment analysis — political framing, entity alignment, and source stance all carry ideological signal. We trace the pipeline from basic sentiment labeling to multidimensional orientation detection, covering the unique challenges Arabic presents: dialectal variation, implicit framing, and geopolitical media alignment patterns.
#### Content
The rise of political bias problem across several news anchors presents a real threat to free and independent journalism and a major factor in shifting the populace conception of the world. Several NGO's, research centres and private organizations are working on monitoring and limiting the spread of this type of content, especially in news outlets, usually in a manual fashion.
Most of the above-mentioned effort have been centred at detecting a bias towards one end of the political spectrum (Left vs Right), and while this fits to an accurate degree the politics of Europe and the USA, this dichotomy falls short when considering the politics of the Arabic speaking world. In the MINA region, politics are extremely more convoluted, with different parties appealing to a different aspect of the people life including religion, nationality, and traditional ethos. Furthermore, the various religious, racial, and national splits in these communities are explicitly utilized by the politicians.
Following our mission of building a reliable news source and battling Fake news, we, at ALMETA, have devoted our effort to build an automated political bias detection system, that can process thousands of articles and flag out potentially biased content this article gives a general outline one of the factors in our algorithm namely, emotional bias.
To better understand the design of the system we should first look at how the process is done manually. The manual assignment of political bias is usually based on [4 main metrics](https://mediabiasfactcheck.com/methodology/):
1. Biased Wording/Headlines- Does the source use loaded words to convey emotion to sway the reader.
2. Factual/Sourcing - Does the source report factually and back up claims with well-sourced evidence.
3. Story Choices: Does the source report news from both sides or do they only publish one side.
4. Political Affiliation: How strongly does the source endorse a particular political ideology? In other words how extreme are their views.
In this article, we will show how we simulated the first metric.
This article is a part of our series on political bias detection we will hopefully introduce you to the various aspects of our political bias detection system, and you can learn about:
- [How can we predict the political orientation behind a piece of the news?](/blog/political-orientation-detection-ai-and-nlp-approach)
- [What is Stance detection? and what are the different types of it?](/blog/stance-detection-state-of-the-art)
- [What is subjective stance detection? what is distance supervision? and why they make a cute couple?](/blog/subjective-stance-detection-what-is-it-and-how-to-do-it)
- [What is ALSA, ELSA and how can opinion mining save us from political bias?](/blog/aspect-level-vs-entity-level-sentiment-analysis)
- [How to implement an initial political bias detector just from sentiment analysis and some probabilistic distribution? (warning cool visualizations)](/blog/from-sentiment-to-political-bias-in-the-arab-world-and-the-arabic-content)
- How to Visualize a Political Bias Metric
## Entity-level Sentiment Analysis
It is not good enough to find the loaded wording in the article we are processing, it is equally important to finding the target of this polarity (hate or praise). We have already developed an entity-level sentiment analysis system. Another blog about this one is coming. **But in a nutshell, our system can detect meaningful targets in the text and assign to each of them a polarity. Both the target detector and polarity assignment models returns their confidence for each one of the targets.**
## How to model bias based on polarity
The simple intuition we used is that an article that excessively praises or criticizes a certain entity is a biased article, mainly because most of the things in life most of the time tend to be [just normal](https://opentextbc.ca/physicstestbook2/chapter/entropy-and-the-second-law-of-thermodynamics-disorder-and-the-unavailability-of-energy/). Therefore, we opted to use the confidence of our entity-level sentiment analysis to create a measure of article subjectivity.
### Show me the math!
Based on our assumption the unnormalized sentiment model confidence (coupled with the unnormalized target detection confidence) can give an estimate of emotion towards a target as follows:
where **$latex S\_{conf}$** is the NMLL (Negative Marginal Log Likelihood) confidence of the sentiment model, $latex T\_{conf}$ is the NMLL confidence of the Target model. This metric stays the same regardless of the assigned polarity and therefore, can be seen as a subjectivity metric of a certain target t. Note that for a Conditional Random Field (CRF) model with input X and output Y the NMLL score of a sub-sequence K from X to be tagged as a certain class is calculated as follows:
where k is a token in the sequence K and $latex P(Y k=class)$ is the marginal probability of the token k to be assigned the class, this marginal probability is estimated using the forward-backwards routine in the CRF calculation. Based on this both $latex S\_{conf} , T\_{conf}$ are bounded from above by 0 and from
below by $latex -K \* log(eps)$
To create a metric of the whole article we can use the sum of exponents of the previous subjectivity score across all the targets found by ELSA model, and to account for the variation in the number of targets per article which correlates with the article size we can normalize the sum by the number of words in the article, therefore the article score becomes:
However, in this type of analysis, we are assuming a uniform distribution of the scores. And here is where the second assumption comes in hand. The best way to visualize is to imagine a histogram of the news articles. Ideally, the histogram of the scores will be similar to the following figure where most of the article has low subjectivity with the number of articles falling with the increase of subjectivity.

image is taken from the public domain
Such a histogram would mostly be generated from a gamma distribution (red curve in image). By assuming that the scores of the articles follow a certain probabilistic distribution we can use its CDF (cumulative density function) to get a normalized scale that adheres to the second assumption.
For example, the following graphs illustrate the PDF and CDF for gamma distribution. Note how that many values on the linear scale especially near the long tail of the distribution are all mapped to the same probability, (basically of the case of the red pdf [k = 1.0, theta=2.0] all the values above
8 are mapped to a probability of nearly 1 in the CDF), this shows the difference between using a linear scale and a nonlinear one (as in a linear scale the above-mentioned values will occupy the probability range from 0.4 to 1.

image is taken from the public domain

image is taken from the public domain
Therefore the CDF probability of the articleScore can represent a good way to normalize the original article score.
Furthermore, the same analogy can be used if the scores are not distributed following gamma, since the CDFs across all of the probabilistic distributions share this favourable feature. Even in the case of a multi-headed histogram, where the data is generated from a combination of distributions, we can approximate the latent distribution using Mixture Models (for example [GMMs](https://towardsdatascience.com/gaussian-mixture-models-explained-6986aaf5a95)) and then map the scores using the GMMs CDF.

image taken from public domain
### How can I implement this?
1. Collect a large set of news articles (we already have this)
2. Run the model and calculate the unnormalized article score for each of them
3. Find the histogram of the unnormalized scores and fit a distribution to it (this can be any distribution or mixture of them but ideally this should be Gamma)
4. Use the CDF of the fitted PDF to get a normalizer of the scores this can be:
1. A closed-form solution: after fitting to a certain distribution use that distribution CDN equation to get the probability
2. A numerical solution: which is suitable for weird distributions like GMM, then we will have to store a discrete CDN as an array and use it to calculate the prop.
5. While the app works online Aggregate the unnormalized scores to create new sets and periodically re-adapt the distribution to the new data points (using [maximum a posterior](https://towardsdatascience.com/a-gentle-introduction-to-maximum-likelihood-estimation-and-maximum-a-posteriori-estimation-d7c318f9d22d) Algorithm for example)
### How did we implemented it?
At ALMETA we host a plethora of datasets among them a sizable corpora of Arabic News articles that we aimed to use to create our metric. However, applying the aforementioned methodology using this corpora is unfeasible because of its shear size, so before we can start calculating probabilities we needed to create a smaller (yet informative enough) dataset.
#### Selecting a informative dataset
The goal of this step is to select a smaller yet informative subset. For this subset, we will calculate the unnormalized articles scores and then fit them to a probability distribution.
the following figure shows the histogram of articles based on the word count in their text, Note that nearly
no articles have more than 2000 words, furthermore, we see that there is a sizable number of very short articles these includes:
- Video descriptions mainly from BBC
- Articles where the text is the same as the title (Aljadeed)
- Error in the crawler where only part of the text was retrieved (Alarabia)
To elevate this only articles with word count between 60 and 2000 words were considered, this figure shows the new distribution.
Next, we tried to get as many different topics as possible to do so we used the genre assigned by the author, Note that although this is a manual tag this classification is rather fuzzy and not strict, with several articles being erroneously and ambiguously classified, the following figure shows the distribution of articles based on genre. It is easy to see that many articles relate to the same meaning (journalism, press, inthepress) this is caused by the different naming conventions across the news domains these
articles should be combined in a single bucket, furthermore, many articles have a problematic genre:
- Due to errors in crawling (a lot of articles have numerical genres like 4667654321 or even hashes)
- Some articles have weird manual genres (year2013, 1300GMT)
To elevate these issues articles with the problematic genre were discarded and a mapping between the genre and topics of the site was created manually please Note that the building of this mapping was not accurately aimed at creating a topic classification set but to rather diversify the selected set as much as possible, and that the assignment of a certain genre to a common topic was based on reviewing a small random set of that genre articles. Following is the distribution of articles based on the new set of topic.
In order to select an informative (yet smaller) set from these articles, from every “topic” we sampled N random articles were N is:
$latex N=min(topicSize ,max(0.7∗topicSize ,100)) $
where topic size is the number of articles assigned to a certain topic, after this step the size of the selected dataset is 73294 articles, which while is still big is computationally attainable.
#### Data Modeling
After selecting the dataset the articles are passed to the system to calculate the article score. The following figure shows the distribution of articles scores note that as expected the distribution resembles a gamma CDF with
most of the articles having a low bias score.
To find the best fit we tried finding the distribution that fits the data with lowest SSE (sum of squares error) which turned out to be Gaussian following figure shows the PDF and CDF of this distribution, However, using a Gaussian violates the constraint that scores are greater than zero this can be easily seen from the fact that CDF of that distribution gives a considerable weight to negative scores (the probability of CDF of zero score is 5% instead of zero)
Several families were tested manually and we settled with a gamma distribution the final figure shows the PDF and CDF of this distribution, Note how in compassion with the Gaussian CDF the Gamma CDF gives nearly no weight to negative scores (the probability of CDF of
zero score is 0.04%)
#### What about the code?
we have created a general "enough" [python notebook](https://colab.research.google.com/drive/1JFyux5099AF__-34v5YtVAJzzjdT4yIi) for you to test this normalization method on your own hope you like Google's Collaboratory cause we sure do!
##
### Multidimensional Topic Modelling. The What? And the How?
- URL: https://mohammadshaker.com/en/blog/multidimensional-topic-modelling-the-what-and-the-how
- Date: 2020-01-19T00:00:00.000Z
- Tags: ml, nlp, Human-Written
Standard LDA assigns each document a topic distribution across a single latent dimension. Multidimensional topic modeling extends this to learn multiple latent variables simultaneously — such as topic, perspective, and writing style — in a single generative model, giving richer document representations for tasks like stance detection and bias analysis.
#### Content
Yes, understandably you might be thinking is this related to Rick and Morty?
Well unfortunately No. **But** you should really continue reading cause Multidimensional topic modeling is really cool.
In this short piece we will explore the fundamental idea behind multi-dimensional topic modeling and even give you a list of some open-sourced implementations so stay tuned.
## The What?
Well the task of Topic Modelling is the task of modelling the process of documents creation mostly through approximating the (topics, aspects, perspectives, styles, ….) of the document as a probabilistic distribution over the words.
### One Dimension
The aforementioned factors are called Latent variables (hidden factors that influences the document generation process and the choice of words) mostly these distributions are assumed to follow Dirichlet distribution.
LDA [1] is one of the oldest and most successful topic models, it assumes we have a set of Z latent components (usually called “topics” ), and each data point (document) has a discrete distribution over these topics.
The set of latent components usually relates to a single latent variable where the LDA tends to learn distributions which correspond to semantic topics (such as SPORTS or ECONOMICS) which dominate the choice of words in a document, rather than syntax, perspective, or other aspects of document content.
### Two Dimensions
Better modelling can be achieved by using more than a single set of latent components.
Imagine that instead of a one-dimensional vector of Z topics, we have a two-dimensional matrix with Z1 components along one dimension (rows) and Z2 components along with the other (columns).
This structure makes sense if your data is composed of two different factors, and the two dimensions might correspond to factors such as news topic and political perspective (if we are modelling newspaper editorials), or research topic and discipline (if we are modelling scientific papers). Individual cells of the matrix would represent pairs such as (ECONOMICS, CONSERVATIVE) or (GRAMMAR, LINGUISTICS). this is the idea behind the two-dimensional models like TAM [2] and SAGE.
### A Ton of Dimensions
We can expand the idea even further by assuming K factors modeled with a K-dimensional array, where each cell of the array has a pointer to a word distribution corresponding to that particular K-tuple.
For example, in addition to topic and perspective, we might want to model a third factor of the author’s gender in newspaper editorials, yielding triples such as (ECONOMICS, CONSERVATIVE, MALE).
Conceptually, each K tuple functions as a topic in the original LDA (with an associated word distribution ) except that K-tuples imply a structure, e.g. the pairs (ECONOMICS, CONSERVATIVE) and (ECONOMICS, LIBERAL) are related.
### Related algorithms
Other related approaches include the **Contrastive Opinion Summarization** task. The goal of this task is to extract sentences from positive and negative sets of opinions on a topic and generate a comparative summary containing sentence pairs that are both contrastive to each other and representative with respect to the given sets of opinions.
The method reported in [3] models the summarization task as an Optimization problem and it uses Natural Language Processing (NLP) and Optimization techniques to generate a representative and comparative summary from customer reviews about a topic, product or service.
## The How
In the following table we list some of the available implementations of the aforementioned algorithms
| Name | Programming Language: | Description |
| --- | --- | --- |
| Factorial LDA [code](https://cmci.colorado.edu/~mpaul/downloads/flda.php) | Java | Implementation of factorial LDA |
| ccLDA and TAM [code](https://cmci.colorado.edu/~mpaul/downloads/mftm.php) | Java | Implementation TAM [2] and ccLDA |
| [VODUM](https://github.com/tthonet/VODUM) | Java | Implementation of the Viewpoint and Opinion Discovery Unification Model [4] |
| [Contrastive Summarization](https://github.com/otvio/COSummarizer) | Python | Implementation of the Contrastive summarization model from [3] |
| [SeaNMF](https://github.com/tshi04/SeaNMF) | Python | Implementation of the Sea Nonnegative Matrix Factorization from [5] |
| [STTM](https://github.com/qiang2100/STTM) | Java | A Library of Short Text Topic Modeling |
## Conclusion
In this article we explored the idea of expanding topic modeling to cover various aspect and by now, you might be thinking ...
I told you didn't I.
But seriously how am I going to make a million-dollar from this knowledge? Well, if you want to see a cool application check out this [piece](/blog/viewpoint-topic-and-opinion-discovery-in-an-opinionated-document) to find out how it is possible to use multi-dimensional topic modelling to discover different political views in an opinionated text.
##
## Further reading
[1] D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent dirichlet
allocation,” *J. Mach. Learn. Res.*, vol. 3, no. Jan, pp.
993–1022, 2003.
[2] M. Paul and R. Girju, “A two-dimensional topic-aspect model for discovering multi-faceted topics,” in *Twenty-Fourth AAAI Conference on Artificial Intelligence*, 2010.
[3] H. D. Kim and C. Zhai, “Generating comparative summaries of
contradictory opinions in text,” in *Proceedings of the 18th ACM
conference on Information and knowledge management*, 2009, pp.
385–394.
[4] T. Thonet, G. Cabanac, M. Boughanem, and K. Pinel-Sauvagnat,
“Vodum: A topic model unifying viewpoint, topic and opinion
discovery,” in *European Conference on Information Retrieval*,
2016, pp. 533–545.
[5] T. Shi, K. Kang, J. Choo, and C. K. Reddy, “Short-text topic
modeling via non-negative matrix factorization enriched with local
word-context correlations,” in *Proceedings of the 2018 World
Wide Web Conference*, 2018, pp. 1105–1114.
### Viewpoint, Topic and Opinion Discovery in an Opinionated Document
- URL: https://mohammadshaker.com/en/blog/viewpoint-topic-and-opinion-discovery-in-an-opinionated-document
- Date: 2020-01-19T00:00:00.000Z
- Tags: almeta.io, ml, nlp, Human-Written
Two outlets can cover the same event with opposite viewpoints while using different words entirely. Probabilistic topic models like LDA and VODUM can discover these latent viewpoints, topics, and opinions from text without labels. This post explains the theory and how it applies to detecting partisan bias in news coverage.
#### Content
In one of our previous articles, we discussed the idea of multi-dimensional topic modelling, and no it is not related to Star Wars, so if you thought it is, go [here](/blog/multidimensional-topic-modelling-the-what-and-the-how) and give it a good read.
Back from Alderaan. Cool, let us get started. In this article, we are introducing and explaining the concepts of viewpoint, topic and opinion discovery from a text document in an unsupervised manner.
How can we model a viewpoint or a topic and the different techniques to classify documents based on their contrastive viewpoints. Then we discuss their implementations and evaluate their methods on different datasets.
Why does this matter? well in our work here at ALMETA to battle news bias and partisan coverage it is important for us to find and understand how different outlets view the same event and report it in different ways. interested? ok let us head on.
## The What
In an opinionated text, an author expresses her opinions on one or several topics, according to the author’s viewpoint. We define the key concepts of topic, viewpoint, and opinion as follows:
- A topic is one of the subjects discussed in a document collection.
- A viewpoint is the standpoint of one or several authors on a set of topics.
- An opinion is a choice of words that is specific to a topic and a viewpoint.
e.g. "Israel occupied the Palestinian territories of the Gaza strip", the topic is the presence of Israel on the Gaza strip, the viewpoint is pro-Palestine and the opinion is negative. Indeed, when mentioning the action of building Israeli communities on disputed lands, the pro-Palestine side is likely to use the verb to occupy, whereas the pro-Israel side is likely to use the verb to settle. Both sides discuss the same topic, but they use different wording that conveys an opinion.
## The How
Concepts like viewpoint or topic can be considered latent variables (hidden variables) because there are not explicitly shown in a text, so we leverage probabilistic topic models such as LDA [3] or TAM [4] to learn these concepts from the words and other signals itself like POS-tagging, or co-occurrence resolution ...
### Enter VODUM
In [2] the authors introduce a probabilistic topic model based on LDA called VODUM. VODUM identifies topical words and viewpoint-specific, topic-dependent opinion words, using part-of-speech tags by assuming that nouns are topical words; adjectives, verbs and adverbs are assumed to be opinion words while all other words with different tags are removed.
So each word has two options either being a topical word or an opinion word.
VODUM uses a hierarchical dependency structure through a graphical model where every node presents some distribution (viewpoint dist, topic dist,..) and edges present the conditional probabilities between nodes (aka dependencies). We can see the graph model in the following figure.
### How to model
papers related to topic modelling usually differ in structuring the dependencies and the relations between concepts (viewpoint, topic, opinion) and even the definition of the concepts itself, let us take [1] as an example:
Here the authors assume that a document is sampled from a mixture over topics as well as a mixture over viewpoints.
The two mixtures (topic and viewpoint) are drawn independently of each other, and thus can be thought of as two separate clustering dimensions. A word is associated with variables denoting its topic and viewpoint assignments, as well as two binary variables to denote if the word depends on the topic and if the word depends on the viewpoint. A word may depend on the topic, the viewpoint, both, or neither.
The paper [1] demonstrates the structure with a cool example, let’s consider a set of product reviews for a home theater system.
- Content topics in this data might include things like sound quality, usability, etc., while the viewpoints might be the positive and negative sentiments.
- A word like ‘speakers’, for instance, depends on the sound topic but not a viewpoint,
- While ‘good’ would be an example of a word that depends on a viewpoint but not any particular topic.
- A word like loud would depend on both (since it would be considered positive sentiment only in the context of the sound quality topic)
- While a word like think depends on neither.
On the other hand paper [2] describes the generative process for a document as modelled by VODUM.
Here the author writes the document according to her own viewpoint. Depending on her viewpoint, she selects for each sentence of the document a topic that she will discuss. Then, for each sentence, she chooses a set of topical words to describe the topic that she selected for the sentence, and a set of opinion words to express her viewpoint on this topic.
What path to use? well that is left to trail and error.
### How to Implement
There is plenty of resources and datasets to implement any of these algorithms please review our article [here](/blog/multidimensional-topic-modelling-the-what-and-the-how) for a full overview of the available dataset and codebases, as well as an example of the implementation.
## Conclusion
In this article we tried to cover the basic Idea of view point discovery using multidimensional topic modeling, we saw how the viewpoints, topics and opinions can be represented. How can we view the process of writing an Article in light of these concepts. And finally How we can reverse engineer this process to find the view points topics and opinions that are expressed in a piece of text.
As usual if you are interested in this task don't forget to check out the sources below.
##
## Further Reading
[1] M. J. Paul, C. Zhai, and R. Girju, “Summarizing contrastive viewpoints in opinionated text,” in
Proceedings of the 2010 conference on empirical methods in natural language processing,2010, pp.
66–76.
[2] T. Thonet, G. Cabanac, M. Boughanem, and K. Pinel-Sauvagnat, “Vodum: A topic model
unifying viewpoint, topic and opinion discovery,” in European Conference on Information
Retrieval, 2016, pp. 533–545.
### An Overview of the Event Extraction Task in NLP
- URL: https://mohammadshaker.com/en/blog/an-overview-of-the-event-extraction-task-in-nlp
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Event extraction identifies triggers, arguments, and roles from text using NLP, covering ACE-style structured extraction and open-domain approaches.
#### Content
Events possess a rich structure that is important for intelligent information access systems (information retrieval, question answering, summarization, etc.). Without information about what happened, where, and to whom, temporal information about an event may not be very useful.
In light of this importance, the event extraction task emerges. In compare to the [**event detection**](/blog/event-detection-in-media-using-nlp-and-ai) task, which aims to discover related stories in a continuous stream of news articles, **event extraction** tries to extract the information about an event that is reported in one or multiple documents.
## Task Description
The [ACE](https://en.wikipedia.org/wiki/Automatic_content_extraction) program provides annotated data and evaluation tools for a variety of information extraction tasks. There are five basic kinds of extraction targets supported by ACE: entities, times, values, relations, and events.
### The ACE Event Model
- An **event** is something that happens.
- An **event extent** is a sentence within which a taggable event is described.
- An **event trigger** is the word that most clearly expresses the occurrence of an event (main verb, adjective, past-participle, nouns, and pronouns).
- Event **participants** are the Entities that are involved in that event.
- Event **attributes** are frequently entities and values within the scope of an event that are not properly participants.
- **Event arguments** are event participants and event attributes.
- An **event properties** are several properties related to the event e.g., when and if the event really took place. Currently, they are the features (Polarity, Tense, Genericity and Modality).
- A **participant role**: each event type and sub-type will have its own set of potential participant roles for the entities which occur within the scopes of its exemplars.
In the ACE model, only "interesting" events (events that fall into one of 34 predefined categories) are extracted.
As a part of this program corpora that support the tasks (entities, relations, events) were developed for English, Chinese, and Arabic. Moreover, the guidelines to annotate an Arabic event extraction corpus according to the ACE model were described in [1].
## Methodologies
We distinguish between two main approaches for event extraction, in analogy with the classic distinction that is made in the field of modeling.
### Data-Driven Event Extraction
Despite their differences, **all approaches focus on discovering statistical relations**, i.e., facts that are supported by statistical evidence. Examples of discovered facts are words or concepts that are (statistically) associated with one another. **However, statistical relations do not necessarily imply semantically valid relations, nor relations that have proper semantic meaning.**
Several examples of the usage of the data-driven approaches for event extraction can be found in the literature. For instance, [2] broke down the ACE task of extracting events into a series of classification sub-tasks, each of which is handled by a *machine-learned classifier*:
1. Triggers Identification: finding event triggers in text and assigning them an event type. That was modeled as a word classification task.
2. Argument identification: determining which entity mentions are arguments of each event mention. That was modeled as a pair-classification task i.e. Each event mention is paired with each of the entity mentions occurring in the same sentence to form a single classification instance.
3. Attribute assignment: determining the values of the modality, polarity, genericity, and tense attributes for each event mention. A separate classifier was trained for each attribute.
4. Event coreference: determining which event mentions refer to the same event. Each event mention in a document is paired with every other event mention, and a classifier assigns to each pair of mentions the probability that the paired mentions corefer.
Clustering techniques were also employed to solve this task. For instance, [3] aimed for real-time news event extraction, but focus especially on violence and disaster events.
They developed an event extraction engine, which for each detected violent event produces a frame, whose main slots are: date and location, number of killed and injured, kidnapped people, actors, and type of event.
First, their event extraction system, used linear patterns in order to extract entities that have specific semantic roles in a news cluster. Second, they merged the single extracted pieces into event descriptions via application of information aggregation algorithm.
They used machine learning algorithms for the acquisition of the patterns. However, ML approaches are never 100% accurate, therefore they manually filtered out implausible patterns and added hand-crafted ones.
In [4] they simplified the problem into a sentence-level classification problem. **They used machine learning classification methods to differentiate between sentences that describe one or more event and those that do not.**
### Knowledge-Driven Event Extraction
In contrast to data-driven methods, knowledge-driven models are often based on patterns that express rules representing expert knowledge. It is inherently based on linguistic and lexicographic knowledge, as well as existing human knowledge regarding the contents of the text that is to be processed.
In [5] they aimed to recognize events in the Arabic language. They were interested in the annotation of verbal events only. Although they didn't depend on huge lexical resources they constructed a minimal set of general and simple hand-crafted rules for time and location expressions recognition.
Another work in the Arabic language is [6]. They considered only verbs and nouns as events while adjectives are less significant. Their system was built using [GATE](https://en.wikipedia.org/wiki/General_Architecture_for_Text_Engineering). It identified predefined named entities (names of "people", "places", "organization", and "date"), and the relations between the entities and the defined events.
The extraction of an event consisted of the discovery of links between the "trigger" of the event and its arguments. The extraction of the link is established based on a syntactic analysis "dependency analysis" and of extraction rules exploiting this analysis. To implement this task they used JAPE transducer provided by GATE Toolkit. While the identification of triggers was done by the use of manually collected gazetteers.
## Conclusion
The event extraction task aims to extract the information related to the events mentioned in texts. It's considered useful for many NLP tasks including information retrieval, question answering, summarization, etc.
Extracting events is a complex task consisting of multiple sub-tasks of varying difficulty, involving detection of event triggers, assignment of attributes, identification of arguments and assignment of roles, and determination of event co-reference.
The task has many simplifications and variations in literature. Many methods including data-driven and knowledge-driven ones have been explored to solve this problem.
In this article, we introduced the task of event extraction. If you are interested in this task, you can review the rest of our series on the related topics of even detection:
- [What is Event detection? and how is it done?](/blog/event-detection-in-media-using-nlp-and-ai)
- [How can sequential clustering enable event detection from large streams of data?](/blog/an-implementation-of-a-news-stream-sequence-clustering-algorithm)
##
## References
[1] Linguistic Data Consortium. "ACE (Automatic Content Extraction) Arabic Annotation Guidelines for Entities." (2005).
[2] Ahn, David. "The stages of event extraction." *Proceedings of the Workshop on Annotating and Reasoning about Time and Events*. 2006.
[3] Tanev, Hristo, Jakub Piskorski, and Martin Atkinson. "Real-time news event extraction for global monitoring systems." *Joint Research Center of the European Commission, Web and Language Technology Group of IPSC, TP* 267.
[4] Naughton, Martina, Nicholas Kushmerick, and Joseph Carthy. "Event extraction from heterogeneous news sources." *proceedings of the AAAI workshop event extraction and synthesis*. 2006.
[5] Aliane, Hassina, Wassila Guendouzi, and Amina Mokrani. "Annotating events, time and place expressions in arabic texts." *Proceedings of the International Conference Recent Advances in Natural Language Processing RANLP 2013*. 2013.
[6] Hkiri, Emna, Souheyl Mallat, and Mounir Zrigui. "Events automatic extraction from Arabic texts." *Natural Language Processing: Concepts, Methodologies, Tools, and Applications*. IGI Global, 2020. 1686-1704.
## Further Reading
[1] Hogenboom, Frederik, et al. "An Overview of Event Extraction from Text." *DeRiVE@ ISWC*. 2011.
### Analysis of the Readability Metric Results in Almeta News Feed
- URL: https://mohammadshaker.com/en/blog/analysis-of-the-readability-metric-results-in-almeta-news-feed
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, nlp, Human-Written
Applying the AARIBase readability formula to Almeta's Arabic news feed shows a clear pattern: longer articles score lower on readability, and long sentences have an outsized effect. Two articles with identical read-time can differ dramatically in readability score based entirely on sentence length, not word complexity.
#### Content
In this post, we're analyzing the results returned by the readability metric in our news feed. If you haven't checked our post about "*How to measure the readability of a text?*" before, you can read about it [here](/blog/how-to-measure-text-readability).
## How Are We Measuring the Readability?
The main part of analyzing a metric is to know how does it work. In the current version, we're depending on the AARIBase metric for measuring the readability. So, let's have a look first on how does AARIBase work.
Here's the AARIBase formula:
**AARIBase** = (3.28 × NOC) + (1.43 × ACW) + (1.24 × AWS)
Where:
**NOC**: Number of characters.
**ACW**: Average character per word.
**AWS**: Not that [AWS](http://www.amazon.com/aws), It's the Average words per sentence.
According to the
considered factors in the previous formula and their weights, the
following results can be concluded:
- The NOC factor which has the largest weight causes a high sensitivity to the text length, longer texts will be less readable.
- Texts with longer words will be less readable, this is because of the ACW factor in the formula, the NOC factors may have a little effect in this too.
- Because of the AWS factor, texts with longer sentences will be less readable.
## Almeta News Feed Results Analysis
One thing that's remarkable in the results **is that the readability decreases when the read-time increases. Articles with very high read-time have very low readability and vice versa.**
Here are some examples:
This kind of results is reasonable according to our previous discussion on AARIBase formula, which is sensitive to the text length.
However, sometimes the articles may have the same read-time but they vary a lot in terms of readability. In this situation, mostly the length of the sentences is the factor that's behind this contrast in the readability scores.
Here is an example of this situation:
[The link to the article](https://www.aljazeera.net/news/scienceandtechnology/2019/8/27/%D8%BA%D9%88%D8%BA%D9%84-%D9%85%D8%AA%D9%87%D9%85%D8%A9-%D8%A7%D9%84%D8%AA%D9%84%D8%A7%D8%B9%D8%A8-%D9%86%D8%AA%D8%A7%D8%A6%D8%AC-%D8%A7%D9%84%D8%A8%D8%AD%D8%AB-%D8%B5%D8%A7%D9%84%D8%AD-%D8%AE%D8%AF%D9%85%D8%A7%D8%AA%D9%87)
[The link to the article](https://www.aljazeera.net/news/cultureandart/2019/8/27/%D9%84%D8%A3%D8%AC%D9%84-%D9%81%D9%84%D8%B3%D8%B7%D9%8A%D9%86-%D8%A7%D9%84%D9%85%D9%88%D8%B3%D9%8A%D9%82%D9%89-%D9%81%D9%8A-%D8%AE%D8%AF%D9%85%D8%A9-%D8%A7%D9%84%D8%AA%D8%B1%D8%A7%D8%AB-%D8%A7%D9%84%D8%A8%D8%AF%D9%88%D9%8A-%D8%A8%D8%BA%D8%B2%D8%A9)
In the previous example, you can examine that the sentences of the first article are longer than the sentences of the second article, which is reflected in the readability scores.
## Conclusion
In this post, we showed the analysis of our readability metric. We first analyzed the formula of the AARIBase metric which is the used readability metric, then we reflected this analysis on the results returned in our news feed.
##
### Aspect Detection and Named Entity Linking (NEL): Using SPARQL and DBpedia
- URL: https://mohammadshaker.com/en/blog/aspect-detection-and-named-entity-linking-nel-using-sparql-and-dbpedia
- Date: 2020-01-18T00:00:00.000Z
- Tags: almeta.io, engineering, ml, nlp, Human-Written
Named entity linking (NEL) connects entity mentions in news text to structured knowledge bases like DBpedia via SPARQL queries, enabling richer aspect detection. We use this pipeline to capture how different publishers cover the same entity — person, organization, location — and surface cross-article interaction patterns.
#### Content
In our effort to provide the best news feed out there, **one of the goals we are trying to achieve here at Almeta is to capture the interaction between different news outlets and how the coverage of the same event is presented by different views and different publishers.**
Imagine having the ability to combine all the articles from various news outlets that are talking about a certain event in one place that gives you a summary of all of the coverage while at the same time providing you with the different views on this event.
The process of Named Entity Linking (NEL) is a crucial part of our progress towards this goal. this article will introduce you to this amazing task in NLP and give an initial solution to it.
## The What
To understand NEL we must first understand the core concept of Information Extraction (IE.)
While IE encompasses various applications from DNA mapping to weather and time series forecasting, the common aspect in all of these tasks is the fact that we are trying to extract structured information from unstructured data.
The original data can be measurements from a telescope or texts from Wikipedia but the goal remains the same.
### NLP Example
Let us assume that out of a paragraph we have the following sentence
> ... and it is believed that Tim was born in London on the 8th of June 1955 ...
The process of IE on this sentence is shown in the following figure see how that the unstructured text data is converted into a structured semantic graph. A broad goal of information extraction is to extract knowledge from unstructured data and use that obtained knowledge for various other tasks.

Now let us take a deeper look at this "information extraction" algorithm.
The process can be done in three simple steps:
1. Find the entities in the text (in our case Tim, London and the date)
2. Disambiguating entities: for each of the entities found in the text we need to know what is the physical entity they correspond to for example the disambiguator have found from the context that we are talking about [Tim Berners Lee](https://en.wikipedia.org/wiki/Tim_Berners-Lee) and have linked it to a resource in a knowledge base like Wikipedia.
3. The last step is to find relations between the entities, for example, the relation between Tim and London is that he was born there
The process seems simple enough for a human annotator, but how would a computer be able to do it and do it efficiently?
The aforementioned list represents 3 different tasks in NLP:
1. Named Entity Recognition (NER): given a text find all named entities and assign to each of them a class (Person, Place, ...)
2. Named Entity disambiguation (NEL): given a text and a list of entities from it assign each entity to a resource from a knowledge base
3. Relation Extraction (RE): given a text and a list of entities from it assign find relations that link these entities
### NER VS NEL
A named entity is a real-world object, such as persons, locations, organizations, etc. NER identifies and classify named entity occurrences in text into pre-defined categories. NER is modelled as a task of assigning a tag to each word in a sentence. NER will tell us what words are entities and what are their types.
On the other hand, NEL will assign a unique identity to entities mentioned in the text. In other words, **NEL is the task to link entity mentions in text with their corresponding entities in a knowledge base**. The target knowledge base depends on the application, but we can use knowledge bases derived from Wikipedia for open-domain text.
## The Why
The space of possible applications of the NEL is simply massive, here is a list of what you can do:
- By disambiguating the entities we can have more accurate search services,
- We can extract more rich information from our articles and compose them into a semantic web, see the figure below, this allows us to answer questions like (whos is the wife of whom) or (how many children does X have)
- And way way more
But how does it help our goal: consider the following phrase from a news article:
> وقال الجبير في مؤتمر صحفي بالعاصمة البريطانية لندن "مقتنعون من خلال الأدلة الموجودة لدينا بتورط الجيش الإيراني في هجمات أرامكو"
The quote translates to :
> Ajoubair said during a press conference in the british capital London "we are convenced based the evidences we have that the Iranian Army had a role to play in Aramko attacks"
One of the entities extracted from this text would be "الجيش الإيراني" (the Iranian Army) and we can see that the sentence convenes a negative sentiment towards this entity. Now consider the following excerpt describing the same event just in a different way:
> و جدد الجبير اتهامه للقوات العسكرية الإيرانية بالوقوف وراء حادثة أرامكو
Which translates to:
> Aljoubair repeated his acqusations of the Iranian armed forces in being behined the Aramko incident
In order to have a true understanding of these articles, it is important for our platform to be able to say that the entities "الجيش الإيراني" (the Iranian Army) and "للقوات المسلحة الإيرانية" (the Iranian armed Forces) represents the same physical thing.
This is a basic building block in finding relations between different articles.
## The How
Hopefully, you now understand what is NEL and why it is a good idea to have it around now let us find out how, wow, we can do it.
### The Lazy Way
If you are as lazy as me and happens to be working on texts for either English, German or Portuguese, then you are in for a treat. You basically need to do nothing since the people at DBPedia Spotlight have a service ready for you that will do all of the fuss and wikify your text directly. Don't believe me? test the [demo](https://www.dbpedia-spotlight.org/demo/). Technically, the service well be supporting other languages for the future.
### The Naive Way
Follow this if you are lazy and you are working on other languages like Arabic. We will also be able to use DBPedia to link our data by a very simple approximation:
Here, we are assuming that the NER step is already implemented. There is plenty off the shelf packages to achieve this, cool? let's talk NEL.
Now you have a system that takes in a text and spits out a list of entities with their classes, the simplest way to disambiguate these entities goes in 3 steps for every entity do the following:
1. Find a ranked list of resources from a knowledge base that matches your entity, in case of DPBedia you can use their amazing text search service, [here](https://wiki.dbpedia.org/lookup). The ranking is usually based on the occurrence on commonly used metric is inverse candidate frequency measure (ICF), which is used to weight the words in the context based on their whole frequency.
2. Filter out the list based on the named entity class. If your entity is a person, you won't need resources that are places or dates.
3. Select the most frequent candidate
The shortcoming of this approach is pretty clear since we are linking any mention of the entity with the most frequent disambiguation, but if you can accept the hit to the performance then [Wikifier](http://wikifier.org/) is your guy.
### The Right Way, aka the, "Really missed up I have no life" Way
OK don't freak out it is not that hard.
The fact is: NEL is a really wide research area and usually the task is converted into an ML problem. The papers in the reference section are some of the best options to start with. All of the aforementioned papers are cross-lingual or multi-lingual and some of them have their own implementations open-sourced.
It would really hard for us to cover all of the fields in this article and therefore we will focus on one particularly amazing work. DeepType [1] this paper by open.ai is the current state of the art on this task.
In this approach, the authors rely on very detailed NER that can find very specific classes like Animal, Road, vehicle or Region. And from these classes they add symbolic constraints on the output of their ML model by splitting the learning process in 2 steps:
1. Finding the list of classes to be used based on DBpedia classes and types;
2. Training a constrained ML model using those types;
I am obviously glossing over a ton of details and if you like math then you should totally give the paper a read.
## Conclusion:
In this article, we introduced the task of named entity linking, explained what is it and why you should but an effort to incorporate it in your system. We have as well introduced some ways to tackle this issue while leaving you with a reading list to delve deeper. check the references list.
##
## References:
[1] Raiman, Jonathan Raphael, and Olivier Michel Raiman. "DeepType: multilingual entity linking by neural type system evolution." *Thirty-Second AAAI Conference on Artificial Intelligence*. 2018.
[2] Le, Phong, and Ivan Titov. "Improving entity linking by modeling latent relations between mentions." *arXiv preprint arXiv:1804.10637* (2018).
[3] Ganea, Octavian-Eugen, and Thomas Hofmann. "Deep joint entity disambiguation with local neural attention." *arXiv preprint arXiv:1704.04920* (2017).
[4] Kolitsas, Nikolaos, Octavian-Eugen Ganea, and Thomas Hofmann. "End-to-end neural entity linking." *arXiv preprint arXiv:1808.07699* (2018).
### Automatically Tagging Data for Content Informativity Scoring
- URL: https://mohammadshaker.com/en/blog/automatically-tagging-data-for-content-informativity-scoring
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, informativity, ml, nlp, Human-Written
Training a supervised informativeness classifier requires labeled data, but manually annotating thousands of Arabic articles is impractical. We use summarization-based similarity as a proxy label — comparing each article's lead paragraph to its summary — and validate this approach against the Kalimat and EASC Arabic datasets.
#### Content
In this task the goal is to assign a given piece of text a tag (or
number) representing the level of informativeness or detail this text
holds usually by training a model to do that.
Here, we rely on the intuitive idea suggested in [1] which states that usually in any news article the lead paragraph can be used as a good summary especially if the author utilizes the [inverted pyramid style](https://en.wikipedia.org/wiki/Inverted_pyramid_(journalism)). And therefore lead paragraphs that uses this style are more informative than those that does not based on this the authors suggested that a corpora for informativeness prediction can be automatically generated using a corpora of human summarized news articles simply by comparing the human generated summary with the lead paragraph and assigning the lead paragraph a tag of (informative/ creative) based on this similarity.
You can read more about this idea at our previous post [Summarization for Informativeness](/blog/can-you-measure-a-text-informativeness-using-its-summary)
## What Constitutes a Lead Paragraph
The boundaries of the lead paragraphs across the different
datasets is not clear with no simple separator, the only assumption
we have is that the lead marks the start of the article. Based on
this we decided to empirically select the first N sentences from the
start as the lead paragraph the choice of N was done empirically
using the elbow method by observing the change in the overall
standard deviation of the similarity (averaged across the 5 metrics)
for different values of lead size between 1 and 10 sentences and we
found that 4 sentences gave the highest relative reduction in overall
std without impacting the relative mean similarity.
## What Similarity Metric to Use
In our previous article We have decided to go with a simple similarity metric for our dataset selection process, [Jaro-winkler similarity](https://en.wikipedia.org/wiki/Jaro%E2%80%93Winkler_distance) is a normalized text similarity function similar to edit-distance, but we are now using a different similarity function based on [1], The score is computed as the fraction of words in the lead that also appear in the summary.
## What Dataset to Use
We have carried out a detailed exploration of the available summarization datasets in Arabic You can read more about this idea at our previous post [Summarization for Informativeness](/blog/can-you-measure-a-text-informativeness-using-its-summary) for now we will detail the datasets we are currently using:
### Kalimat
Kalimat is a very large (20K) multi-purpose dataset, the original documents are scraped from several newspapers websites, it covers several features including morphological analysis, single and multi-doc summarization and NER as well as the topic. The summaries in here are extractive summaries generated automatically and thus are of lower quality than other datasets. The dataset is composed of 5 genres of articles ('Culture', 'Economy', 'International', 'Local', 'Religion', 'Sports').
To create an informativeness dataset from it we carried out the following steps:
- Filtering by compression ratio: There is a large number of very short articles, to which the summary is basically a permuted version of the original text to resolve this we selected only articles that have a compression ratio greater than 0.6 where the compression ratio is given by:
$latex CR = 1 - numWordsCompressed/numWordsOriginal $
- Filter by similarity: It is reasonable to expect that our indirect annotations will be noisy. To obtain cleaner data for training of our model, for each genre, we only use the leads with scores that fall below the 20th percentile and above the 80th percentile.
The following table shows the results of the filtering process and the resulting dataset size:
| Dataset | Culture | Economy | International | Local | Sports | Religion | total |
| --- | --- | --- | --- | --- | --- | --- | --- |
| original | 2495 | 3265 | 1689 | 3237 | 3475 | 4095 | 18256 |
| After CR filter | 916 | 1045 | 556 | 939 | 1889 | 741 | 6086 |
| After similarity filter | 368 | 418 | 224 | 379 | 761 | 298 | 2448 |
| Low percentile | 0.147 | 0.185 | 0.949 | 0.159 | 0.157 | 0.179 | - |
| High percentile | 0.508 | 0.621 | 0.973 | 0.469 | 0.45 | 0.63 | - |
here
are some of the issues we observed:
- The filtering process has drastically reduced the size of the dataset
- The distribution of the articles based on the similarity value is nearly the same for all the genres (with the exception of international news) in all of these genres it is easy to see that the distribution is biased towards smaller similarity values with 80th percentile being lower than 0.6 for all of the genres (with the exception of the international news) the following figure shows the distribution of the articles based on the similarity value for the culture genre top and the economy buttom.
- For the international news we noted that the similarity values are concentrated mostly above 0.9, the distribution is shown in the following graph, for this genre, the similarity value is not of significance.
- After random inspection of some of the results we found some irrational tags, to investigate further we manually annotated 5% of each genre to measure the accuracy of the automatic tags. The annotation process was done using the lead paragraph only with the goal of measuring if the lead follows the inverted triangle style or not. The following table shows the results of this annotation process. Note that the religion genre was not considered
| Genre | Culture | Economy | International | Local | Sport |
| --- | --- | --- | --- | --- | --- |
| f1\_pos | 0.68 | 0.73 | 0.88 | 0.67 | 0.62 |
| f1\_neg | 0.65 | 0.67 | 0.67 | 0.53 | 0.50 |
| f1\_macro | 0.67 | 0.70 | 0.77 | 0.60 | 0.56 |
Note
that the overall results are not very good, furthermore, we noted
that the positive class had superior results across all the genres
while the negative class results were worst this was caused by the
lower precision and is understandable due to the inherent bias we saw
in the distribution of the similarity metric values.
- Furthermore the manual inspection have shown
that not all of the genres have a comparable language to the news
articles processed by us, most specifically the genres of religion
and to some extent culture and local news genres
### EACS
EASC is a small **extractive** dataset it consists
of 153 short articles extracted from wikipedia and Alwatan and Alrai
newspapers, with each of them 5 reference manual summaries named {A
to E}, and although most of the articles comes from Wikipedia nearly
all of them adhere to the inverted triangle scheme and thus can be
utilized for the summaries are of news articles it is possible to use
them to generate informativeness tags, however the quality of the
reference summaries are of varying degrees since they are generated
using Mturk.
In this dataset, we only used the text-similarity filter as this dataset didn’t have the problem of long summaries seen in KALIMAT. Note that although the dataset does include genre annotations the step was carried out on the whole dataset since it is already pretty small. Following is the results of this step:
| reference | A | B | C | D | E |
| --- | --- | --- | --- | --- | --- |
| Low percentile | 0.244 | 0.25 | 0.196 | 0.17 | 0.232 |
| High percentile | 0.602 | 0.555 | 0.527 | 0.527 | 0.568 |
Note that the
resulting dataset size in all of the cases is 62 samples, we can also
see that while all of the references showed a bias towards the
negative class with the 80th percentile lower than 0.6,
To ensure the quality of the resulting dataset we manually inspected the resulting dataset and found that nearly all the leads were informative since the inverted triangle style is prominent in Wikipedia entries we don’t really believe the automated annotation scheme is viable for this dataset.
### MULTI-Ling 2011, 2013
Both these datasets are extremely small for us to
carry out this analysis and through manual examination of the 30
leads found in these 2 articles we found that all of them are
informative since they are as well entities of Arabic Wikipedia.
## Conclusion
Overall we found that the usage of summarization datasets for informativeness training in Arabic is problematic, overall the datasets of EASC and MULTI-Ling can be faithfully considered informative, however we lack samples of non-informative samples the usage of the KALIMAT dataset for this purpose is risky and even if considered as an option the only genres that can be used are international and economy parts. And even in this case, the resulting dataset would still be too small.
##
# References
[1] Y. Yang and A. Nenkova, “Detecting information-dense texts in
multiple news domains,” in *Twenty-Eighth AAAI Conference on
Artificial Intelligence*, 2014.
### Can You Measure a Text Informativeness Using Its Summary?
- URL: https://mohammadshaker.com/en/blog/can-you-measure-a-text-informativeness-using-its-summary
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, nlp, Human-Written
If a summary captures the essential information in an article, then similarity between the full text and its summary should proxy for informativeness. We test this hypothesis using ROUGE, cosine similarity, and other metrics against Arabic summarization datasets, examining whether high similarity reliably predicts that an article is informative rather than creative.
#### Content
In a previous article (see next paragraph) we explored how to approximate an article informativeness in a supervised fashion, such a method would require training data, in this article we will explore on way to get this data, one very unusual way to do it.
This article is a part of our research on informativeness, and we have A LOT of it if you are interested in more details you can check each individual piece to learn, or [review the gist of it.](/blog/informativity-detection-almetas-research-gist)
- [What makes an article informative in our view and how can we measure this quantitatively?](/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of)
- [How to detect cliches in text? and why do they matter?](/blog/how-to-detect-cliches-in-text)
- [How to measure the informativeness of each individual word?](/blog/term-informativeness-estimation-in-the-arabic-language)
- [How to train a supervised model to measure informativeness?](/blog/supervised-article-informativeness-prediction-the-what-and-the-how)
- [Can we get data to train such a model from summaries?](/blog/can-you-measure-a-text-informativeness-using-its-summary) ?
- [Can we get such data in Arabic?](/blog/automatically-tagging-data-for-content-informativity-scoring)
- [How can we rank articles based on how informative they are?](/blog/how-to-rank-articles-based-on-how-informative-they-are-using-snorkel)
# What is the goal?
In this task, the goal is to assign a given piece of text a tag (or number) representing the level of informativeness or detail this text holds usually by training a model to do that.
Here we rely on the intuitive idea suggested in [1] which states
that usually in any news article the lead paragraph can be used as a
good summary especially if the author utilizes the [inverted
pyramid style](https://en.wikipedia.org/wiki/Inverted_pyramid_(journalism)), an example of a good lead paragraph looks
something like this:
“The European Union’s chief trade negotiator, Peter Mandelson,
urged the United States on Monday to reduce subsidies to its farmers
and to address unsolved issues on the trade in services to avert a
breakdown in global trade talks. Ahead of a meeting with President
Bush on Tuesday, Mr. Mandelson said the latest round of trade talks,
begun in Doha, Qatar,in 2001, are at a crucial stage. He warned of a
”serious potential breakdown” if rapid progress is not made in
the coming months.”
however, this generalization does not apply on all authors as some
uses a more creative approach, e.g.
“ART consists of limitation,” G. K. Chesterton said.” The most beautiful part of every picture is the frame.” Well put, though, the buyer of the latest multi-million dollar Picasso may not agree.
But there are pictures – whether sketches on paper or oils on canvas – that may look like nothing but scratch marks or listless piles of paint when you bring them home from the auction house or dealer. But with the addition of the perfect frame, these works of art may glow or gleam or rustle or whatever their makers intended them to do.
It is clear that the first lead paragraph is more informative than
the second. The authors relied on the idea that leads like the first
example constitute a more plausible summary of the article than the
second example, and thus suggest that a corpora for informativeness
prediction can be automatically generated using a corpora of human
summarized news articles simply by comparing the human generated
summary with the lead paragraph and assigning the lead paragraph a
tag of (informative/ creative) based on this similarity.
# What Similarity Metric to Use?
We have decided to go with a simple similarity metric for our dataset selection process, [Jaro-winkler similarity](https://en.wikipedia.org/wiki/Jaro%E2%80%93Winkler_distance) is a normalized text similarity function similar to edit-distance. This choice is based on 2 factors:
- While, ideally, we should use more informative distances like cosine similarity between the Doc2Vec representations of the summary and the lead paragraph, the usage of this simpler metric can represent a stronger constraint on the similarity and is acceptable since the summaries we are using are extractive.
- On the other hand, this is a normalized metric which can account for the variations in summary and article length.
This metric value ranges between 0 and 1, where 1 represent a total match and 0 total difference (No match at all).
# Available Summarization datasets:
## Arabic:
- The [wiki-how](https://github.com/mahnazkoupaee/WikiHow-Dataset) dataset was built using the wikiHows website, it is a very large **extractive summaries** dataset, the English version contains around 500k articles. It is based on aggregating the titles sentences from the steps of a method to create an overall summary of that method, since the site is multi-lingual the same method can be applied to the Arabic version of the site (or any other language for that matter)
- [DUC2004](https://duc.nist.gov/duc2004/) is an instance of the Document Understanding Conference Datasets for **extractive summarization**, this version includes noisy machine translated Arabic documents along side there summaries, the DUC dartasets are thoroughly studied throughout the literature and they are based on newswires. However, because the Arabic documents are machine translated they are relatively very noisy, further more acquiring this datasets requires a formal request and several bureaucratic forms to be filled.
- [RTLTDS](https://rtltds.github.io/index.html) is a large collection of Iranian scientific publications along with their abstracts, keywords, and authors
- [MULTILING2011](http://multiling.iit.demokritos.gr/file/view/353/tac-2011-multiling-pilot-dataset-all-files-source-texts-human-and-system-summaries-evaluation-data) is a small dataset of **abstractive summaries** based on wikinews where the events are aggregated and summarized it is composed of 10 events each covers 10 short articles the data is manully translated ans summarized in 7 languages (Arabic, Czech, English, French, Hebrew, Hindi, Greek) with a limit of 250 words for each event summary, over all the data is small and the quality is OK with some small problems such as spelling errors, more importantly the dataset contains multiple references, enabling the usage of metrics such as rouge.
- [MULTILING2013](https://drive.google.com/file/d/0B31rakzMfTMZRTZiM29UR3VxYmc/view) is an updated version of multiling2011, here the source of the articles is Wikipedia entries and not events and the summaries are **extractive** , the dataset includes 40 languages instead of 7 and for each language there is 30 articles with a single source manual summary, the summaries quality is relatively high. However, the original articles have scraping issues and needs near manual cleaning.
- [MULTILING2019Task1](http://multiling.iit.demokritos.gr/pages/view/1648/task-financial-narrative-summarization) is another dataset from multi ling that focuses on financial documents, although it is very large the quality is questionable since the documents are OCRed from pdf files and there is a lot of issues with the source text
- [EASC](https://sourceforge.net/projects/easc-corpus/) is a small **extractive** dataset it consists of 153 short articles extracted from wikipedia and Alwatan and Alrai newspapers, with each of them 5 reference manual summaries named {A to E}, and although most of the articles comes from Wikipedia nearly all of them adhere to the inverted triangle scheme and thus can be utilized for the summaries are of news articles it is possible to use them to generate informativeness tags, however the quality of the reference summaries are of varying degrees since they are generated using Mturk, after inspection the main issue we found with this dataset is inconsistency in both the articles length and the summaries length mainly because there was no restriction on the annotators with regards to the summary length, the articles have an average word length of 383 words but the standard deviation of 180 is relatively large, the same can be said to all the reference summaries whose average word length is around 130 across the 5 references but with an std that can reach 88 with an average compression rate of 0.32, overall the most consistent of the annotators seems to be E with the lowest std of 74, However if the we consider only the summaries shorter than 350 words, nearly all the references have the same std of 70, with the Exception of reference E which still have the lowest std of 65, However manual inspection of the summaries reveal no significant difference (except is individual cases) between the annotators behavior and thus any of them can be used, furthermore to asses the usability of the data set for informativeness annotation we used the first 4 sentences of every article as the lead paragraph and then calculated the distance between the lead paragraph and each of the reference summaries using the similarity metric. The choice of lead size is done empirically using the elbow method by observing the change in the overall standard deviation of the similarity (averaged across the 5 metrics) for different values of lead size between 1 and 10 sentences and we found that 4 sentences gave the highest reduction in overall std without impacting the similarity value. Following is the histogram of overall similarity.
- [Kalimat](https://sourceforge.net/projects/kalimat/)
is a very large (20K) multi-purpose dataset, the original documents
are scraped from several newspapers websites, it covers several
features including morphological analysis, single and multi-doc
summarization and NER as well as topic. The summaries in here are
extractive summaries generated automatically and thus are of lower
quality than other datasets, we conducted manual inspection of the
sports part of the dataset and the following are our observations:
- the length of the generated summaries is relatively
consistent across all the articles with a mean of 174 words and a
small std of 23 words, compare this with the massive variance of
235 found in the articles lengths, the following figure shows the
histogram of the summaries and articles lengths. This means that
there is great loss in information especially for longer articles.
- There is a large number of very short articles, to which the summary is basically a permuted version of the original text in case of the sport genre out of 4100 articles nearly 1400 articles have the same length as their summaries.
- While the generated summaries have low fluency, the main reason for this low quality is not omission of meaningful sentences in the article but rather because of the erroneous re-ordering of the sentences by the summarizer, this phenomenon appears clearly in short documents, while this issue is critical for the task of automatic summarization, it should be possible to use this dataset for the task of informativeness if an order agnostic distance measure was used like the one we are using. The following figure shows the distribution similarity function between the lead paragraph (4 sentences again) and the summary.
## English:
View [this
awesome list](https://github.com/mathsyouth/awesome-text-summarization#corpus), however here are some honorable mentions that are
truly eye opener:
- [TLDR
2017](https://zenodo.org/record/1168855) is massive 3 million document dataset that is collected
from Reddit site using the TL;DR tag from the posts, [this](https://www.reddit.com/r/bestofTLDR/)
is a whole community for the TLDR see [2]
## Which dataset to use?
| Dataset | genre | Pros | Cons |
| --- | --- | --- | --- |
| [wiki-how](https://github.com/mahnazkoupaee/WikiHow-Dataset) | Tips and life hacks | - multi-lingual - relatively large | - summaries are too short - this data set can’t be used for informativeness detection |
| [DUC2004](https://duc.nist.gov/duc2004/) | News wires | - there is previous literature on it - based on news wire and thus can be used for informativeness | - relatively small - is based on machine-translation and thus very noisy - requires form filling and no less than 7 working days for response |
| [RTLTDS](https://rtltds.github.io/index.html) | Scientific publications | | |
| [MULTILING2011](http://multiling.iit.demokritos.gr/file/view/353/tac-2011-multiling-pilot-dataset-all-files-source-texts-human-and-system-summaries-evaluation-data) | News articles about events | - multi-lingual - multiple reference summaries - quality is ok | - very small 10 events per language where each event is segmented into 10 very short articles - it is not trivial to utilize for informativeness since the events are not 100% articles are represents a chronological order of facts rather than an inverted triangle scheme |
| [MULTILING2013](https://drive.google.com/file/d/0B31rakzMfTMZRTZiM29UR3VxYmc/view) | Various entities from Wikipedia that ranges from cities to characters to sites | - multi-lingual - quality is high | - very small 30 articles per langauge - it is not trivial to utilize for informativeness since the wikipedia entries does not necessarily follows the inverted triangle scheme |
| [MULTILING2019Task1](http://multiling.iit.demokritos.gr/pages/view/1648/task-financial-narrative-summarization) | financial narrative disclosures | - multi-lingual - relatively large | - quality is questionable since the documents are OCRed from PDF files and there is a lot of issues with the source text - the domain of the text is very narrow - is not applicable to informativeness |
| [EASC](https://sourceforge.net/projects/easc-corpus/) | Wikipedia entries and news stories | - high quality - multi reference | - too small only 153 articles - nearly all of the lead paragraphs have the same distance (further investigation using other similarity metrics is needed) |
| [Kalimat](https://sourceforge.net/projects/kalimat/) | NewsWires | - very large - single and multi-document | - questionable quality since the summaries are automatically generated |
##
# References:
[1] Y. Yang and A. Nenkova, “Detecting information-dense texts in
multiple news domains,” in *Twenty-Eighth AAAI Conference on
Artificial Intelligence*, 2014.
[2] M. Völske, M. Potthast, S. Syed, and B. Stein, “Tl; dr: Mining
reddit to learn automatic summarization,” in *Proceedings of the
Workshop on New Frontiers in Summarization*, 2017, pp. 59–63.
### Clickbait Detection Using Word2Vec Representation
- URL: https://mohammadshaker.com/en/blog/clickbait-detection-using-word2vec-representation
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, nlp, Human-Written
Representing clickbait headlines as averaged word2vec vectors captures semantic similarity between sensational phrases better than bag-of-words models. We use t-SNE projections to validate whether clickbait and non-clickbait headlines cluster separately in embedding space before training a classifier on Arabic news headlines.
#### Content
In a previous article, [How to Detect Clickbait Headlines using NLP?](/blog/how-to-detect-clickbait-headlines-using-nlp) We introduced the task of clickbait detection and explored how it can be modeled within the domain of machine learning and NLP. If you are not familiar with the concept of clickbait detection, make sure to review it before continuing.
In this post, we're building a classifier for clickbait detection in the news headlines depending on a pre-trained Arabic Word2Vec model and we're validating this solution. If you are not familiar with the Word2Vec concept you can refer to [this](https://en.wikipedia.org/wiki/Word2vec) Wikipedia article for more information.
## **News Headlines Representation**
In order to get a vectorized representation for a headline using the word2vec, we aggregated the vectors of all the words in that headline by averaging them. The headline can be represented in this manner since it's a too short text, while this method is not applicable to other longer text parts like the content of the article.
## **Representation Effectiveness**
As the word2vec features are not easily interpret-able and trying to get an insight into the effectiveness of its representation for this problem we investigated the projection of the training set headlines vectors on the two-dimensional space using the t-SNE algorithm.
The following graph shows the projection:
**neg**: the negative class, indicates "not-clickbait".
**pos**: the positive class, indicates "clickbait".
As we can see in the projection, all the clickbait headlines are nearly grouped together in a cluster separating them from the not-clickbait headlines. While the occurred overlapping may be caused by three sources:
- Annotation error.
- Projection error.
- Representation error, which will be a part of our classification error.
Hence, we can say that this representation can effectively represent our problem, but not perfectly. So, we can proceed in our implementation.
Another noticeable thing from this projection is that the classes are not linearly separable.
**Now.. is it? It can represent the problem but they're definitely not 100% separable since there's an area of intersection.**
Well, It's as I think, after inspecting the results, and in comparison to our first implementation that relied on hand-crafted features at least. Moreover, I think most of our error using this representation is because of the shortage in coverage in the training set, not because of the used representation itself. Another point is that I think it shouldn't be perfect to be effective.
## Testing on Real-World Data
After training a clickbait GBoost classifier on our training dataset using the discussed representation, in this section, we're analyzing its performance on real-world data.
### **The Test Dataset**
The dataset consists of a collection of articles headlines collected from ALMETA's news feed data plus other headlines scrapped from the news outlets' websites. It contains ~14,000 headlines, the following pie graph shows these headlines distribution over the news outlets they were collected from:
For a valid evaluation, none of these domains were used to collect the training dataset.
### **The Probability Distribution of Being Clickbait**
Let's investigate the probability distribution of being clickbait that's produced by this classifier:
Most of the headlines seem to have low probabilities of being clickbait, which makes sense, since clickbait is an anomaly characteristic, thus most of the headlines in the real world should not be clickbait.
### **Investigating The Learned Features**
To be more confident in the model let's try to discover the features that were used by it to separate the classes, it would be a good sign if they make sense to us as humans.
The following method was applied to discover the types of words that are used by the classifier to detect clickbait:
1. Clustering the Arabic words (the vocabularies of the used pre-trained word2vec model) according to their word2vec representation, such that semantically related words \_from the model viewpoint\_ are grouped together.
2. Finding the word-clusters that mostly appeared in the headlines that have a high probability of being clickbait.
3. Highlighting the words that belong to each of these clusters in those headlines.
Following are examples from our test set with the discovered words in bold:
| **The words-cluster interpretation** | **Examples** |
| --- | --- |
| **Adjective** | بريد جستون تتلقي صدمه **جديده** ثغره **خطيره** في انستغرام صورك وفيديوهاتك ليست خاصه بك بعد انستغرام فيسبوك تختبر ميزه **صادمه** للمستخدمين |
| **Words common in clickbait** | باختصار خمس **حقائق** مثيره عن اصحاب العيون الخضراء **اسرار** جمال منزليه .. لا تعرفها سوي المراه التركيه |
| **Time Adverbs** | **بعد** ان حير العلماء الكشف عن اصل وموطن الحشيش مغنيه اميركيه تعتنق الاسلام **بعد** حادث نيوزلندا اشعر ببراءه الطفوله |
| **Stopwords common in clickbait** | من الاعتزال الي احضان الغوريلا حلا شيحه تثير الانتقادات **مجددا** ناسا تقرر العوده برواد فضاء الي القمر **مره اخري** |
| **Question words** | **هل** يمكنك البحث عن معلومه داخل فيديو ؟ غوغل ترد ثلاث صور غدت ايقونات ل 11 سبتمبر 2001 .. **فماذا** يقول اصحابها ؟ |
| **Time adverbs, demonstrative pronouns, etc.** | فانتازيا بشائر طيبه **حينما** تتفوق الدراما علي الروايه الاصليه لعشاق الايفون **هكذا** سيصبح سعر الهاتف لو صنع في اميركا السيارات الكهربائيه واقع يشبه الخيال العلمي **ولكن** |
| **Currencies** | ارتداء الشورت في السعوديه يكلفك خمسه الاف **ريال** وسم غرامه البيجامه الف **دينار** يتصدر بالكويت والسبب |
| **Superlative adjective** | **اطول** رحله جويه في العالم تهبط في نيويورك فكم دامت هل يساريو اليد **اذكي** **وانجح** من يمينيي اليد |
Recall the handcrafted features used in the clickbait problem from [How to Detect Clickbait Headlines using NLP?](/blog/how-to-detect-clickbait-headlines-using-nlp) In comparison, these features seem sensible.
### Measuring The Error Rate
We're determining the rate of the error made by the model in different sub-ranges from the probability range of being clickbait, based on a manual inspection of a set of randomly sampled examples, where for each sub-range we sampled examples that are classified with probabilities from this sub-range and count the occurred errors by labeling these examples manually.
#### **The False-Positive Error**
These are headlines classified as being clickbait while they're not. The following graph shows the distribution of this error over the probability sub-ranges:
So, ~6% of the not-clickbait headlines may be classified with a high probability of being clickbait. As the probability of being clickbait approaches to 50%, this error rate increases.
#### The **False-Negative Error**
These are headlines classified as being not-clickbait while they're clickbait, in other words, clickbait headlines that were classified as being clickbait with very low probabilities. The following graph shows the rate of this error in the low probability range [0, 50[:
So, ~10% of the clickbait headlines may be classified with a low probability of being clickbait.
#### **Analyzing The Error Sources**
We manually checked the false-positive errors searching for frequent patterns in them, and we found some words-types that are frequently appeared there:
| **Type** | **Examples** |
| --- | --- |
| **Person-related words** | يونيسيف **الاطفال** يشكلون ثلث ضحايا تجاره البشر في العالم الصحف المصريه تركز الاهتمام علي مصرع **ابن** كاهانا **وزوجته** ومقررات القمه الخليجيه |
| **Cars-related words** | سياره **هيونداي** سانتا في بجيلها الرابع تحسين موديل اكس 4 من **بي ام دبليو** |
| **Art related words** | ساره **والموسيقي** مجموعه من **الاوركسترات** الرائعه وفاه **عازف** **البيانو** الشهير فيكتور بورغ |
These word-types are clearly neutral and shouldn't be indicators of clickbait.
The following graph shows the rate of the presence of each of that words-type in the false-positive errors:
To determine the cause of this error type, we investigated the presence of these words-types in the training set using the following method:
1. Clustering the Arabic words (the vocabularies of the used pre-trained word2vec model) according to their word2vec representation, such that semantically related words \_from the model viewpoint\_ are grouped together
2. We find the words-cluster corresponds to each words-type.
3. We calculate the rate of the headlines that contain words from this cluster in both clickbait and not-clickbait classes in the training set.
The following graph shows the rates of words-types in both classes in the training set:
So, it seems the training set is biased in terms of these word-types, as they are over-appeared in the clickbait class examples while they're rare in the not-clickbait class examples, thus the classifier learned to give them high weights as being clickbait indicators.
Examples of the presence of these word-types in the clickbait class in the training set:
| **Type** | **Examples** |
| --- | --- |
| **Person-related words** | تعرفوا على سبب اكل هذه **طفله** لسجاد والرمل بالفيديو: لحظه انقاذ **فتاه** حاولت الانتحار من اعلى مبنى صور **شاب** يحول نفسه الى اميرات ديزني بمكياج يستحيل ان يكتشفه احد |
| **Car-related words** | صور: سعودي يحصل على سياره **تويوتا** هديه .. تعرف على السبب فيديو **سوبارو** معدله تخرج من حفره عميقه بطريقه مذهله! شاهد بالصوره ... عمر البشير بملابس بيضاء وسياره " **لاند كروزر** "... اول ظهور للبشير منذ الاطاحه به |
| **Art-related words** | شمبانزي بدرجه **فنان** **يعزف** على **الجيتار** بطريقه رائعه بالفيديو : اب يشارك طفلته **رقصه** مجنونه برفقه المكنسه رد فعل مفاجئ لانثى اسد لم يعجبها **عزف** احد الزوار |
## Conclusion
In this post, we proposed and validated our solution for the clickbait problem using the word2vec representation of the headlines of the articles. We also tried to interpret the behavior of the produced model by discovering the learned features and trying to analyze its error. The model seems to be affected by the domains of the articles that was trained on. To improve the results we need to build more diverse and unbiased dataset.
I don't have any other proposals at all. I believe the article content is not useful in this problem, moreover, it can't be represented using the word2vec vectors average because it's long, and representing it with a different representation then concatenating it with the title seems a mess. The only solution I have is to improve the dataset.
##
### Comparison of Available TTS Services
- URL: https://mohammadshaker.com/en/blog/comparison-of-available-tts-services
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, nlp, research, Human-Written
A comparison of text-to-speech services for Arabic, evaluating Google Cloud TTS WavNet and other APIs on voice quality, latency, and cost.
#### Content
Text-to-speech (TTS) is a type of [assistive technology](https://www.understood.org/en/school-learning/assistive-technology/assistive-technologies-basics/assistive-technology-what-it-is-and-how-it-works) that reads digital text aloud. It’s sometimes called “read aloud” technology.
text-to-speech applications are offering an innovative solution for users to interact with content by taking it out of books and computer screens and integrating it into any environment that the user finds convenient.
Most people lead a busy lifestyle and complain that they don’t have time to read, which is totally cool, because with text-to-speech, you don’t really have to carry a book or scroll infinitely through your favorite blogs or publications when you could be spending that time doing something else.
There is a lot of companies offering TTS APIs, for Arabic languages we took three of the best APIs for now:
- *Google Cloud TTS API WavNet*
- *Google Cloud TTS API Non-WavNet*
- *Read speaker API*
This table shows the comparison between these APIs:
| | **GCP TTS API** **WaveNet voices** | **GCP TTS API Non-WaveNet voices** | **Read speaker** |
| --- | --- | --- | --- |
| Close to real speech | 8/10 | 5/10 | 5/10 |
| Handles numbers | Very Good | Very Good | Very Good |
| Handles proper names and places | Good | Good | Good |
| Handles ambiguous words | Bad | Bad | Very Bad |
| Number of Voices | 3 | 3 | 2 |
| Pricing | $16.00 USD / 1 million characters | $4.00 USD / 1 million characters | starting at $4/month |
## **Google Cloud TTS API**
Google Cloud Text-to-Speech converts text into human-like speech in more than 180 voices across 30+ languages and variants. It applies groundbreaking research in speech synthesis (WaveNet) and Google's powerful neural networks to deliver high-fidelity audio. With this easy-to-use API, you can create lifelike interactions with your users that transform customer service, device interaction, and other applications.
### 1 - Pricing:
Cloud Text-to-Speech is priced monthly based on the amount of characters to synthesize into audio sent to the service.
| Feature | Monthly free tier | Paid usage |
| --- | --- | --- |
| Standard (non-WaveNet) voices | 0 to 4 million characters | $4.00 USD / 1 million characters |
| WaveNet voices | 0 to 1 million characters | $16.00 USD / 1 million characters |
### 2 - Voices and Language:
It supports more than 180 voices across 30+ languages and variants, including the Arabic language with six voices.
You can listen to the voices from [here](https://cloud.google.com/text-to-speech/docs/voices?authuser=1).
### 3 - Max size of the request and number of the requests per minute:
Content limit: 5,000 Total characters per request
Requests limit: 300 Request limit per minute, 150.000 Characters per minute
You can try it by this [code](https://colab.research.google.com/drive/1Yo19n0k4g43lssuAaTLnDfDSRbn585V-).
## **Read Speaker**
### 1 - Pricing:
From individual complete subscriptions starting at $4/month to institutional licenses, ReadSpeaker TextAid is the most cost-effective solution available today. [Contact us](https://www.readspeaker.com/contact/) about multi-user licenses or click here for more information about ReadSpeaker TextAid for Individuals and to sign up for a free trial.
### 2 - Voices and Languages:
It supports about 30+ languages including the Arabic language with tow voices (Male and Female). You can list to the voice samples from [here](https://www.readspeaker.com/languages-voices/#lv-voice-samples).
> Note:
>
> IBM Watson Text to Speech and Microsoft Azure and alot don't support the Arabic language.
### How we can use TTS in app, natively by using the Mobile OS (Android or iOS) itself.
There might be several ways to use TTS offline, like:
- Flutter plugins like [flutter\_tts](https://pub.dev/packages/flutter_tts) or [sytody](https://github.com/rxlabz/sytody) or others.
- Java libraries like [FreeTTS](https://freetts.sourceforge.io/) or [AndroidMaryTTS](https://github.com/AndroidMaryTTS/AndroidMaryTTS) or others.
- IOS libraries like [iphone-tts](https://bitbucket.org/sfoster/iphone-tts/src/default/) or [TTSOverview-iOS](https://github.com/meredian/TTSOverview-iOS) or others.
but it depends on the device support and even if it is achieved Arabic is not supported.
## Conclusion
For Arabic language TTS APIs, it is still not good enough like English language TTS APIs. That's because it has a lot of ambiguous words like " التقى الرئيس الأفغاني مع نظيره الإيراني بعيد قمة الأربعين " in this example the word بعيد should be pronounced بُعَيد not بَعِيد .
I think the best voice for news in Arabic Language is Google Cloud Text-to-Speech WaveNet Type, ar-XA-Wavenet-C voice name, it's the best for reading numbers and ambiguous words and it's more close to the human being.
##
### How to Detect Cliches in Text
- URL: https://mohammadshaker.com/en/blog/how-to-detect-cliches-in-text
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Cliches reduce text informativeness because they carry no new information — the reader already knows what comes next. Detecting them computationally means extracting collocations with high PMI scores, then cross-referencing against known idiom lists. High cliche density is a reliable signal that an article is formulaic rather than informative.
#### Content
In a previous article, we talked about the various factor that makes an article more informative, using cliches was not one of them, this article is a part of our research on measuring text informativeness, if you are interested [jump directly to the gist](/blog/informativity-detection-almetas-research-gist) of our research or review other parts:
- [What makes an article informative in our view and how can we measure this quantitatively?](/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of)
- [How to measure the informativeness of each individual word?](/blog/term-informativeness-estimation-in-the-arabic-language)
- [How to train a supervised model to measure informativeness?](/blog/supervised-article-informativeness-prediction-the-what-and-the-how)
- [Can we get data to train such a model from summaries?](/blog/can-you-measure-a-text-informativeness-using-its-summary) ?
- [Can we get such data in Arabic?](/blog/automatically-tagging-data-for-content-informativity-scoring)
- [How can we rank articles based on how informative they are?](/blog/how-to-rank-articles-based-on-how-informative-they-are-using-snorkel)
You are still here? Cool! Let us start.
Cliches are overused, unoriginal expressions that appear in a context where something more novel might have reasonably been expected. They are predominantly multi-word expressions (MWEs) and are closely related to the idea of formulaic language. They are an aspect of the more general formulaic language which also include:
- idioms: a phrase or expression whose meaning can't be understood from the ordinary meanings of the words in it e.g. “Get off my back!” is an idiom meaning “Stop bothering me!”
- binomials: pair is an expression containing two words which are joined by a conjunction (usually “and” or “or”) e.g. “rock and roll“, “more or less”
- collocations: implicit constrains on the attachment of words. e.g. “Heavy rain” vs “Strong rain”
Overall cliches can be classified into three types based on placement:
- Fixed expressions: “An eye for an eye”
- Semi-fixed expressions: can include small changes in prepositions: “Throw gas (gasoline) on (to) fire”
- Syntactically-flexible: “To rake someone **over the coals**" or "to **haul** someone **over the coals**". Means to **reprimand** them **severely**.
# Should writers “avoid cliches like the plague”?
In [1] the authors collected a large set of dutch (and Dutch translated) novels, used crowdsourcing to rate them based on quality, then studied the relationship between cliches usage and the reported quality by users. The following figure demonstrates their results. With a Pearson correlation coefficient r=-0.32 there is a small linear negative correlation (articles with more cliches tend to have lower ratings), and with a p-value of 5.6e-10 the probability of a null hypothesis is too small (which means, we are confident of the results).
image is taken from [1]
Furthermore, The detection of cliché and more generally formulaic language can go well beyond the simple issue of writing quality to include many other tasks like improving the quality of machine translation, since the ability to handle MWEs in any language-related tasks will have a positive impact on their final output quality.
# How other People Do it? What are clichés like, linguistically?
In [1] the authors studied several indicators to detect cliches
and those included:
- Mean sentence length(number of tokens)
- Common vocabulary: the percentage of tokens part of the 3000 most common words in a large reference corpus
- Direct speech: the percentage of sentences with direct speech punctuation. e.g. 'he said that bla' vs 'He: “bla” '
- Compression ratio the number of bytes when the text is compressed divided by the uncompressed size.
They found out that all of the above correlates significantly with the density of cliches, that basically cliches stand out as having simpler language: they consist almost exclusively of common words, are more repetitive (a lower compression ratio indicates more repetitiveness), and contain shorter sentences.
## Using Words Count
In [2] the authors worked on measuring how much a song is clichéd
as a factor in measuring the quality of it based on its lyrics and
rhythm (the rhythmatic last word in each sentence like say,day,away),
for lyrics they used n-grams and for the rhythm they used a ranked
list of pairs (stay, away) … , they have annotated a very tiny test
set and then suggested several equations of a cliché metric for
songs. Nonetheless, there are several issues with this works:
- The test set is very small and is annotated by a single person thus is subjective and under-representative
- This method isn’t applicable for our goal of determining how cliched an arbitrary text since in this case rhyming is not a typical feature of the texts.
- Repetition in song lyrics motivated their n-grams score, but this is not a salient feature of the texts we consider
In [3] the authors used the frequencies of n-grams to assess whether a text is clichéd or not, they found out that clichéd text tends to have their higher n-grams 3,4,5-grams more frequent the following figure illustrates the difference in n-grams distribution between clichéd and original text. However, similar to [2] the authors used a very small test set to demonstrate their method.
## Using Dictionaries and Lists
In [4] the authors try to detect Portuguese proverbs by collecting a closed list of them om sources like Wikipedia. The authors in [1] as well suggest using a large lexicon of formulaic language.
## Formulaic Language Detection
A part of the effort to detect formulaic language, the authors in [5] built a list of common collocations empirically for Modern Standard Arabic by employing several metrics to rank bi-grams including t-score, log-liklihood, etc.
further effort to detect longer sequences of Formulaic Language is discussed in [6] and [7], in [7] the authors used ranked n-grams in a similar manner to [5] to semi-automatically create a list of common Arabic collocations and long Formulaic sequences.
## Supervised Models
While it seems possible to annotate a large corpus for clichés and train a system to extract them, we didn’t come upon any study that did that.
# How Can We Do it?
On one hand, we can use the words counts, based on the results of [3] it is possible to use the following algorithm:
- create a histogram of n-grams distribution of original and
clichéd text
- for each new text find the histograms of high n-grams
- use the distance between the text histogram and those of
original and clichéd text (using KL divergence for example) to
choose the closest distribution as well as the confidence.
On the other hand it is also possible to scrap/build a closed list
of cliché expressions and detect them in text.
The question of using closed lists vs n-grams distributions is basically a precision-recall trade-off. A manually curated dataset may have limited recall, but will yield higher precision (i.e., will contain fewer false positive). Moreover, the n-gram technique cannot be used to detect whether a particular set of clichés is present in a large text, and the clichés cannot be located; the n-gram method is, therefore, coarse-grained.
| Method | Pros | Cons |
| --- | --- | --- |
| Closed list | Possibility of detecting the actual clichéd text.
High precision. | Needs constant updates.
Low to mid recall based on list accuracy.
Might be hard to acquire based on language. |
| Words counts | Easier to apply. | Depends on corpora size.
Low precision. |
| Semi-automatic list building | Can combine the best of both worlds. | Extremely harder to build.
Has the same shortcomings as a closed list. |
# Any Available resources?
- [This](http://www.clichelist.net/) site contains a comprehensive manually maintained list of common cliché expressions
- [This](https://simple.wikiquote.org/wiki/Arabic_proverbs) is a list of Arabic proverbs from Wikiquote with their translation to English
##
# References:
[1] A. van Cranenburgh, “Cliché Expressions in Literary and Genre
Novels,” in *Proceedings of the Second Joint SIGHUM Workshop on
Computational Linguistics for Cultural Heritage, Social Sciences,
Humanities and Literature*, 2018, pp. 34–43.
[2] A. Smith, C. Zee, and A. Uitdenbogerd, “In your eyes:
Identifying clichés in song lyrics,” in *Australasian Language
Technology Workshop 2012 (ALTW 2012)*, 2012, pp. 88–96.
[3] P. Cook and G. Hirst, “Automatically assessing whether a text
is clichéd, with applications to literary analysis,” in
*Proceedings of the 9th Workshop on Multiword Expressions*,
2013, pp. 52–57.
[4] A. P. Rassi, J. Baptista, and O. Vale, “Automatic detection of
proverbs and their variants,” in *3rd Symposium on Languages,
Applications and Technologies*, 2014.
[5] A. A. O. Alghamdi, E. Atwell, and C. Brierley, “An empirical
study of Arabic formulaic sequence extraction methods,” in
*Proceedings of the LREC 2016*, 2016, pp. 502–506.
[6] A. Alghamdi and E. Atwell, “Towards Comprehensive Computational
Representations of Arabic Multiword Expressions,” in *International
Conference on Computational and Corpus-Based Phraseology*, 2017,
pp. 415–431.
[7] A. Alghamdi and E. Atwell, “Constructing a corpus-informed list
of Arabic formulaic sequences (ArFSs) for language pedagogy and
technology,” *Int. J. Corpus Linguist.*, vol. 24, no. 2, pp.
202–228, 2019.
### How to Detect Clickbait Headlines Using NLP?
- URL: https://mohammadshaker.com/en/blog/how-to-detect-clickbait-headlines-using-nlp
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Clickbait detection in NLP uses a combination of linguistic features: exaggerated sentiment, forward-reference patterns, question-form headlines, and punctuation abuse. We survey the key NLP approaches — from rule-based methods to machine learning classifiers — and explain how each maps to observable headline characteristics.
#### Content
Clickbait is a type of hyperlink on a web page that has catchy or provocative headlines difficult for most users to resist, they tell you exactly what you’re about to see, with just enough of a tease at the end to intrigue you into clicking. Mostly, the content of these pages is disappointed in comparison to their headlines.
The aim of the clickbait detection task is to detect those clickbait pages. A related field of research is identifying bad content on the web, such as spam and fake websites, where features like link-structure and blacklists of URLs, hosts, and IPs have been found to be useful. However, clickbait articles are not spam or fake pages. They can be hosted on reputed news sites.
If you already familiar with the task of click-bait detection and wish to know the details of our system then check out our next article, if you wish to know how clickbait can be formulated as a machine learning task then keep reading.
## Clickbait Detection in NLP
The task is usually treated as a binary classification problem, where the classes are “is clickbait” and “is not clickbait”. It can be treated as a regression problem too, where the target denotes the degree of being a clickbait in %.
We can consider
the task in terms of social media or normal websites, where the
difference is in the available information:
- For a normal website, we have the article and its headline.
- For a social media post in addition to the article and its headline we have the post text and other meta-information related to the post (number of shares, number of comments, number of likes, the time when it was posted, the hashtags, etc...) and to the user (number of followers, number of following, etc…).
## From Where Can We Get Data to Train a Clickbait Detector?
The most popular dataset in English is the Clickbait Challenge dataset [15], some competitors worked on extending it manually, other researchers annotated their own datasets manually depending on websites famous in using the clickbait technique in their marketing strategies like [buzzfeed.com](file:///C:%5CUsers%5CASUS%5CDesktop%5Cbuzzfeed.com), and other official websites like cnn.com where clickbait supposed to be rare.
**To our knowledge, there is no available Arabic clickbait detection dataset.** However, many websites are filled with clickbait headlines and can be considered as good data sources.
## Which Features Refer to Clickbait?
The clickbait detection task is known for its huge feature space, which can be separated into three categories.
### Textual Features
Features related to the headline, the main content in the article itself, the text in the post if we are solving the problem in the social media domain, and the URL of the article.
Following are kinds of textual features:
#### Title-based
- The presence/number of some features in the title: numbers, punctuations (exclamation marks, question marks, etc...), question words, stopwords, etc...
- POS-based: the presence of superlative adverbs and adjectives
- Sentiment-based: presence/number of negative/positive sentiment words, sentiment polarity.
- Lexicon-based: presence/number of specific phrases (click here, exclusive, won’t believe, happens next, don’t want, you know, etc...)
- word2vec-based: the word2vec representation of the headline.
- Words-based: N-gram, TFIDF.
- Char-based: N-gram.
#### Content-based
- Length-based: average words per sentence.
- Words-based: N-gram, TFIDF
#### Similarity-based
The textual similarity between the headline and the content (all the content, first lines, summary, or meta description).
#### Informality & Readability Based
Measuring the informality level, and the reading difficulty of the text using metrics like LIX, RIX, formality measure, etc...
#### Forward-reference Based
The presence/number of four grammatical categories in headlines:
- Demonstratives (this, that, those, these).
- Third-person personal pronouns (he, she, his, her, him).
- The definite article (the).
- Whether the title starts with an adverb.
#### URL-based
Frequencies of the dash, ampersand, upper case letters, comma, period, equal-to sign, percentage sign, plus, tilde, underscore, and the URL depth (no. of forwarding slashes)
### Visual Features
These features are elated to the lead image in the article, or the post image if we are solving the problem in the social media domain, and are presented in the following forms:
- Using pre-trained object recognition models in two ways:
- Semantic features: the presence of an object in the image, we take the name of the class as a feature.
- The embedding of the image taken from the last layer of the model.
- Using an OCR to extract and analyze the text in the image.
### Behavioral and MetaData Features
These features are considered only if we are solving the problem in the social media domain, they are related to the behavior of the user and the metadata in the post, and can be extracted from the following sources:
- The number of the following and the followers in the publisher profile.
- The number of likes, comments, and shares
- The time of publishing
- Hashtags and mentions
### Modeling the Curiosity
A remarkable work to be considered is [3], Where they proposed modeling the curiosity presented in the clickbait headlines in terms of novelty, surprising, and information gap.
#### Novelty
To model the novelty the following steps were taken:
- An LDA topic model was trained on 200 topics from the ABC news headlines dataset.
- A probability distribution over the 200 topics was generated for each headline in clickbait and non-clickbait samples.
- The information distance between the headline topics that the users were exposed to and the topics that were present in clickbait and non-clickbait was calculated.
They found that clickbait is significantly more novel than non-clickbait.
#### **Surprising**
To model the surprising factor the following steps were taken:
- They took words bi-grams from clickbait and non-clickbait headlines.
- Measured the frequency with which these bi-grams occurred in the ABC news headlines corpus.
- Each headline was represented by the frequency of each bigram in it, which was called the surprise frequency vector.
- More frequent the occurrence of a bigram, lesser is the perceived surprise value of it when encountered by a reader.
#### **Information Gap**
Some features that we talked about previously in this report were used to represent the information gap in the headlines like the presence of a question, pronouns, etc...
## Conclusion
In this post, we discussed some of the proposed methods in the literature to detect clickbait pages using NLP, and ML techniques. We considered the problem as a binary classification and presented a bunch of the features used to characterize the clickbait headlines.
Are you intrigued? check how we implemented this task in Almeta in [our next article](/blog/clickbait-detection-using-word2vec-representation).
##
## References
[1] Biyani, Prakhar, Kostas Tsioutsiouliklis, and John Blackmer. "" 8 Amazing Secrets for Getting More Clicks": Detecting Clickbaits in News Streams Using Article Informality." Thirtieth AAAI Conference on Artificial Intelligence. 2016.
[2] Potthast, Martin, et al. "The clickbait challenge 2017: towards a regression model for clickbait strength." arXiv preprint arXiv:1812.10847 (2018).
[3] Venneti, Lasya, and Aniket Alam. "How Curiosity can be modeled for a Clickbait Detector." arXiv preprint arXiv:1806.04212 (2018).
## Further Reading
[1] Geckil, Ayse et al. “A Clickbait Detection Method on News Sites.” 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM) (2018): 932-937.
[2] Papadopoulou, Olga, et al. "A two-level classification approach for detecting clickbait posts using text-based features." arXiv preprint arXiv:1710.08528 (2017).
[3] Rony, Md Main Uddin, Naeemul Hassan, and Mohammad Yousuf. "BaitBuster: A Clickbait Identification Framework." Thirty-Second AAAI Conference on Artificial Intelligence. 2018.
[4] Ha, Yui, et al. "Characterizing clickbaits on instagram." Twelfth International AAAI Conference on Web and Social Media. 2018.
[5] Potthast, Martin, et al. "Clickbait detection." European Conference on Information Retrieval. Springer, Cham, 2016.
[6] Khater, Suhaib R., et al. "Clickbait Detection." Proceedings of the 7th International Conference on Software and Information Engineering. ACM, 2018.
[7] Elyashar, Aviad, Jorge Bendahan, and Rami Puzis. "Detecting Clickbait in Online Social Media: You Won't Believe How We Did It." arXiv preprint arXiv:1710.06699 (2017).
[8] Wiegmann, Matti, et al. "Heuristic Feature Selection for Clickbait Detection." arXiv preprint arXiv:1802.01191 (2018).
[9] Grigorev, Alexey. "Identifying clickbait posts on social media with an ensemble of linear models." arXiv preprint arXiv:1710.00399 (2017).
[10] Cao, Xinyue, and Thai Le. "Machine learning based detection of clickbait posts in social media." arXiv preprint arXiv:1710.01977 (2017).
[11] Chen, Yimin, Niall J. Conroy, and Victoria L. Rubin. "Misleading online content: Recognizing clickbait as false news." Proceedings of the 2015 ACM on Workshop on Multimodal Deception Detection. ACM, 2015.
[12] Chakraborty, Abhijnan, et al. "Stop clickbait: Detecting and preventing clickbaits in online news media." 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM). IEEE, 2016.
## Frequently Asked Questions
### How does NLP detect clickbait?
NLP detects clickbait by analyzing linguistic features such as excessive punctuation, sensational words, forward-referencing pronouns ("this", "what"), and emotional sentiment scores. Machine learning classifiers trained on these features can distinguish clickbait from legitimate headlines with over 90% accuracy.
### What makes a headline clickbait?
Clickbait headlines typically use curiosity gaps, exaggerated claims, emotional triggers, and vague pronouns to entice clicks without delivering proportional value. Common patterns include listicles with superlatives, questions that imply shocking answers, and phrases like "you won't believe."
### Can clickbait detection be automated?
Yes. Automated clickbait detection combines NLP feature extraction with supervised learning. Features include word embeddings, syntactic patterns, sentiment polarity, and reading difficulty scores. Models like SVM, Random Forest, and neural networks achieve strong results on benchmark datasets.
### How to Measure Text Readability?
- URL: https://mohammadshaker.com/en/blog/how-to-measure-text-readability
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, nlp, Human-Written
Text readability formulas like Flesch-Kincaid and Gunning Fog Index estimate difficulty from sentence length and syllable count — but they were designed for English. Arabic readability requires adapted metrics like AARIBase, which weighs character count, average word length, and average sentence length to score Arabic text difficulty.
#### Content
Readability is the ease with which a reader can understand a written text, which accordingly indicates how effectively the text will reach the target audience.
The readability of text depends on its content (the complexity of its vocabulary and syntax), its presentation (such as typographic aspects like font size, line height, and line length), and factors related to the reader himself (his experience, interest, and motivation). However, the most important set of factors that affect readability are the factors that are related to the text itself.
If you are familiar with how readability can be measured for Arabic content you can jump directly to [our next article](/blog/analysis-of-the-readability-metric-results-in-almeta-news-feed) to discover how we at Almeta have implemented this feature.
To discover how we at Almeta have implemented this feature.
## Why Measure The Readability?
Here are several goals of measuring the readability:
- The easier a text is to read and the clearer the ideas it contains, the more it is likely to attract and retain the attention of the reader.
- Readability has been widely used in education in order to write and select the appropriate books and assessments for students’ level.
- It has been widely used in industry for writing manuals and user instructions in a language’s level appropriate for the average end-users.
- Several official agencies require their forms to be written in a manner that meets a specific readability level, in order to better the spread of information among society members, especially among the ones with a lower level of education and limited literacy.
- In medicine, the readability of instructions and other important forms, like consent forms, is considered vital to assure better medical treatments and accountability towards patients and their families.
- Many researchers use text readability for web applications and information retrieval systems, where they give priority for displaying pages that best match a user's reading level.
## How to Measure the Readability?
Judging readability aims to measure the grade level a person must have to read and comprehend a text. The approaches stated in the research can be divided into two types: traditional approaches and data-driven approaches.
### Traditional Approach
The common traditional approach to predict readability is the use of readability formulas, which work by measuring certain features of a text-based on mathematical calculations. We base these readability measures on a handful of factors, that tend to be simple and clear, and can be distinguished as language-dependent (number of syllables in a word, etc...) and language-independent (mean number of words per sentence, the mean number of characters per word, etc...) factors.
*What about Arabic?*
Over 200 mathematical formulas have been published to help to assess the level of text’s readability in different languages, while a limited amount of research has been conducted on the Arabic language.
#### Applying other languages formulas on the Arabic text
The most common formulas that are used to measure the readability of Arabic are:
**El-Heeti** [1] readability formula which refers to a grade level required to comprehend the text. This formula uses the average word length in characters as the only feature.
**Heeti** = (AWL × 4.414) - 13.468
Where:
***AWL***: The average of words lengths in the text.
The formula does not work as a good indicator of Arabic text readability, especially given that Arabic is a highly inflectional and derivational language and word length by itself does not reflect difficulty.
ARI [2] readability formula ARI which refers to the U.S. grade level needed to comprehend the text.
**ARI Grade Level** = (4.71 × ACW) + (0.5 × AWS) - 21.43
Where:
***ACW***: Average number of characters per word.
***AWS***: Average number of words per sentence.
**LIX** [3] readability formula which produces a score that determines the difficulty of a given text according to the following:
**Score**
0 - 24
25 - 34
35 - 44
45 - 54
55 and above
**Meaning**
Very easy
Easy
Standard
Difficult
Very difficult
And is calculated as follows:
**LIX** = W/S + 100 × WD/W
*Where*:
***W***: Number of words in the text.
***S***: Number of sentences in the text.
***WD***: Number of difficult words in the text, where difficult words are defined as words consisting of more than six letters.
These formulas were selected for their simplicity and more importantly because their parameters can be easily applied to Arabic texts. They are also was chosen for the fact that they do not use language-dependent features like the number of syllables in a word.
[4] conducted a test to check the reliability of the previously mentioned formulas to measure the readability of Arabic text. Their test results showed that:
1. Even these readability formulas which don't contain language-dependent features are language-dependent. Therefore, the constants of these formulas should be adjusted appropriately in order to adapt them to the Arabic language.
2. Average sentence length is a good indicator of Arabic text readability.
3. Al-Heeti formula focuses on one factor that is the average word length. This factor is unreliable alone to assess the readability of the Arabic text.
4. LX produced the most acceptable results among these formulas for the Arabic text.
According to the previous reasons a better technique to measure Arabic text readability is needed.
#### Arabic-based Readability Formulas
There are only two readability formulas for Arabic: AARI and Osman.
#### 1. AARI Metric
**AARI** [5] readability formula:
**AARIBase** = (3.28 × NOC) + (1.43 × ACW) + (1.24 × AWS)
*Where*:
***NOC***: Number of characters.
***ACW***: Average character per word.
***AWS***: Average words per sentence.
The authors showed in their experiments that the results obtained applying their formulas overcome the results obtained applying the previously mentioned non-Arabic based formulas.
#### 2. Osman Metric
OSMAN [6] readability formula is an open-source Arabic metric for text readability written in Java. The formula calculation process used Mishkal2 to diacriticize Arabic text, to extract the word syllables. The use of diacriticized texts is problematic because of Arabic texts often are not diacriticized and the process of introducing diacritics if done automatically, can introduce errors. So this is one of the weak points of the formula. The author introduced a set of frequently misspelled letters and referred them as “Faseeh”. Misspelling those letters could result in prosaic Arabic “Rakeek” –weak pronunciation– and therefore affect text readability.
**OSMAN** = 200.791 - (1.015 × A/B) - 24.181 × ( C/A + D/A + G /A + H /A)
*Where*:
***A***: the total number of words.
***B***: the total number of sentences.
***C***: the total number of hard words (words with more than 5 letters).
***D***: the number of syllables per word.
***G***: the number of complex words (words with more than 4 syllables).
***H***: the number of “Faseeh” words (complex words with any of the “Faseeh” letters)
Here too, the author showed in his experiments that the results obtained applying his formulas overcome the results obtained applying the previously mentioned non-Arabic based formulas. However, he didn’t compare his results with the results of the AARI metric.
#### Corpus-Based Arabic Formula
In [7] they applied a corpus-based approach to build an Arabic readability formula. Based on their claim that frequently used words are usually easier than rarely used ones.
They used King Abdulaziz City for Science and Technology Arabic Corpus to characterize a state of the Arabic language. In the corpus, the word with the highest number of frequencies ranked last. The ranking in the corpus is reversed so that the easiest is ranked the first and so on. The difficulty level based on this ranking is taken into consideration when the mean is computed.
#### Traditional Approach Limitations
According to [4], extensive research has shown that the popular readability formulas are not 100% accurate, yet these formulas provide a "rough estimate" of the readability of a text.
### Data-Driven Approach
Machine learning-based methods, treat the problem as a binary classification problem. We need a training set that is annotated with the class to which each example belongs, where the classes usually grade levels. A commonly used data source for Arabic and other languages is GLOSS3 which contains only 230 MSA annotated texts.
## Conclusion
In this post, we discussed the text readability measurement task in NLP, its benefits, and methods to implement. We presented an approach based on pre-defined formulas, and another based on ML.
Hope you have enjoyed this discussion if you are interested you review [our next article](/blog/analysis-of-the-readability-metric-results-in-almeta-news-feed) to see How do these metrics work in real-world with real Arabic Content.
##
## References
[1]. Flesch, Rudolf Franz. “A new readability yardstick.” The Journal of applied psychology 32 3 (1948): 221-33.
[2]. Smith, E. Anthony and R. J. Senter. “Automated readability index.” AMRL-TR. Aerospace Medical Research Laboratories (1967): 1-14.
[3]. Kootstra, G. "Project on exploratory Factor Analysis applied to foreign language learning." Accessed via: http://www. let. rug. nl/~ nerbonne/teach/remastats-meth-seminar/Factor-Analysis-Kootstra-04. PDF. 2004.
[4]. Al-Ajlan, Amani A. et al. “Towards the development of an automatic readability measurements for arabic language.” 2008 Third International Conference on Digital Information Management (2008): 506-511.
[5]. Tamimi, Abdel Karim Al et al. “AARI: automatic arabic readability index.” Int. Arab J. Inf. Technol. 11 (2014): 370-378.
[6]. El-Haj, Mahmoud, and Paul Rayson. "OSMAN―A Novel Arabic Readability Metric." Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016). 2016.
[7]. Daud, Nuraihan Mat, Haslina Hassan, and Normaziah Abdul Aziz. "A corpus-based readability formula for estimate of Arabic texts reading difficulty." World Applied Sciences Journal 21 (2013): 168-173.
## Further Reading
[1]. Nassiri, Naoual, Abdelhak Lakhouaja, and Violetta Cavalli-Sforza. "Modern standard arabic readability prediction." International Conference on Arabic Language Processing. Springer, Cham, 2017.
## Frequently Asked Questions
### What is the best readability formula?
No single formula is best for all contexts. Flesch-Kincaid is widely used for general English text, SMOG is preferred for health communications, and Gunning Fog works well for business writing. Modern NLP approaches using sentence embeddings outperform traditional formulas on diverse text types.
### What reading level should web content be?
Most web content should target a 6th-8th grade reading level (Flesch-Kincaid Grade Level 6-8). This ensures accessibility for the broadest audience. News articles typically target grade 8-10, while academic papers may reach grade 16+.
### How do readability formulas work?
Readability formulas calculate text difficulty based on measurable features like average sentence length, syllable count per word, and word frequency. More complex words and longer sentences produce higher difficulty scores. Each formula weights these factors differently.
### How to Rank Articles Based on How Informative They Are - Using Snorkel
- URL: https://mohammadshaker.com/en/blog/how-to-rank-articles-based-on-how-informative-they-are-using-snorkel
- Date: 2020-01-18T00:00:00.000Z
- Tags: ml, nlp, Human-Written
Ranking articles by informativeness is easier than scoring them. Humans naturally excel at pairwise comparisons, and AI models mirror that advantage. Snorkel's weak supervision framework lets you build an informativeness ranker without expensive manual labels by turning human heuristics into labeling functions.
#### Content
Let's start with a simple question, what constitutes an informative article? based on Oxford's dictionary.
informative/ɪnˈfɔːmətɪv/ *adjective*: **informative**
providing useful or interesting information
However, this is still an abstract concept. Yes, it is much simpler to flag an article as spammy or unprofessional to see our previous article [here](/blog/automatically-extracting-valuable-content-from-news-streams)
but when it comes to normal article the matter becomes more complicated.
imagine you are given 10 articles and asked to rank each on a scale from 1 to 10 based on how "informative" they are how long will it take you? an hour? half?
If you ever went through such an experience you will remember struggling with exact numbers, yes you will know when an article is zero or 10 but what about deciding between 7 and 6.
Let's try another task given 5 pairs of articles, for each pair you should choose the article that seems more informative to you. it is clear that the second task is much simpler.
This is the difference between the task of regression the former and ranking the latter. We as humans are much suitable for the latter than the former, it is much easier to choose between chocolate and vanilla ice cream than to rank chocolate on a scale of 10.
The same analogy can be carried out to the AI models. mainly since building datasets by human annotators for pairwise ranking is simpler than direct mapping.
Let us assume that you have built a system that given 2 articles that can decide which of them is more informative. In this article, we will explore a single question, who can we use this system to build an informative feed (sort the articles of your feed based on their informativeness).
This article is the latest part of our research on informativeness, and we have A LOT of it if you are interested in more details you can check each individual piece to learn, or [review the gist of it.](/blog/informativity-detection-almetas-research-gist)
- [What makes an article informative in our view and how can we measure this quantitatively?](/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of)
- [How to detect cliches in text? and why do they matter?](/blog/how-to-detect-cliches-in-text)
- [How to measure the informativeness of each individual word?](/blog/term-informativeness-estimation-in-the-arabic-language)
- [How to train a supervised model to measure informativeness?](/blog/supervised-article-informativeness-prediction-the-what-and-the-how)
- [Can we get data to train such a model from summaries?](/blog/can-you-measure-a-text-informativeness-using-its-summary) ?
- [Can we get such data in Arabic?](/blog/automatically-tagging-data-for-content-informativity-scoring)
## Offline Ranking
### What is it?
Let us start with a simpler task, assuming that articles represent a closed set, let us assume we have N articles and for these articles, we have M pairwise comparisons, how can we sort the N elements:
One option is to use ordinary sorting algorithms like quickSort so if we want to sort the articles using pairwise comparisons we will need M=Nlog(N), apart from needing a lot of comparisons the main problem with this method is that it assumes the comparisons to be consistent, meaning that we assume that:
If $latex a > b, b > c$ then $latex a > c$
where $latex a, b, c$ are 3 elements to be compared and > represents our ranking system.
However, this constraint of transitivity might not be necessarily correct, due to noise, errors in the model, etc and if this constraint is not satisfied the whole logic of the sorting algorithm fails.
A method to tackle this issue is to use the Bradley–Terry model in this very simple model that states given a pair of individuals *i* and *j* the probability that the $latex i > j$ is:
$latex P(i>j)={\frac {p\_{i}}{p\_{i}+p\_{j}}}$
Where *pi* is a positive real score assigned to the individual *i*. The comparison *i* > *j* can be read as "*i* is preferred to *j*", "*i* ranks higher than *j*", or "*i* beats *j*", depending on the application.
This model can be parameterized as follows:

Meaning that given a limited number of comparisons M we can fit the Bradley–Terry using these comparisons to estimate each of $latex beta{i}$ for i from 1 to N which, in turn, represents the strength (score) of the element i, and thus we can sort the elements based on this score
However, there is still an issue, in the start we assumed that our set of news articles is closed, however in our real application, the news articles will be added periodically, in that case, the abovementioned models can't assign a score to these new articles, and will require refitting the model every time we add a new article which is computationally very expensive this leads us to the next section.
### How can I implement it?
There are a lot of ready toolkits to be used: use the very simple [elo](https://pypi.org/project/elo/) in case of trasitive comparisons or employ [choix](https://pypi.org/project/choix/) if you need to use models like Bradley–Terry
## Online Ranking
### What is it?
The goal here is to be able to use the pairwise comparison to get a score of each individual article, several methods were introduced in the literature to handle this issue but let us talk about a single line of research RankNet [1] then LambdaRank [2] a full simple mathematical introduction to these algorithms is presented in [4] and this amazing [blogpost](https://mlexplained.com/2019/05/27/learning-to-rank-explained-with-code/) also a good read, here we will try to provide a simplified version of the thing.
#### RankNet:
Let us start by defining an ML model that can take an article represented as an input vector and return a score representing how "Good" this article is, the underlying model can be any model for which the output of the model is a differentiable function of the model parameters RankNet training works as follows. The training data is partitioned by a query. The query here means a dichotomy of the data for example in the case of our application in means comparing articles that talk about the same event. At a given point during training, RankNet maps an input feature vector $latex X from R^n $ onto a number f(x) where n represents the size of the feature vector and f(x) the score of the article. For a given pairwise comparison between elements i and j that belongs to the same query we can calculate the scores $latex s\_i = f(x\_i) and s\_j = f(x\_j)$ RankNet models the probability of i being better than j as :
$latex P(\textrm{rank}(i) > \textrm{rank}(j)) = \frac{1}{1+e^{-(s\_i-s\_j)}} $
Based on this modeling we can apply negative log-likelihood to create a cost function that can be optimized using gradient descent, Note that this optimization process will affect the weights of the original ML model that we are extracting the scores from. After this optimization process is finished this model can be deployed as it cant take a new article and assign a new score to it.
#### LambdaRank:
The main problem with this model is that in our equation we are only using the difference $latex s\_i - s\_j$ to optimize the ranking, while this would work in many cases, most of the time we are more interested in the top results than in the bottom ones. In other words, the loss is the same for any pair of items i , j regardless of whether i and j are ranked in 1st and 10th place or if they are in 100th and 110th place. since the difference will be the same and therefore the modification of the weights will also be the same.
The algorithm of LambdaRank adds a small detail to handle this issue basically we will give more importance to higher ranks than to lower ones by modifying the ranking probability to be:
$latex P(\textrm{rank}(i) > \textrm{rank}(j)) = \frac{-|\Delta(i,j)|}{1+e^{s\_i-s\_j}} $
where $latex -|\Delta(i,j)|$ represents a penalty related to ranks of i and j, several metrics have been proposed to implement delta most famous of them is NDCG (normalized discounted cumulative gain). NDCG is a metric that is commonly used in information retrieval and measures the quality of a ranking. the details of NDCG is out of the scope know and you can read more about it in [here](http://fastml.com/evaluating-recommender-systems/).
### How to do it?
It is possible and may be beneficial to implement the algorithm using your favorite ML toolkit, for example, we tried to implement it in [keras](http://keras.io) by defining a specific loss function (LambdaRank loss), However, the best off the shelf solution we found is [LGBMRanker](https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html#lightgbm.LGBMRanker) from Microsoft's [LightGBM](https://github.com/microsoft/LightGBM)
#### Is this the only way to do it?
Absolutely not. Several methods have been devised to generate full rankings from a limited amount of pairwise comparisons, we have only explored a single line of research other methods includes ranks embedding [5] others based on modified heuristics like [6]. [7] represents a good initial overview of the field.
## Few Pairwise Comparisons
Based on our discussion so far the main thing that is missing to build your amazing feed ranking system is to have enough pairwise comparisons between items to be used as training data for LambdaRank or any other algorithm. **While later on in the life of the system it would be easy to harness users feedback and their interaction to enhance this dataset it is necessary to have an initial data to build the model, and this means manual annotation :(**
### Enter Snorkel
Snorkel is a system for *programmatically* building and managing training datasets **without manual labeling**.
In Snorkel, users can develop large training datasets in hours or days rather than hand-labeling them over weeks or months.
Snorkel currently exposes three key programmatic operations:
- **Labeling data**: e.g., using heuristic rules or distant supervision techniques
- **Transforming data**: e.g., rotating or stretching images to perform data augmentation
- **Slicing data**: into different critical subsets for monitoring or targeted improvement
Snorkel then automatically models, cleans and integrates the resulting training data using novel, theoretically-grounded techniques. Don't believe me? read the paper [8].
Snorkel is such a big deal that we are planning to create a full article dedicated just for it, but if you are out of patience, visit their amazing [site](https://www.snorkel.org/) and jump into the tutorials. For now, let us focus on how we can use snorkel for our ranking system. And for now, we are only interested in **labeling the data.**
#### Labeling Functions
Snorkel is based on the idea that we can create several Labeling functions and then combine them to programmatically annotate the dataset, these Labeling functions can be seen as simple rules that vote on the label of the data point. For example, in our task the input to a labeling function would be 2 articles, A and B, and the output will be one of 3 labels -1 (A>B), +1 (A
#### What can be used as a labeling function?
The short answer is anything. The long answer is that because their only requirement is that they map a data point a label (or abstain), they can wrap a wide variety of forms of supervision. Examples include, but are not limited to:
- *Keyword searches*: looking for specific words in a sentence
- *Pattern matching*: looking for specific syntactical patterns
- *Third-party models*: using a pre-trained model (usually a model for a different task than the one at hand)
- Distant supervision: using an external knowledge base
- Crowdworker labels: treating each crowd worker as a black-box function that assigns labels to subsets of the data
More importantly, these LFs can be Noisy (not always accurate) and even have conflicts and correlations between them and it is the role of snorkel to combine them and build a probabilistically annotated dataset.
#### Are all labeling functions the same?
Obviously, no. Assume that for our task of articles comparison we used 3 LFs all of them takes 2 articles as input and should output a label in the set {-1,0, 1}:
- LF1: ranks the 2 articles based on the number of cliche terms in each of them.
- LF2: ranks the 2 articles based on the readability of the 2 articles.
- LF3: ranks the 2 articles based on the ratio of named entities in each of them. (to measure how detailed the article is)
It is easy to see that the 3 LFs will have different behavior and might as well have different accuracy when used to label data, but how can we evaluate them and hopefully improve upon them. based on snorkel: the typical LF development cycles include multiple iterations of ideation, refining, evaluation, and debugging. A typical cycle consists of the following steps:
1. Look at examples to generate ideas for LFs
2. Write an initial version of an LF
3. Spot check its performance by looking at its output on data points in the training set (or development set if available)
4. Refine and debug to improve coverage or accuracy as necessary
And here is the catch, while theoretically, it is possible to use Snorkel to build your dataset with 0 labeled data in practice it is important to have at least a very small development dataset this can be in the magnitude of 100 data points. the development data will be used to help you build and evaluate the performance of the LFs. However, annotating such a small number of data points by hand is not a big deal and can be accomplished in hours to days.
## To Wrap it All Up
OK, so here is the simple plan to build your amazing online ranking system:
1. Start by collecting a large amount of unlabeled data in form of pairs. Let us call this set Training Set
2. Manually annotate at least 2 small sets of data one for developing Labeling Functions with Snorkel (this can be around 100 data points) and a larger test data to evaluate the overall performance of the pairwise ranker
3. Develop LFs using your development set and imagination then use them with Snorkel to automatically annotate your Training Set
4. Use any of the aforementioned online ranking algorithms to train a ranking model
5. Enjoy.
In this article, we simply explained the workflow of creating a feed ranking system in the next articles we will go with you through the implementation details of every step of this plan.
##
## Further Reading
[1] Burges, Christopher, et al. "Learning to rank using gradient descent." *Proceedings of the 22nd International Conference on Machine learning (ICML-05)*. 2005.
[2] Burges, Christopher J., Robert Ragno, and Quoc V. Le. "Learning to rank with nonsmooth cost functions." *Advances in neural information processing systems*. 2007.
[3] Wu, Qiang, et al. "Adapting boosting for information retrieval measures." *Information Retrieval* 13.3 (2010): 254-270.
[4] Burges, Christopher JC. "From ranknet to lambdarank to lambdamart: An overview." *Learning* 11.23-581 (2010): 81.
[5] Jamieson, Kevin G., and Robert Nowak. "Active ranking using pairwise comparisons." *Advances in Neural Information Processing Systems*. 2011.
[6] Shah, Nihar B., and Martin J. Wainwright. "Simple, robust and optimal ranking from pairwise comparisons." *The Journal of Machine Learning Research* 18.1 (2017): 7246-7283.
[7] Jamieson, Kevin G., and Robert Nowak. "Active ranking using pairwise comparisons." *Advances in Neural Information Processing Systems*. 2011.
[8] Ratner, Alexander J., et al. "Data programming: Creating large training sets, quickly." *Advances in neural information processing systems*. 2016.
## Related Questions
### Q: What is Snorkel and how does it help with data labeling?
Snorkel is a framework for programmatically building training datasets without manual labeling. It allows users to write labeling functions—simple rules or heuristics that vote on data labels—and then combines them using a noise-aware model to produce probabilistically annotated datasets in hours instead of weeks.
### Q: How does pairwise ranking differ from regression for article quality?
Pairwise ranking asks which of two articles is better, while regression assigns an absolute score to each article. Humans are significantly better at making comparative judgments than assigning precise numerical scores, which makes pairwise ranking annotations more reliable and faster to produce.
### Q: What is the Bradley-Terry model used for in ranking?
The Bradley-Terry model estimates the relative strength of items from pairwise comparison data. Given a limited number of comparisons, it fits probability parameters for each item, producing scores that can be used to create a full ranking even when comparisons are noisy or intransitive.
### Q: How do RankNet and LambdaRank work for online article ranking?
RankNet trains a neural network to output article scores by optimizing pairwise comparison probabilities using gradient descent. LambdaRank improves on RankNet by weighting errors at higher ranks more heavily using NDCG, ensuring that mistakes among top-ranked articles are penalized more than mistakes among lower-ranked ones.
### Informativity Detection - Almeta's Research Gist
- URL: https://mohammadshaker.com/en/blog/informativity-detection-almetas-research-gist
- Date: 2020-01-18T00:00:00.000Z
- Tags: almeta.io, arabic, nlp, Human-Written
Measuring article informativeness requires breaking an abstract concept into quantifiable features: readability, cliche density, term-level informativeness, and skimmability. Almeta's approach treats informativeness as a supervised learning problem — training a model on proxy-labeled data from human summaries to rank Arabic news articles by quality.
#### Content
Let's start with a simple question, what constitutes an informative article? based on Oxford's dictionary.
informative/ɪnˈfɔːmətɪv/ *adjective*: **informative**
providing useful or interesting information
However, this is still an abstract concept. The question of measuring How informative a piece of news is not really a simple one. And we at Almeta have been debating this issue and testing multiple things for quite some time, in this article we will provide an overview of our plan to implement a service that can measure an article informativeness and provide you with the best possible news feed.
If you don't feel like reading a lot you can jump directly to the final paragraph for a gist of the gist.
This article represents a gist of all of our research on informativeness, and we have A LOT of it if you are interested in more details you can check each individual piece to learn:
- [What makes an article informative in our view and how can we measure this quantitatively?](/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of)
- [How to detect cliches in text? and why do they matter?](/blog/how-to-detect-cliches-in-text)
- [How to measure the informativeness of each individual word?](/blog/term-informativeness-estimation-in-the-arabic-language)
- [How to train a supervised model to measure informativeness?](/blog/supervised-article-informativeness-prediction-the-what-and-the-how)
- [Can we get data to train such a model from summaries?](/blog/can-you-measure-a-text-informativeness-using-its-summary) ?
- [Can we get such data in Arabic?](/blog/automatically-tagging-data-for-content-informativity-scoring)
- [How can we rank articles based on how informative they are?](/blog/how-to-rank-articles-based-on-how-informative-they-are-using-snorkel)
## Properties of the Informative Text
One way to think about informativeness is to see it as a collection of features that comes together in a news piece and determines it's worth.
Imagine a beautiful building, can we define what makes it beautiful? it is really hard to pin down, it might be the structure, the distribution of light, or maybe it is the selection of colours. All of these aspects are simpler and hopefully can be quantitatively measured.
The same analogy can be taken when dealing with text, here the readability, the skimability or the usage of cliches can constitute measurable features, in our previous [article](/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of) we tried to explore what type of features will mark an article as informative, but at the same time be measurable by our systems.
## A student-teacher approach
While the above-mentioned reasoning might seem sound, let us explore a different assumption, here we are assuming that the informativity is an abstract concept that is hard to pin down exactly and that we can measure it using common sense just like the coherence of a text or a musical composition.
These common-sense rules must have been constructed in our mind through years of experience. If this assumption is valid wouldn't it be possible to transfer this experience to information systems? In this article, we are exploring exactly this possibility.
In AI terminology this means that we should deal with the problem of informativeness measurement as a learning problem, basically train a model in a supervised manner to take a piece of text and spit out a number that represents how informative an article is. Obviously this system can utilize the features from [here.](/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of)
The bottleneck in such a scheme is the ability to create a large amount of data to train such a model, we have explored this issue in details. [How we can do it and](/blog/supervised-article-informativeness-prediction-the-what-and-the-how) how can we harness data for this task ([here](/blog/can-you-measure-a-text-informativeness-using-its-summary) and [here](/blog/automatically-tagging-data-for-content-informativity-scoring)) but overall the main issue when dealing with this problem is our failure in harnessing data for Arabic articles.
## Humans Point of View
The failure in the supervised informativity detection made us rethink our steps, is this task really what we think it is.
Any task that we should try to automate we must firstly try to understand. it is not possible to ask an information system to measure the worth of an article if we don't know how humans do that.
Yes, it is much simpler to flag an article as spammy or unprofessional, in our previous article [here](/blog/automatically-extracting-valuable-content-from-news-streams) we have explored those limits that make an article totally not worthy of your time, but apart from the obvious what makes an article more informative and more noteworthy than others.
Imagine you are given 10 articles and asked to rank each on a scale from 1 to 10 based on how "informative" they are how long will it take you? an hour? half?
Let's try another task given 5 pairs of articles, for each pair you should choose the article that seems more informative to you. it is clear that the second task is much simpler.
This is the difference between the task of regression the former and ranking the latter. We as humans are much suitable for the latter than the former. It is much easier to choose between chocolate and vanilla ice cream than to rank chocolate on a scale of 10. The same analogy can be carried out to the AI models. Mainly since building datasets by human annotators for pairwise ranking is simpler than direct mapping.
While this seems like a sensible proposition, it is important for it to be applicable, this is what we explored in our previous [article](/blog/how-to-rank-articles-based-on-how-informative-they-are-using-snorkel) and how we can hopefully deal with the lack of data using [Snorkel](https://www.snorkel.org/).
## The Road Ahead
Overall our current plan for building an informativity detection system goes as follows:
- Treat the task as a ranking rather than a regression problem
- Build a pairwise ranker of articles as follows:
- create a small manually annotated evaluation and development dataset
- utilize snorkel to expand this dataset
- train a pairwise ranker like LambdaRank on the created dataset
- Reuse the pairwise ranker to give a score for each of the fetched articles effectively converting the model from a ranking model into a regression model.
This is what we will do.
##
### Political Orientation Detection - AI and NLP Approach
- URL: https://mohammadshaker.com/en/blog/political-orientation-detection-ai-and-nlp-approach
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, nlp, Human-Written
Political orientation detection automatically classifies news articles along ideological dimensions using NLP — left, right, or center. The main approaches use stance classification, topic modeling, and framing analysis. Arabic media presents additional challenges because political alignment often maps to geopolitical affiliation rather than a simple left-right spectrum.
#### Content
While some news anchors try to stay professional and subjective in all of their articles, most of the news we consume are published to push a specific agenda especially when it comes to politics.
In our effort to battle news bias, it is crucial for our system at Almeta to understand this hidden Agenda, in this article we will explore one path to do this.
This article is a part of our series on political bias detection we will hopefully introduce you to the various aspects of our political bias detection system, and you can learn about:
- [How can we predict the political orientation behind a piece of the news?](/blog/political-orientation-detection-ai-and-nlp-approach)
- [What is Stance detection? and what are the different types of it?](/blog/stance-detection-state-of-the-art)
- [What is subjective stance detection? what is distance supervision? and why they make a cute couple?](/blog/subjective-stance-detection-what-is-it-and-how-to-do-it)
- [What is ALSA, ELSA and how can opinion mining save us from political bias?](/blog/aspect-level-vs-entity-level-sentiment-analysis)
- [How to implement an initial political bias detector just from sentiment analysis and some probabilistic distribution? (warning cool visualizations)](/blog/from-sentiment-to-political-bias-in-the-arab-world-and-the-arabic-content)
- How to Visualize a Political Bias Metric
## The What
Let us start with a very crucial assumption, we assume that **if an article is extremely in favor (or against) a certain entity or idea then this article is biased.** This is the simplest axiom we can start our discussion from.
Although this task does not strictly fit within the boundaries of the opinion mining family it fits the assumption we had, here the entity is the political ideology, in this task, given an article, the system should predict the political orientation of that article (in favor of which ideology is it).
In western politics, the political spectrum is more uniform with a clear dichotomy between the left and right, in that case, the system boils down to a binary classifier. However, in Arabic politics, the vision is more blurred with various orientations.
## The How (The Data)
For the western politics the congressional and parliament debates make for a great source of data since by simply knowing the name of the representative it is possible to know his/her political orientation. For example (democrats usually hold liberate left ideologies while republicans have a conservative right ideology) this dichotomy helps creating large (and fairly clean) datasets automatically, see [1] for an example.
However, this type of data collection does not work in Arabic politics since no such dichotomy exists. Since the role of the party is rather negligible and no such consistency is found.
The only work we found working on this in Arabic was [2] here the authors consider only 5 types of political ideology present in the Arabic politics and scraped data for them from news sites that hold the same ideology, the authors focused on articles that talked about a certain news in order to consider the way the same news is seen by the different ideological sites, while discarding any articles that directly talks about the ideology and its history. The extracted data is [here](https://github.com/malayyoub/Arabic-Articles-Political-Orientation) the reported size of the data set is fine yet we didn’t assess its quality, and in case more data is needed, applying the method outlined in the paper is straight forward (either by using the same sites or similar ones), the following table shows the ideologies considered in the dataset and the source sites chosen by the authors.
| Ideology | Sites |
| --- | --- |
| Arab Nationalism | |
| Islamic Brotherhood | [This](https://www.ikhwanwiki.com/index.php?title=%D8%A7%D9%84%D8%B5%D9%81%D8%AD%D8%A9_%D8%A7%D9%84%D8%B1%D8%A6%D9%8A%D8%B3%D9%8A%D8%A9) |
| Islamic Shia | |
| Liberals | |
| Socialists | |
## The How (Code)
The task boils down to simple text classification, and while the authors in [2] used feature engineering, a simple neural model relaying on words embeddings such as Facebook’s [fasttext](https://pypi.org/project/fasttext/) can give better results. And the SOTA is pretty high.
## Conclusion
Hopefully by now you should have an essential Idea of what is the political orientation and how it is possible to distinguish between different orientations of different text pieces. and as always don't forget to check the resources.
##
## References
[1] M. Iyyer, P. Enns, J. Boyd-Graber, and P. Resnik, “Political
ideology detection using recursive neural networks,” in *Proceedings
of the 52nd Annual Meeting of the Association for Computational
Linguistics (Volume 1: Long Papers)*, 2014, vol. 1, pp. 1113–1122.
[2] R. Abooraig, S. Al-Zu’bi, T. Kanan, B. Hawashin, M. Al Ayoub,
and I. Hmeidi, “Automatic categorization of Arabic articles based
on their political orientation,” *Digit. Investig.*, vol. 25,
pp. 24–41, 2018.
### Search Service Frameworks Evaluation
- URL: https://mohammadshaker.com/en/blog/search-service-frameworks-evaluation
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, nlp, Human-Written
Choosing between Lucene, Elasticsearch, and Solr for an Arabic NLP application depends on more than raw indexing speed. Arabic language support — tokenization, diacritization, stemming — is often the deciding factor. We evaluate all three frameworks on indexing performance, query capabilities, Arabic-language handling, and horizontal scalability.
#### Content
Search engines use indexing to store information about web pages, enabling them to quickly return relevant, high-quality results.
Indexing is the process by which search engines organize information before a search to enable super-fast responses to queries.
Searching through individual pages for keywords and topics would be a very slow process for search engines to identify relevant information. Instead, search engines (including Google) use an inverted index, also known as a [reverse index](https://allotment.digital/learn/technical-seo/how-search-engines-work/indexing/).
The following libraries and engines are services for search, we will discuss each individually, then we'll make a comparison between them based on several factors:
### **1 - Lucene**
Apache Lucene is a free and open-source search engine software library, originally written completely in Java.
It is supported by the Apache Software Foundation and is released under the Apache Software License.
Lucene has been ported to other programming languages including Object Pascal, Perl, C#, C++, Python, Ruby and PHP.
### 2 - **Solr**
Solr (pronounced "solar") is an open-source enterprise-search platform, written in Java, from the Apache Lucene project. It uses the Lucene Java search library at its core for full-text indexing and search, and has REST-like HTTP/XML and JSON APIs that make it usable from most popular programming languages.
### **3 - Elasticsearch**
Elasticsearch is a search engine based on the Lucene library. It provides a distributed, multitenant-capable full-text search engine with an HTTP web interface and schema-free JSON documents. Elasticsearch is developed in Java. Following an open-core business model, parts of the software are licensed under various open-source licenses (mostly the Apache License), while other parts fall under the proprietary (source-available) Elastic License. Official clients are available in Java, .NET (C#), PHP, Python, Apache Groovy, Ruby and many other languages. According to the DB-Engines ranking, Elasticsearch is the most popular enterprise search engine followed by Apache Solr, also based on Lucene.
### 4 - **Sphinx**
Sphinx can be used either as a stand-alone server or as a storage engine ("SphinxSE") for the MySQL family of databases. When run as a standalone server Sphinx operates similar to a DBMS and can communicate with MySQL, MariaDB and PostgreSQL through their native protocols or with any ODBC-compliant DBMS via ODBC. MariaDB, a fork of MySQL, is distributed with SphinxSE.
If Sphinx is run as a stand-alone server, it is possible to use SphinxAPI to connect an application to it. Official implementations of the API are available for PHP, Java, Perl, Ruby and Python languages. Unofficial implementations for other languages, as well as various third party plugins and modules are also available. Other data sources can be indexed via pipe in a custom XML format.
### 5 - Amazon CloudSearch
Amazon CloudSearch is a scalable cloud-based search service that forms part of Amazon Web Services (AWS). CloudSearch is typically used to integrate customized search capabilities into other applications. According to Amazon, developers can set a search application up and deploy it fully in less than an hour.
### 6 - Amazon **Elasticsearch** Service (Amazon ES)
With Amazon Elasticsearch Service, you pay only for what you use. There is no minimum fee or usage requirement. You are charged only for Amazon Elasticsearch Service instance hours, Amazon EBS storage (if you choose this option), and data transfer.
This table shows the comparison between these Frameworks:
| | **Lucene** | **Solr** | **Elasticsearch** | **Sphinx** | **Amazon CloudSearch** | **Amazon Elasticsearch** |
| --- | --- | --- | --- | --- | --- | --- |
| Autocomplete | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| Auto-suggestion | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| Recommendation | ✔️ | ✔️ | ✔️ | - | ✔️ | ✔️ |
| Support Arabic | ✔️ | ✔️ | ✔️ | - | ✔️ | ✔️ |
| Memory size per million document | 125% - 150% of docs size | 125% - 150% of docs size | 125% - 150% of docs size | - | 125% - 150% of docs size | 125% - 150% of docs size |
| Disk size | size of docs \* 2 | size of docs \* 2 | size of docs \* 2 | size of docs \* 2 | - | - |
| Cost | Free | Free | Standard-16$/month | Free | [Pay only for what you use](https://aws.amazon.com/cloudsearch/pricing/) | [Pay only for what you use](https://www.amazonaws.cn/en/elasticsearch-service/pricing/) |
*The response time dependence on the hardware you use.*
#### Why Elasticsearch is paid?
Actually, the code of elastic is open-source, so if you want managed hosting from elastic.co, they charge you according to several variables. You can find the pricing [here](https://www.elastic.co/cloud/elasticsearch-service/pricing).
If you want to use the open-source version, stand up your own servers and manage your own deployment, the code is at no cost and can be found [here](https://github.com/elastic/elasticsearch).
#### AWS ELASTICSEARCH VS AWS CLOUDSEARCH
There is some useful comparison [here](https://optimalbi.com/blog/2016/02/16/aws-elasticsearch-vs-aws-cloudsearch/).
Useful comparison: [Amazon CloudSearch vs ElasticSearch vs Apache Solr Comparison in detail](http://Amazon CloudSearch vs ElasticSearch vs Apache Solr Comparison in detail).
## Conclusion
The purpose of storing an index is to optimize speed and performance in finding relevant documents for a search query. Without an index, the search engine would scan every document in the corpus, which would require a considerable time and computing power.
According to my search, I think that Amazon CloudSearch, Amazon Elasticsearch and maybe Elasticsearch paid-version.
##
### Stance Detection - State of the Art
- URL: https://mohammadshaker.com/en/blog/stance-detection-state-of-the-art
- Date: 2020-01-18T00:00:00.000Z
- Tags: almeta.io, ml, nlp, Human-Written
Stance detection classifies a text's position toward a specific target — agree, disagree, or neutral — and splits into two branches: objective stance for fact verification and subjective stance for opinion classification. State-of-the-art methods range from feature-engineered classifiers to attention-based deep learning models fine-tuned on task-specific datasets.
#### Content
This article is a part of our series on political bias detection we will hopefully introduce you to the various aspects of our political bias detection system, and you can learn about:
- [How can we predict the political orientation behind a piece of the news?](/blog/political-orientation-detection-ai-and-nlp-approach)
- [What is Stance detection? and what are the different types of it?](/blog/stance-detection-state-of-the-art)
- [What is subjective stance detection? what is distance supervision? and why they make a cute couple?](/blog/subjective-stance-detection-what-is-it-and-how-to-do-it)
- [What is ALSA, ELSA and how can opinion mining save us from political bias?](/blog/aspect-level-vs-entity-level-sentiment-analysis)
- [How to implement an initial political bias detector just from sentiment analysis and some probabilistic distribution? (warning cool visualizations)](/blog/from-sentiment-to-political-bias-in-the-arab-world-and-the-arabic-content)
- How to Visualize a Political Bias Metric
Stance detection is an important task in NLP concerned with evaluating the text stance towards a specific predefined target.
The task of Stance *Detection* and Stance *Analysis* can be used as a broad term to cover many sub-tasks there are 2 most dominant types of stance detection:
## Objective Stance Detection
In this type given a piece of text (usually an news article but
can be smaller such as a comment in a community thread) and a claim
(usually the claim would be a full sentence such as a news headline)
the task is to return on of the following tags:
- Disagree: the article disagrees with the given claim
- Agree: the article agrees with the given claim
- Discuss: the article is related to the claim yet does not
impose any sentiment towards the claim’s validity
- Unrelated: this piece of text is totally unrelated to the
claim
this type is usually used for fact-checking claims to find trust-worthy resources that either agrees or disagrees with the claim. Furthermore, in this type the articles are usually objective i.e. the article does not have an opinion regarding the claim, it simply states i it is true or false. Finally, this task is very much related to fact checking
This type of stance detection was the focus of [Fake
news challenge](http://www.fakenewschallenge.org/), [subtask
A of semeval 2019 task 7](file:///home/simn/Downloads/factmata/stance%20detection/SemEval%202019%20Task%207), [task
8 of semeval 2017](http://alt.qcri.org/semeval2017/task8/) and is related to [subtask
A of semeval 2019 task 8](https://competitions.codalab.org/competitions/20022).
[Emergent](http://www.emergent.info/) is an online system that captures viral claims and collects articles that are relevant to it, then verifies them (I am not quit sure weather the verification process is done automatically or not) yet it mostly relays on manual sites such as [Snops](https://www.snopes.com/) .
See [1], [2] and [3]from the FNC, [4] is related to click bait,
and most importantly [this
unpuplished paper](ftp://download.hrz.tu-darmstadt.de/pub/FB20/Dekanat/Publikationen/UKP/2018_NAACL_AnH_AvP_BeS_FeC_DeC_IG_submission.pdf) in which the authors use both the FNC data and
also utilizes data from the less related [semeval
2018 task 12](https://competitions.codalab.org/competitions/17327).
Regarding Arabic: the most promising work was done through the “Check that!” labs at CLEF 2018 and 2019, [5] is the 2018 version paper and [here](https://groups.csail.mit.edu/sls/downloads/factchecking/downloads.cgi) is the corpus. [This](http://alt.qcri.org/clef2019-checkthat/index.php?id=overview) is the 2019 version and it is still an ongoing challenge thus any work we do might be submitted as a paper, [6] is the 2019 version overview paper. I really like the proposed pipeline as it is very suitable to both fact-checking and stance detection.
## Subjective Stance Detection
In this type given a shorter piece of text representing an opinion
(usually a comment in social media) and a target which in comparison
with the previous type is a single named entity instead of a full
claim. The model should give one of three categories (Favor, Against,
None) for example:
Target:
legalization of abortion
Text: The pregnant are more than walking incubators, and have rights
Stance:
Favor
[Semeval 2016 task 6](http://alt.qcri.org/semeval2016/task6/) tackles this issue here is the challenge paper [7] and [8], [9] are examples of submissions for this task. [10] is an amazing reference to understand the task of stance detection in tweets and its relation to sentiment analysis. [This](http://stel.ub.edu/Stance-IberEval2017/) challenge tackles the same task but only on one target “the independence of Catalonia”. [11] applies the same technology to news articles rather than tweets see [This](https://www.youtube.com/watch?v=WYckOr2NhFM) demo. didn’t find any Arabic papers or dataSets.
A more advanced version of this task is to first identify the
targets and then generate the stance, this field of target
identification is still in its infancy an example is [subtask
C of semeval 2019 task 6](file:///home/simn/Downloads/factmata/stance%20detection/semeval%202019%20task%206) and [subtask
A of semeval 2016 task 5](http://alt.qcri.org/semeval2016/task5/)
Finally this Idea is very similar to the Idea of Aspect level sentiment analysis tackled by semeval following tasks: [2014 task 4](http://alt.qcri.org/semeval2014/task4/), [2015 task 12](http://alt.qcri.org/semeval2015/task12/), and [2016 task 5](http://alt.qcri.org/semeval2016/task5/) which includes Arabic dataset. See [12] from semeval 2016 task 5 provides a general introduction while [13] provides an older (far more comprehensive) review of the field. To simply understand the task view [This](https://developer.aylien.com/text-api-demo) demo.
## Conclusion
By now I hope that you have an initial understanding of the tasks of stance detection.
If you are interested in subjective stance detection and you are planning to start your customer analysis service then jump to [our piece](/blog/subjective-stance-detection-what-is-it-and-how-to-do-it) on this task.
We are planning to publish a new article about objective stance detection and fact checking.
In any case don't forget to check out the references below.
##
## Further Reading
[1] A. Hanselowski *et al.*, “A retrospective analysis of the
fake news challenge stance detection task,” *ArXiv Prepr.
ArXiv180605180*, 2018.
[2] B. Riedel, I. Augenstein, G. P. Spithourakis, and S. Riedel, “A
simple but tough-to-beat baseline for the Fake News Challenge stance
detection task,” *ArXiv Prepr. ArXiv170703264*, 2017.
[3] C. Conforti, M. T. Pilehvar, and N. Collier, “Towards Automatic
Fake News Detection: Cross-Level Stance Detection in News Articles,”
in *Proceedings of the First Workshop on Fact Extraction and
VERification (FEVER)*, 2018, pp. 40–49.
[4] P. Bourgonje, J. M. Schneider, and G. Rehm, “From clickbait to
fake news detection: an approach based on detecting the stance of
headlines to articles,” in *Proceedings of the 2017 EMNLP
Workshop: Natural Language Processing meets Journalism*, 2017, pp.
84–89.
[5] R. Baly, M. Mohtarami, J. Glass, L. Màrquez, A. Moschitti, and
P. Nakov, “Integrating stance detection and fact checking in a
unified corpus,” *ArXiv Prepr. ArXiv180408012*, 2018.
[6] T. Elsayed *et al.*, “CheckThat! at CLEF 2019: Automatic
Identification and Verification of Claims,” in *European
Conference on Information Retrieval*, 2019, pp. 309–315.
[7] S. Mohammad, S. Kiritchenko, P. Sobhani, X. Zhu, and C. Cherry,
“Semeval-2016 task 6: Detecting stance in tweets,” in *Proceedings
of the 10th International Workshop on Semantic Evaluation
(SemEval-2016)*, 2016, pp. 31–41.
[8] I. Augenstein, A. Vlachos, and K. Bontcheva, “Usfd at
semeval-2016 task 6: Any-target stance detection on twitter with
autoencoders,” in *Proceedings of the 10th International Workshop
on Semantic Evaluation (SemEval-2016)*, 2016, pp. 389–393.
[9] G. Zarrella and A. Marsh, “Mitre at semeval-2016 task 6:
Transfer learning for stance detection,” *ArXiv Prepr.
ArXiv160603784*, 2016.
[10] S. M. Mohammad, P. Sobhani, and S. Kiritchenko, “Stance and
sentiment in tweets,” *ACM Trans. Internet Technol. TOIT*,
vol. 17, no. 3, p. 26, 2017.
[11] S. Ruder, J. Glover, A. Mehrabani, and P. Ghaffari, “360
stance detection,” in *Proceedings of the 2018 Conference of the
North American Chapter of the Association for Computational
Linguistics: Demonstrations*, 2018, pp. 31–35.
[12] M. Pontiki *et al.*, “Semeval-2016 task 5: Aspect based
sentiment analysis,” in *Proceedings of the 10th international
workshop on semantic evaluation (SemEval-2016)*, 2016, pp. 19–30.
[13] K. Schouten and F. Frasincar, “Survey on aspect-level
sentiment analysis,” *IEEE Trans. Knowl. Data Eng.*, vol. 28,
no. 3, pp. 813–830, 2015.
### Subjective Stance Detection: What Is It? And How to Do It?
- URL: https://mohammadshaker.com/en/blog/subjective-stance-detection-what-is-it-and-how-to-do-it
- Date: 2020-01-18T00:00:00.000Z
- Tags: almeta.io, arabic, engineering, nlp, Human-Written
Subjective stance detection extracts sentiment toward a specific target — not the whole document — enabling fine-grained opinion mining. Given a review saying 'The RAM is small, but the price is low,' it can identify opposing sentiments for each aspect. This post explains the task and covers both feature-based and neural approaches.
#### Content
If you don't know what is stance detection make sure to check our [article](/blog/stance-detection-state-of-the-art) on it. Are we on the same page? Cool let's go.
First a motivational example:
Many products on the internet allow the user to leave some feedback this feedback is usually reviewed manually to figure out what are the users likes or dislikes in the product, what are the features they desire, and what are the problems they are facing. However wouldn't it be amazing if we can find what specific aspect of the product does the review like or dislike? for example: in a review of a smartphone like the following:
> The RAM is really small, the price is low though
Would it be possible to figure out that the author hates the small RAM but is happy about the cheap price? well, keep that thought because this is exactly what we are gonna discuss today.
This article is a part of our series on political bias detection we will hopefully introduce you to the various aspects of our political bias detection system, and you can learn about:
- [How can we predict the political orientation behind a piece of the news?](/blog/political-orientation-detection-ai-and-nlp-approach)
- [What is Stance detection? and what are the different types of it?](/blog/stance-detection-state-of-the-art)
- [What is subjective stance detection? what is distance supervision? and why they make a cute couple?](/blog/subjective-stance-detection-what-is-it-and-how-to-do-it)
- [What is ALSA, ELSA and how can opinion mining save us from political bias?](/blog/aspect-level-vs-entity-level-sentiment-analysis)
- [How to implement an initial political bias detector just from sentiment analysis and some probabilistic distribution? (warning cool visualizations)](/blog/from-sentiment-to-political-bias-in-the-arab-world-and-the-arabic-content)
- How to Visualize a Political Bias Metric
## The What
In this task given a piece of text and a target extract the sentiment polarity towards that target, the target can be explicitly mentioned within the text or it can be inferred from it, usually the target would be drawn from a close list however some methods accepts ant string as a possible target.
## The How
To achieve this task using a Machine Learning methodology we need 2 main things, a model to train and a data source to feed that model, let us start with the data first
### Where Can I Get Some Training Dataset
There are some datasets in English. See [3] for a complete list and the sources of these data sets vary.
Many rely on online debates sites such as [this](http://www.createdebate.com/debate/show/The_Democrat_Presidential_candadates_all_want_to_get_rid_of_the_Hyde_amendment_3), which are available only in English these topics include both serious and funny debates and cover a wide variety of targets.
Congressional and parliamentary debates regarding controversial resolutions have also been used, such means can allow the creation of very large and relatively clean datasets in an automatic or semi-automatic manner.
The authors in [4] use the argumentative essays written by TESOL students to create their data sets.
When it comes to social media, manual annotation is a must, however, some tricks can simplify the process like utilizing hashtags, emojis,… see [5] for the details of the sem-eval 2016 twitter stance data creation.
Finally, the authors in [6] also manually annotate their multi-target news stance detection dataSet. In order to simplify the process of annotation of full articles, the annotators use the following scheme to annotate for polarity:
- In case the target is explicitly mentioned within the article: a small context of 2-3 sentences before and after the target mention
- In case the target is not found in the article: they use the title and the first paragraph, however simple text summarization schemes such as [textrank](https://pypi.org/project/summa/) [7] can be used.
In Arabic there is basically no data, the only dataSet we found was from [8] yet it deals with objective stance detection (i.e. fact-checking), this mainly because most of the data sources mentioned above in English are not available in Arabic.
This means that in order to tackle this problem we need to build our own dataset. but can we make this task easier?
### Enter Distant Supervision
Distant supervision is a method of supervised text classification
wherein the training data is automatically generated (mostly expanded
using an already trained model) using certain indicators present in
the text. There are multiple indicators that can be used:
- Social media indicators: hashtags, emojies, posts reactions, and so on. The authors in [5] use this scheme to expand their sem-eval2016 data, and this new data was use in 2 main ways
- as a new data set that can be added to the original set: in this case the model performance degraded mainly because of the noisy nature of this data in comparison with the manually annotated data
- to train feature extractors such as words embeddings and the authors concluded that this can boost the model performance
- Social media graph analysis: [9]–[11] here the task is taken to a higher level by finding the stance of a user towards a certain target not the stance of a single post or tweet, to achieve this type of supervision the authors in [11] for example build a multi-level similarity graph between the users based on multiple similarity metrics, then using a seed of users whom stance towards a target is already known they annotate the other users. However there are multiple problems with the practical application of this approach:
- these methods assume that the stance of a user towards a certain target is always constant i.e. people does not change their opinions of targets
- some of the similarity measures such as geographical distance, gender and age can create prejudice in the system
- Community extraction: this is mainly done through platforms like Reddit where it is safe to assume that the users of a certain narrow community such as [r/SandersForPresident](https://www.reddit.com/r/SandersForPresident/) and for less serious matters such as [r/grandpajoehate](https://www.reddit.com/r/grandpajoehate/), that the opinion of the users towards the target is consistent and thus the comments can constitute a valid dataset. Alongside debate sites such communities are extremely useful data sources for multi-target stance detection. However this type of distance supervision is not viable in Arabic due to the the lack of such sources, the only site similar to Reddit in Arabic is [hsoub](https://www.hsoub.com/) which is a community centered around programming.
- Voice prosody: finally some people sought to predict stance using the prosodic features of the talk shows hosts in order to create large data sets see [12] however the performance using speech features only is not satisfactory.
### Now let us Jump to the Implementation
- Overall simple shallow models are more favorable than deep ones as they have much simpler implementation and the gain in performance is negligible, in the semeval2016 task 6 dataSet [13] the shallow baseline of SVM with feature engineering surpassed all of the competitors models, and it took the research community 2 years to beat it with a very low margin [14]
- In most of the cases uni-target models (models trained only on a single target) [13], [15] have a much higher performance than multi-target models like [16], [17] (models trained using data for multiple targets and thus can incorporate any target), the only exception is [14] yet this comes with huge overhead with regards to training data and model complexity
- In multi-target models [14], [16]–[18] the target information is usually added to neural models using a recurrent network that goes over the embeddings of the target phrase in order to create a target representation. These information can be incorporated through simple linear transforms [18] or by using attention mechanism in higher layers [14] , see figures 2 from [14] and 1 from [18] for comparison.
- Overall the SOTA in stance detection is pretty low for sources such as tweets and online debates [13] the best performance is about 70% in f-measure and this is measured on a very restricted set of targets, for more regular sources like the students essays [4] the performance higher and can reach up to 82% f-measure. However, this is not the only factor for example the news articles the performance reported by [6] is less than 60% and this is mainly because the authors used a multi-target model. This low performance is attributed to the following issues:
- the implicit mention of targets mainly in tweets and reviews either using hashtags or more often using pronouns, the authors in [5] compared the performance of the their baseline between the set of tweets that explicitly mention the target and the set of tweets that don’t and found a difference in performance up to 30% f-measure absolute. See figure 3 for a detailed table.
- The use of sarcasm and jokes to convoy negative sentiment mainly in t
- the need of common-sense models to group target together for example capturing the fact that a positive sentiment towards (Bernie Sanders, Hilary Clinton, Democrats) is equal to having a negative sentiment towards (Donald Trump, Republicans), some efforts to resolve this issue in the case of congressional debates were carried out like in [19] but with limited success.
- The lack of data even in English with the semeval corpus being used in nearly all the stance detection publications after 2016
- Finally, based on the assumption we had at the introduction we should use the confidence of the classifiers as an indicator of the bias. However, we can’t confirm the validity of this assumption for stance detection mainly because the confidence of the trained models for this task (mainly the neural models) is usually within the vicinity of 50%, see figure 4 from [6] for example. This phenomenon is attributed to the fact that usually the articles are not biased and hold a neutral stance most of the time [20](at least in western media), the comparison with the more polarized Arabic media is not clear and we don’t have data to assert or deny the applicability of this scheme.
figure 1 linear transformation based target incorporation from [18]
figure 2 attention based target incorporation from [14]
figure 3 loss in performance due to absence of target in text from [5]
figure 4 results from [6] that shows the prevalence of neutral articles and low confidence
## Conclusion
In this article, we explored the idea of subjective stance detection and how we can find the sentiment of the text towards various targets. we have seen how we can find data for such a system in English and how can distance supervision help us in doing it for low resourced languages like Arabic.
If you feel intrigued, make sure to check out our article on [aspect and entity level sentiment analysis](/blog/aspect-level-vs-entity-level-sentiment-analysis) which is basically the next step.
##
## References
[1] M. Iyyer, P. Enns, J. Boyd-Graber, and P. Resnik, “Political
ideology detection using recursive neural networks,” in *Proceedings
of the 52nd Annual Meeting of the Association for Computational
Linguistics (Volume 1: Long Papers)*, 2014, vol. 1, pp. 1113–1122.
[2] R. Abooraig, S. Al-Zu’bi, T. Kanan, B. Hawashin, M. Al Ayoub,
and I. Hmeidi, “Automatic categorization of Arabic articles based
on their political orientation,” *Digit. Investig.*, vol. 25,
pp. 24–41, 2018.
[3] R. Wang, D. Zhou, M. Jiang, J. Si, and Y. Yang, “A Survey on
Opinion Mining: From Stance to Product Aspect,” *IEEE Access*,
vol. 7, pp. 41101–41124, 2019.
[4] A. Faulkner, “Automated classification of stance in student
essays: An approach using stance target information and the Wikipedia
link-based measure,” in *The Twenty-Seventh International Flairs
Conference*, 2014.
[5] S. M. Mohammad, P. Sobhani, and S. Kiritchenko, “Stance and
sentiment in tweets,” *ACM Trans. Internet Technol. TOIT*,
vol. 17, no. 3, p. 26, 2017.
[6] S. Ruder, J. Glover, A. Mehrabani, and P. Ghaffari, “360 stance
detection,” in *Proceedings of the 2018 Conference of the North
American Chapter of the Association for Computational Linguistics:
Demonstrations*, 2018, pp. 31–35.
[7] R. Mihalcea and P. Tarau, “Textrank: Bringing order into text,”
in *Proceedings of the 2004 conference on empirical methods in
natural language processing*, 2004.
[8] R. Baly, M. Mohtarami, J. Glass, L. Màrquez, A. Moschitti, and
P. Nakov, “Integrating stance detection and fact checking in a
unified corpus,” *ArXiv Prepr. ArXiv180408012*, 2018.
[9] R. Dong, Y. Sun, L. Wang, Y. Gu, and Y. Zhong, “Weakly-guided
user stance prediction via joint modeling of content and social
interaction,” in *Proceedings of the 2017 ACM on Conference on
Information and Knowledge Management*, 2017, pp. 1249–1258.
[10] J. Ebrahimi, D. Dou, and D. Lowd, “Weakly supervised tweet
stance classification by relational bootstrapping,” in *Proceedings
of the 2016 Conference on Empirical Methods in Natural Language
Processing*, 2016, pp. 1012–1017.
[11] O. Fraisier, G. Cabanac, Y. Pitarch, R. Besançon, and M.
Boughanem, “Stance Classification through Proximity-based Community
Detection.,” in *HT*, 2018, pp. 220–228.
[12] N. G. Ward, J. C. Carlson, O. Fuentes, D. Castan, E. Shriberg,
and A. Tsiartas, “Inferring Stance from Prosody.,” 2017.
[13] S. Mohammad, S. Kiritchenko, P. Sobhani, X. Zhu, and C. Cherry,
“Semeval-2016 task 6: Detecting stance in tweets,” in *Proceedings
of the 10th International Workshop on Semantic Evaluation
(SemEval-2016)*, 2016, pp. 31–41.
[14] C. Xu, C. Paris, S. Nepal, and R. Sparks, “Cross-Target Stance
Classification with Self-Attention Networks,” *ArXiv Prepr.
ArXiv180506593*, 2018.
[15] J. Mitrovic, B. Birkeneder, and M. Granitzer, “nlpUP at
SemEval-2019 Task 6: A Deep Neural Language Model for Offensive
Language Detection.”
[16] Y. Zhou, A. I. Cristea, and L. Shi, “Connecting targets to
tweets: Semantic attention-based model for target-specific stance
detection,” in *International Conference on Web Information
Systems Engineering*, 2017, pp. 18–32.
[17] I. Augenstein, T. Rocktäschel, A. Vlachos, and K. Bontcheva,
“Stance detection with bidirectional conditional encoding,” *ArXiv
Prepr. ArXiv160605464*, 2016.
[18] J. Du, R. Xu, Y. He, and L. Gui, “Stance classification with
target-specific neural attention networks,” 2017.
[19] M. Lai, D. I. H. Farías, V. Patti, and P. Rosso, “Friends and
enemies of Clinton and Trump: using context for detecting stance in
political tweets,” in *Mexican International Conference on
Artificial Intelligence*, 2016, pp. 155–168.
[20] M. Drissi, P. Sandoval, V. Ojha, and J. Medero, “Harvey Mudd
College at SemEval-2019 Task 4: The Clint Buchanan Hyperpartisan News
Detector,” *ArXiv Prepr. ArXiv190501962*, 2019.
## Frequently Asked Questions
### What is stance detection in NLP?
Stance detection is the task of automatically determining whether a piece of text expresses a favorable, unfavorable, or neutral position toward a specific target or claim. It is closely related to sentiment analysis but focuses on opinion direction toward a specific topic rather than general sentiment.
### How is stance detection different from sentiment analysis?
Sentiment analysis classifies text as positive, negative, or neutral in general tone. Stance detection determines the author's position toward a specific target. A text can have positive sentiment but oppose a target (e.g., “I'm happy the bad policy was rejected” — positive sentiment, against the policy).
### What are the applications of stance detection?
Stance detection is used in fake news verification (checking source agreement), political analysis (mapping ideological positions), rumor detection (identifying supporting vs denying claims), and debate mining (extracting pro/con arguments on issues).
### Supervised Article Informativeness Prediction - The What and the How
- URL: https://mohammadshaker.com/en/blog/supervised-article-informativeness-prediction-the-what-and-the-how
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, ideas, ml, nlp, Human-Written
Supervised informativeness prediction trains a classifier to score text quality using features like readability, cliche density, content-to-function word ratio, and term informativeness. We explain the full pipeline — feature extraction, label generation, model selection, and evaluation — applied to Arabic news articles in the Almeta project.
#### Content
Let's start with a question, given 2 articles A and B that talks about the exact same thing, what makes one of them more informative than another?
Is it the ease of reading? the amount of details? or is it the correct structuring of the text?
Well, maybe. In our previous [article](/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of) we have explored various metrics like this and saw how these metrics can be calculated quantitatively.
But let us explore a different assumption: Here we are assuming that the informativity of a piece of content (text) is an abstract concept that is hard to bin down exactly and that we can measure it using common sense just like the coherence of a text or a musical composition.
These common sense rules must have been constructed in our mind through years of experience. If this assumption is valid, wouldn't it be possible to transfer this experience to an information systems? In this article we are exploring exactly this possibility.
This article is a part of our research on informativeness, and we have A LOT of it. Uf you are interested in more details, you can check each individual piece to learn about:
- [What makes an article informative in our view and how can we measure this quantitatively?](/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of)
- [How to detect cliches in text? and why do they matter?](/blog/how-to-detect-cliches-in-text)
- [How to measure the informativeness of each individual word?](/blog/term-informativeness-estimation-in-the-arabic-language)
- [How to train a supervised model to measure informativeness?](/blog/supervised-article-informativeness-prediction-the-what-and-the-how)
- [Can we get data to train such a model from summaries?](/blog/can-you-measure-a-text-informativeness-using-its-summary) ?
- [Can we get such data in Arabic?](/blog/automatically-tagging-data-for-content-informativity-scoring)
- [How can we rank articles based on how informative they are?](/blog/how-to-rank-articles-based-on-how-informative-they-are-using-snorkel)
## The What
In this task, the goal is to assign a given piece of text a tag (or a number) representing the level of informativeness or detail this text holds usually by training a model to do that.
The variations in this task is based on the size of the text (article, paragraph, sentence) and the system architecture (supervised vs unsupervised), another name for this task is Specificity Prediction.
## The How
Usually the methodologies are split based on the size of text that the model processes and we can categorize them into a paragraph or a sentence level.
### Paragraph Level
These methodologies will yield a model that can process large chunks of text at once, which is suitable for demanding applications.
#### Using Lead Paragraphs
In [1] the authors relay on the intuitive idea that usually in any news article the lead paragraph can be used as a good summary especially if the author utilizes the inverted pyramid style, an example of a good lead paragraph looks something like this:
> “The European Union’s chief trade negotiator, Peter Mandelson, urged the United States on Monday to reduce subsidies to its farmers and to address unsolved issues on the trade-in services to avert a breakdown in global trade talks. Ahead of a meeting with President Bush on Tuesday, Mr. Mandelson said the latest round of trade talks, begun in Doha, Qatar, in
> 2001, are at a crucial stage. He warned of a ”serious potential breakdown” if rapid progress is not made in the coming months.”
However, this generalization does not apply to all authors as some uses a more creative approach, e.g.:
> “ART consists of limitation,” G. K. Chesterton said.” The most beautiful part of every picture is the frame.” Well, put, although the buyer of the latest multi-million dollar Picasso may not agree. But there are pictures – whether sketches on paper or oils on canvas – that may look like nothing but scratch marks or listless piles of paint when you bring them home from the auction house or dealer. But with the addition of the perfect frame, these works of art may glow or gleam or rustle or whatever their makers intended them to do.
It is clear that the *first* lead paragraph is more informative than the *second*: i.e. *first > second*.
The authors relied on the idea that leads like the first example constitute a more plausible summary of the article than the second example and automatically created a corpora for informativeness prediction using a corpora of human summarized news articles simply by comparing the human generated summary with the lead paragraph and assigning the lead paragraph a tag of (informative/ creative) based on this similarity.
After generating the dataset they trained an SVM on several features including some psycholinguistic attributes such as the age of acquisition, imagery, concreteness, familiarity and ambiguity of each word.
They report a binary accuracy between 0.7 and 0.8 across the different genres of news articles and indicate that genre-specific models have a high advantage over general ones.
#### Using Extractive Summaries
In [2], the authors carry out a similar task to [1], they use extractive summaries and label every paragraph in the original texts as either positive (if it appears in the summary) or unlabeled (if it does not appear in the summary).
The second class is unlabeled because of the fact that annotators bias does not allow for faithful labelling of out-of-summary paragraphs as unimportant.
Therefore the authors use a simple semi-supervised model (logistic regression) and very simple features similar to [1] (only lexicon features) their reported results on the DUC2002 dataset and on hand labeled set from NYT is around 0.69 f1 score however their recall is nearly 0.9 in both sets while their accuracy is as low as 0.59 this indicate the simplicity of
the used features and the inability of the model to fit the training data.
#### Available Resources
[This](https://github.com/Franck-Dernoncourt/summarization-corpora) is a comprehensive list of English summarization corpora for section 2.1, Arabic summarization datasets include [EASC](https://sourceforge.net/projects/easc-corpus/), [DUC2004/2001/2002](https://duc.nist.gov/duc2004/), and [Kalimat](https://sourceforge.net/projects/kalimat/) which would be the closest to the proposed method.
Please check [this article](/blog/can-you-measure-a-text-informativeness-using-its-summary) for an in-depth evaluation of these resources and [this one](/blog/automatically-tagging-data-for-content-informativity-scoring) for our results of the implementation of the aforementioned task on Arabic Language.
### Sentence level (sentence specificity)
This similar task is based on the idea that more specific (detailed) sentences are more informative e.g.
> This brand is very popular and many people use its products regularly.
> Mascara is the most commonly worn cosmetic, and women will spend an average of $4,000 on it in their lifetimes.
There are several applications of this from summarization to IR. Things like argument detection … but most importantly in essay evaluation:
- The task was originally introduced in [3] where the authors used the [pdtb](https://www.seas.upenn.edu/~pdtb/) dataset to generate their corpora and trained on it a **binary classifier** based on several *lexical* features including:
- Sentence length (shorter sentences tend to be more specific)
- Average polarity using a lexicon (English MPQA)
- The average estimate of specificity based on hypernym relations between words using WordNet
- Lexical counts of (numbers, money amounts, dates, …) as they can serve as an indicator of specificity as well
- n-grams and pos
They report an accuracy of 76% on instantiation discord relations (when a certain thing is an instance of a larger concept) and 60% on specification relations (when one concept is a special case of another concept)
- In [4] the authors used the same dataset as above but utilized simpler features including:
- Counts: number of words in the sentence, number of numbers, capital letters and nonalphanumeric symbols in the sentence, average words length and number of stop words
- IDF of the words
- Lexicon features the same ones used in [3]
- Psycholinguistic features from [2]
- Counts vectors and Word2Vec
Furthermore the authors co-train their model using a larger un-annotated dataset from WSJ, these new improvements pushes the model performance up to 81% accuracy, 79% F1.
- Finally, in [5] the authors propose a neural model based on LSTM running on words embeddings and hand-crafted features from [3]. This model is trained on news data from [3] and [4]. However the authors provide a method for unsupervised domain adaptation. They hand labelled using MTruck 3 sets from yelp, Twitter, and movie reviews (which is very different than news articles) and tested their method on them. Both the target sets (yelp, twitter and movie reviews) as well as the the code is available on [Github](https://github.com/wjko2/Domain-Agnostic-Sentence-Specificity-Prediction).
## Conclusion
- Overall this task is highly correlated with the goal of informativeness score as this is what it measures explicitly (especially in the cases [1] and [2])
- the implementation of [1] / [2] for both Arabic and English is viable as datasets are available to start with in the 2 languages and that the proposed methods are fairly simple.
##
## References
[1] Y. Yang and A. Nenkova, “Detecting information-dense texts in multiple news domains,” in Twenty-Eighth AAAI Conference on Artificial Intelligence, 2014.
[2] Y. Yang, F. S. Bao, and A. Nenkova, “Detecting (un) important content for single-document news summarization,” ArXiv Prepr. ArXiv170207998, 2017.
[3] A. Louis and A. Nenkova, “Automatic identification of general and specific sentences by leveraging discourse annotations,” in Proceedings of 5th International Joint Conference on Natural Language Processing, 2011, pp. 605–613.
[4] J. J. Li and A. Nenkova, “Fast and accurate prediction of sentence specificity,” in Twenty-Ninth AAAI Conference on Artificial Intelligence, 2015.
[5] W.-J. Ko, G. Durrett, and J. J. Li, “Domain agnostic real-valued specificity prediction,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2019, vol. 33, pp. 6610–6617.
[6] J. J. Li, B. O’Daniel, Y. Wu, W. Zhao, and A. Nenkova, “Improving the annotation of sentence specificity,” in Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), 2016, pp. 3921–3927.
[7] L. Lugini and D. Litman, “Predicting specificity in classroom discussion,” ArXiv Prepr. ArXiv190901462, 2019.
[8] J. Li, “From Discourse Structure To Text Specificity: Studies Of Coherence Preferences,” 2017.
### Term Informativeness Estimation in the Arabic Language
- URL: https://mohammadshaker.com/en/blog/term-informativeness-estimation-in-the-arabic-language
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Not all words contribute equally to a document's meaning. Term informativeness estimation assigns a score to each word or phrase based on how much unique content it carries — using TF-IDF variants, semantic similarity, and corpus-level statistics. We apply these methods to Arabic text, where morphological richness makes term boundaries harder to define.
#### Content
In our effort at Almeta to provide the articles with the highest informative value to the Arabic readers, we have employed several methods to measure the informativeness of a piece of news, in this article we will shed light to one of the algorithms we are using, specifically term informativeness.
This article is a part of our research on measuring text informativeness if you are interested [jump directly to the gist](/blog/informativity-detection-almetas-research-gist) of our research or review other parts:
- [What makes an article informative in our view and how can we measure this quantitatively?](/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of)
- [How to detect cliches in text? and why do they matter?](/blog/how-to-detect-cliches-in-text)
- [How to train a supervised model to measure informativeness?](/blog/supervised-article-informativeness-prediction-the-what-and-the-how)
- [Can we get data to train such a model from summaries?](/blog/can-you-measure-a-text-informativeness-using-its-summary) ?
- [Can we get such data in Arabic?](/blog/automatically-tagging-data-for-content-informativity-scoring)
- [How can we rank articles based on how informative they are?](/blog/how-to-rank-articles-based-on-how-informative-they-are-using-snorkel)
## The What
This task is based on the idea that some words carry more semantic content than others, which leads to the notion of term specificity, or informativeness.
A closely related task is term weighting (for search engines, indexing, summarization, ...)
The goal of this task is to assign every term a certain metric representing its importance. This type of methods was dominant before the neural explosion as assigning a measure of informativeness to every term can be used in many NLP and IR applications including (key phrase extraction, summarization, ... ) basically by ranking the terms using this informativeness metric.
Ideally,
with such metric we can estimate the article informativeness by the
average informativeness of its terms.
There is no simple manner for evaluating this task (because it is mostly metric based). Interesting ways to do it includes:
- Observing the avg, median and relative stats of metrics across several test sets
- Measuring the relative informativeness gain between parents and children in lexicons like WordNet (based on the assumption that children in such hierarchies tend to be more specific and thus more informative than their parent terms)
- Embedding the metrics as features in other tasks most notably key phrase extraction.
## The How
The methods can be used to measure a term informativeness can be categorized in three major sets:
### Statistics-based
These metrics collect some statistics about each of the terms from a relatively large corpus of texts. Many of these methods are based on some variant of term frequency, and they are especially common in the IR community. Note that all of these methods are language ***in***dependent:
- TF-IDF (term frequency inverted document frequency): the TF-IDF is a very popular algorithm in NLP it weights the terms in a corpus, based on their occurrence frequencies, the details of TF-IDF is beyond the scoop of this article and the reader is referred to [this](http://www.tfidf.com/) fairly simple explanation.
- Term variance [1]: is a measure of the variance of the term frequency from the mean. this metric is given by:
where D is the size of training corpus, is the frequency of term w in document d and
Here $latex \bar t\_w$ is the the total number of occurrence of the word w in the corpus
- Burstiness [1]: this metric basically compares the collection frequency and document frequency directly, this metric is given by
- IDF based metrics: these are metrics that represents a variation over the Inverse document frequency. The IDF is a very popular algorithm in the field of Information Retrieval here are some of these metrics:
- Residual IDF [2]: this metric is given by:
where
- IDF gain [2] : another IDF based metric given by
- Mixture score[3] : this score is based on the idea that although topic-centric words are somewhat rare (across the whole corpora). But they also exhibit two modes of operation:
1. A high frequency mode, when the document is relevant to the word, and
2. A low (or zero) frequency mode, when the document is irrelevant. Based o this they suggest modeling this fact through a mixture model of binomial distributions. The score is given by:
Where $latex P\_mix$ is the mixture model probability of the word appearances in the whole corpora while $latexP\_uni$ is the uni-gram probability of the word in the corpora. Where the mixture model is estimated using EM.
However there is a
couple of issues with this metric:
- The probabilities is calculated for each individual word (great overhead)
- Using a mixture of binomial distribution also greatly increase the complexity as for every single word each document must be represented as a one-hot vector
- The final results of the score have negligible to no improvement over simpler metrics like RIDF
### Semantics based
These are metrics that creates some sort of representation of the term meaning and harness it to measure the term informativeness.
LSA informativity [4]: is a very simple and plausible metric, LSA here represents the Latent Semantic Analysis also a very popular algorithm in the field of NLP and is also out of the scope of our article but again we refer you to [9] for a general introduction to LSA,
Here, LSA is used to model the corpora and for each word a vector is generated, usually, the de-facto applications of LSA are clustering, similarity measure, … all of which are based on measuring the cosine similarity between words vectors (or weighted sums of them for the case of document similarity).
However, the authors utilizes the length of the word vector (which is usually used as the weighting terms in the document weighted sum) to measure term informativeness. This is motivated by the fact that “Intuitively, the vector length tells us how much information LSA has about this vector. [...] Words that LSA knows a lot about (because they appear frequently in the training corpus[...]) have greater vector lengths than words LSA does not know well."
Function words that are used frequently in many different contexts have low vector lengths - LSA knows nothing about them and cannot tell them apart since they appear in all contexts.” which greatly overlaps with the motivations of the mixture scores mentioned above. Their simple metric is as follows:
Interestingly enough, the same analogy can be carried out to Word2Vec, see this question and therefore a similar measure can be based on Word2Vec.
However, the main problem with such approaches just like statistics methods is their dependence on the training corpora and the issue of OOV (out of vocabulary). Furthermore, it seems that using genre specific corpora is superior to building the systems using generic ones like Wikipedia.
### Model Reuse
As stated above these metrics have been originally used to rank
the terms in tasks like keyphrase extraction and summarization, but
as these tasks are currently dominated by neural supervised models
one option is to use the confidence of such models (specifically
key-phrase extraction) to rank the words based on their importance.
However there are 2 issues with this scheme:
- These models confidence tend to be non-smooth with non key words having nearly 0 confidence
- Such scheme will cause issues of error propagation and will by tightly coupled with the training set
### Term Weighting
The task is really similar to the task of term weighting in IR. this task is out of the scope of our discussion Now and we will hopefully present a separate article for this however if you are out of patience you can review the following papers (in order of relevance): [1], [5]–[8]
### Final remarks regarding implementation
- Nearly all of the presented methods are language agnostic and are therefore any monolingual dataset can be utilized,
- Some of the metrics like TF-IDF, IDF, and LSA stuff are really simple and efficient to compute, and given the proper training set can be implemented in 1 to 2 workdays.
- Most of the methods depend directly or indirectly (the case of model reuse) on the training corpora and thus focusing on a restricted domain for a start is advised.
- Some metrics like LSA or Word2Vec have a heavy memory foot print and indexing frameworks and structures might be of relevance
## Conclusion
In this article we tried to explore a simple question of how can we measure the importance of a single word in text, and how such a simple measure can be a proxy to measuring the informativeness of the whole article.
##
### References
[1] Z. Wu and C. L. Giles, “Measuring term informativeness in
context,” in *Proceedings of the 2013 conference of the north
american chapter of the association for computational linguistics:
human language technologies*, 2013, pp. 259–269.
[2] K. Papineni, “Why inverse document frequency?,” in
*Proceedings of the second meeting of the North American Chapter
of the Association for Computational Linguistics on Language
technologies*, 2001, pp. 1–8.
[3] J. D. Rennie and T. Jaakkola, “Using term informativeness for
named entity detection,” in *Proceedings of the 28th annual
international ACM SIGIR conference on Research and development in
information retrieval*, 2005, pp. 353–360.
[4] K. Kireyev, “Semantic-based estimation of term
informativeness,” in *Proceedings of Human Language
Technologies: The 2009 Annual Conference of the North American
Chapter of the Association for Computational Linguistics*, 2009,
pp. 530–538.
[5] N. Nanas, V. Uren, A. De Roeck, and J. Domingue, “A
comparative study of term weighting methods for information
filtering. KMi-TR-128,” *Knowl. Media Institue Open Univ.*,
2003.
[6] G. Murray and S. Renals, “Term-weighting for summarization of
multi-party spoken dialogues,” in *International Workshop on
Machine Learning for Multimodal Interaction*, 2007, pp. 156–167.
[7] M. Shirakawa, T. Hara, and S. Nishio, “N-gram idf: A global
term weighting scheme based on information distance,” in
*Proceedings of the 24th International Conference on World Wide
Web*, 2015, pp. 960–970.
[8] C. Lioma and R. Blanco, “Part of Speech Based Term Weighting
for Information Retrieval,” *ArXiv Prepr. ArXiv170401617*,
2017.
[9] Landauer, Thomas K., Peter W. Foltz, and Darrell Laham. "An
introduction to latent semantic analysis." *Discourse
processes* 25.2-3 (1998): 259-284.
### What Makes an Article Informative - and How Computers Can Measure Informativity of Text Content
- URL: https://mohammadshaker.com/en/blog/what-makes-an-article-informative-and-how-computers-can-measure-informativity-of
- Date: 2020-01-18T00:00:00.000Z
- Tags: arabic, ideas, ml, nlp, Human-Written
An informative article is easy to recognize but hard to define computationally. Measurable proxies include readability score, skimmability via header and list structure, content-to-function word ratio, cliche density, and term-level informativeness. We explore how to combine these features into a quantitative informativeness signal for Arabic news.
#### Content
The Concept of an informative text is really abstract and it is hard to come up with a definitive formula to measure it, in this article we will explore some of the features that we believe can make an article more worthy and show how can these features can be measured quantiatively.
This is our first piece on informativeness but there is A LOT more if you are intrigued you can [jump directly to the gist](/blog/informativity-detection-almetas-research-gist) or read the details of individual articles:
- [How to detect cliches in text? and why do they matter?](/blog/how-to-detect-cliches-in-text)
- [How to measure the informativeness of each individual word?](/blog/term-informativeness-estimation-in-the-arabic-language)
- [How to train a supervised model to measure informativeness?](/blog/supervised-article-informativeness-prediction-the-what-and-the-how)
- [Can we get data to train such a model from summaries?](/blog/can-you-measure-a-text-informativeness-using-its-summary) ?
- [Can we get such data in Arabic?](/blog/automatically-tagging-data-for-content-informativity-scoring)
- [How can we rank articles based on how informative they are?](/blog/how-to-rank-articles-based-on-how-informative-they-are-using-snorkel)
### Readability
This is based on the assumption that an informative article should be readable, possible metrics include LIX, FOG, … and are heavily used in the literature. We, at Almeta, have built our own readability measuring system. You can read more about our system [here](/blog/how-to-measure-text-readability).
### Skimmability
This is based on the assumption that an informative text can be skimmed easily i.e. A text is informative when there are a lot of pieces of information one could capture without knowing the context of the author.
So for example: “He is my best friend” is less informative than “Max Mustermann has a best friend called Martin Muster”.
A simple metric for measuring the informativity in a given text is the relative amount of content words to non-content words.
Content words are nouns, proper nouns, verbs, and adjectives. Some definitions include adverbs and some prepositions, but a test showed that those were not useful. The content function ratio (CFR) is calculated like this:
$latex CFR(text) = NumContentWords/NumFunctionWords$
### Lexical Density
Lexical density is defined as the number of lexical words (or ***content*** words) divided by the total number of words.
Lexical words give a text its meaning and provide information regarding what the text is about. More precisely, lexical words are simply nouns, adjectives, verbs, and adverbs. Nouns tell us the subject, adjectives tell us more about the subject, verbs tell us what they do, and adverbs tell us how they do it.
Other kinds of words such as articles (a, the), prepositions (on, at, in), conjunctions (and, or, but), and so forth are more grammatical in nature and, by themselves, give little or no information about what a text is about. These non-lexical words are also called function words. Auxiliary verbs, such as "to be" (am, are, is, was, were, being), "do" (did, does, doing), "have" (had, has, having) and so forth, are also considered non-lexical as they do not provide additional meaning.
With the above in mind, lexical density is simply the percentage of words in the written (or spoken) language which give us information about what is being communicated.
### Lexical Richness (LR)
LR is a set of measures that is based on the lexical nature of the text including:
- Lexical richness is a measure of how many tokens (*individual* words and punctuation) occur in a given text, divided by how many types (*unique* words and punctuation) The higher the number, the more lexically rich it is, meaning, the more unique words (and punctuation) it contains compared to the total number of words and punctuation it contains. A more advanced variant uses stuff like WordNet to group synonyms into the same count;
- Type to Token ratio: (and related metrics) basically the ratio between the size of the pos tags set and the vocabulary of the text (other metrics uses more detailed pos tagging in order to find say the different tenses used, the usage of passive voice, number of nominal types, …);
- Other LR metrics include: it is hard for us to cover all of them here but the MLTD [2], and VOC-D [3] are some of the most common.
### Term Informativeness
This task is based on the idea that some words carry more semantic content than others, which leads to the notion of term specificity, or informativeness.
The goal of this task is to assign every term a certain metric representing its importance. You can get a better idea about this task by reading our [special article about Term informativeness.](/blog/term-informativeness-estimation-in-the-arabic-language)
### Metadiscourse Markers
Metadiscourse markers (such as ‘firstly’ and ‘in conclusion’) are very important since they refer explicitly to aspects of the organisation of a text or indicate a writer’s stance towards the text’s content or towards the reader.
Here we are using the basic Idea that if a text is more structured and properly interlinked then it is easier and more informative.
They are important markers in more academic text styles. The percentage of these markers also can help in measuring the structuring of the text. [This](https://textinspector.com/help/metadiscourse/) is a good explanation of these markers, [4] is a more detailed explanation, [5] talks about their visualization and [6] is a comparison between their usage in the English and Arabic Languages.
### Text Denoising
This is based on the assumption that in a specialized topics like (nuclear physics, genome, ... ) that have less readable passages (ones with longer words and sents) are more informative as they include the details of the article.
This can be especially important for tasks relating to keywords extraction like entity-level sentiment detection and relation extraction. The readability is measured here using Fog-Index and Normalized Fog-Index as they focus on words length rather than other elements, the measures are calculated as follows:
The idea of text denoising is to build a filter to remove sentences that have higher readability that a certain threshold as these is deemed less helpful in tasks related to key-phrase generation, here is [a source code](https://sourceforge.net/p/textdenoisingtool/wiki/Home/) for the denoising thing by the author.
### Cliché Detection
This is based on the assumption that text with cliches is not informative and thus articles with fewer cliches should be more informative, the authors in [8] provide a method of doing this automatically.
### Text Coherence
This is based on the idea that a coherent text should be more informative, there is a full line of research in NLP community interested in automating this task, the publication [9] is a very good example.
### Automated Essay Scoring
There is a very wide field of research on this task and it encompasses many of the ideas previously mentioned, the main problem is the need for large training sets and the focus of the methods on neural models, see for example [11].
### Supervised Informativeness Detection
basically training a binary classifier to predict if a text is informative or not and using its confidence as a metric:
- Cons:
- there will be a need to label a large set of data
- the results of the system will generally be domain dependent
- sensitivity analysis is generally harder
- Pros:
- performance should be generally more consistent and better
The guys in [12] did exactly that on news articles and suggest a method to automatically boost their training data.
### Conclusion
In this article we tried to cover several aspects that marks a piece of text as informative and how can we measure them in an information service, in our view the only way to truly measure the informativeness of an article is to use several measurements as proxy to the more abstract Idea of informativeness.
This was the first article in informativeness series if you are intrigued you can [jump directly to the gist](/blog/informativity-detection-almetas-research-gist) or read the details of individual articles:
- [How to detect cliches in text? and why do they matter?](/blog/how-to-detect-cliches-in-text)
- [How to measure the informativeness of each individual word?](/blog/term-informativeness-estimation-in-the-arabic-language)
- [How to train a supervised model to measure informativeness?](/blog/supervised-article-informativeness-prediction-the-what-and-the-how)
- [Can we get data to train such a model from summaries?](/blog/can-you-measure-a-text-informativeness-using-its-summary) ?
- [Can we get such data in Arabic?](/blog/automatically-tagging-data-for-content-informativity-scoring)
- [How can we rank articles based on how informative they are?](/blog/how-to-rank-articles-based-on-how-informative-they-are-using-snorkel)
##
### Further Reading
[1] R. Shams, “Identification of informativeness in text using
natural language stylometry,” 2014.
[2] P. M. McCarthy, “An assessment of the range and usefulness of
lexical diversity measures and the potential of the measure of
textual, lexical diversity (MTLD),” The University of Memphis,
2005.
[3] P. M. McCarthy and S. Jarvis, “vocd: A theoretical and
empirical evaluation,” *Lang. Test.*, vol. 24, no. 4, pp.
459–488, 2007.
[4] S. G. Sanford, “A comparison of metadiscourse markers and
writing quality in adolescent written narratives,” 2012.
[5] D. Simsek, S. Buckingham Shum, A. Sandor, A. De Liddo, and R.
Ferguson, “XIP Dashboard: visual analytics from automated
rhetorical parsing of scientific metadiscourse,” 2013.
[6] A. H. Sultan, “A contrastive study of metadiscourse in English
and Arabic linguistics research articles,” *Acta Linguist.*,
vol. 5, no. 1, p. 28, 2011.
[7] Z. Wu and C. L. Giles, “Measuring term informativeness in
context,” in *Proceedings of the 2013 conference of the north
american chapter of the association for computational linguistics:
human language technologies*, 2013, pp. 259–269.
[8] P. Cook and G. Hirst, “Automatically assessing whether a text
is clichéd, with applications to literary analysis,” in
*Proceedings of the 9th Workshop on Multiword Expressions*,
2013, pp. 52–57.
[9] Z. Lin, H. T. Ng, and M.-Y. Kan, “Automatically evaluating text
coherence using discourse relations,” in *Proceedings of the 49th
Annual Meeting of the Association for Computational Linguistics:
Human Language Technologies-Volume 1*, 2011, pp. 997–1006.
[10] C. Danescu-Niculescu-Mizil, G. Kossinets, J. Kleinberg, and L.
Lee, “How opinions are received by online communities: a case study
on amazon. com helpfulness votes,” in *Proceedings of the 18th
international conference on World wide web*, 2009, pp. 141–150.
[11] Y. Farag, H. Yannakoudakis, and T. Briscoe, “Neural automated
essay scoring and coherence modeling for adversarially crafted
input,” *ArXiv Prepr. ArXiv180406898*, 2018.
[12] Y. Yang and A. Nenkova, “Detecting information-dense texts in
multiple news domains,” in *Twenty-Eighth AAAI Conference on
Artificial Intelligence*, 2014.
### Google's AutoML Overview
- URL: https://mohammadshaker.com/en/blog/googles-automl-overview
- Date: 2020-01-16T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Google AutoML lets teams with limited ML expertise build custom models for text classification, translation, and image recognition. Its main strength is automating architecture search and training. Its key limitation for Arabic NLP is that several services are English-first, with variable quality on other languages.
#### Content
In this post, we are exploring how Google's AutoML can help us in Almeta in developing automatic Arabic language processing tools.
Before start if you are not familiar with the term AutoML you can refer to our previous post on this topic.
## Who is Google AutoML for? and When to Use It?
The targeted audience by Google's cloud autoML are people who have limited knowledge in machine learning.
The main goal of this cloud service is to let the user build his own AI model that is tailored to his business needs, if the provided services by Google's AI API can't satisfy his needs, even if he doesn't have enough knowledge in machine learning.
In general, anyone can use these services to build a custom AI model on the fly.
## What Kinds of AutoML Services Does Google Provide?
Let's see what kinds of AutoML services are provided by Google in the NLP field, and whether they can be adapted to process the Arabic language.
### Cloud AutoML Natural Language Classification
Enables you to create custom machine learning models to classify content into a custom set of categories.
According to Google's documentation, the current service supports content classification in English language text. We can train a custom model to classify text in other languages including Arabic, but the model quality may vary.
#### How to train the model?
Build your own dataset, and upload it as .csv file. The trained model is automatically deployed, and we can get its predictions using an API.
### Cloud AutoML Natural Language Entity Extraction
Enables you to create custom machine learning models to identify a custom set of entities.
According to Google's documentation, this service currently supports entity analysis in English language text. We can train a custom model using text in other languages, but the model performance is undetermined.
#### How to train the model?
Annotate a dataset, then upload it in JSON format. The annotation can be done before uploading the data or after using Google's AutoML UI, or the user can request annotation from Google's labeling service. The trained model is automatically deployed, and we can get its predictions using an API.
### Cloud AutoML Natural Language Sentiment Analysis
Enables you to create custom machine learning models to analyze attitudes.
According to Google's documentation, the current service supports sentiment analysis in English language text. We can train a custom model to classify text in other languages including Arabic, but the model quality may vary.
The sentiment score is an integer ranging from 0 (relatively negative) to a maximum value of your choice (positive). So, if we want to identify whether the sentiment is negative, positive, or neutral, we would label the training data with sentiment scores of 0 (negative), 1 (neutral), and 2 (positive).
#### How to train the model?
Build your own dataset, then upload it as .csv file. The trained model is automatically deployed, and we can get its predictions using an API.
### Cloud AutoML Natural Language **Pricing**
**Training Cost:** The cost of training a model is $3.00 per hour.
**Prediction Cost**: The usage of AutoML Natural Language is calculated monthly in terms of how many text records were sent for analysis during the billing month, as follows:
If a document contains more than 1,000 characters, it counts as one text record for every 1,000 characters.
| Feature | 0 - 30K | 30K+ - 5M+ |
| --- | --- | --- |
| **AutoML Natural Language Content Classification** | Free | $5.00 |
| **AutoML Natural Language Sentiment Analysis** | Free | $5.00 |
| **AutoML Natural Language Entity Extraction** | Free | $5.00 |
### Cloud AutoML Translation
Enables you to create custom translation models so that translation queries return results specific to a defined domain.
The supported language pairs can be found [here](https://cloud.google.com/translate/automl/docs/languages) which include Arabic to English (and vice versa) translation
#### How to train the model?
AutoML Translation trains custom models using matching pairs of sentences in the source and target languages.
The sentence pairs used to train the custom model must be in Tab-separated values (.tsv) or [Translation Memory eXchange (.tmx)](https://en.wikipedia.org/wiki/Translation_Memory_eXchange) format. A multiple .tsv and .tmx files can be batched into a comma-separated values (.csv) file. AutoML Translation uses the sentence pairs you provide to train, validate, and test the custom model.
The trained model is automatically deployed, and we can get its predictions using an API.
### Cloud AutoML Translation Pricing
**Training Cost**: The cost for training a model is $76.00 per hour, If training fails for any reason other than a user-initiated cancelation, you will not be billed for the time.
**Translation Cost**: Your usage of AutoML Translation is calculated in terms of how many characters you send for translation with an AutoML custom model.
| | 0 - .5 million characters | .5 - 5 million characters |
| --- | --- | --- |
| Translation | Free | $80 per 1,000,000 characters\* |
Price is per character sent for processing, including whitespace characters. Empty queries are charged for one character.
### Cloud AutoML Tables
Enables you to automatically build and deploy state-of-the-art machine learning models on structured data at massively increased speed and scale. Here are its features and capabilities:
#### Data support
Helps in creating clean, effective training data by providing information about missing data, correlation, cardinality, and distribution for each of your features.
#### Feature engineering
Automatically performs common feature engineering tasks, including:
- Normalize and bucketize numeric features.
- Create one-hot encoding and embeddings for categorical features.
- Perform basic processing for text features.
- Extract date- and time-related features from Timestamp columns.
#### Model training
Training for multiple model architectures at the same time. The model architectures AutoML Tables tests include:
- Linear
- Feedforward deep neural network
- Gradient Boosted Decision Tree
- AdaNet
- Ensembles of various model architectures
#### Model evaluation and final model creation
Using a validation set, determine the best model architecture for the data. After that two kinds of models are trained:
1. A model trained with the training and validation sets. this model is used to predict the test set targets to provide the evaluation of this model.
2. A model trained with the training, validation, and test sets. This is the model that is provided to be used to make predictions.
#### Supported Problem Types
- Regression problems
- Classification problems
### Cloud AutoML Tables Pricing
Prices for the usage of AutoML Tables are computed based on the underlying GCP resources required for model training, model deployment, batch prediction, and online prediction. You don't incur charges from AutoML Tables until you start training your model.
**Model training costs**: Model training costs $19.32 per hour of compute resources used to train the model.
**Model deployment costs**: Model deployment costs $0.005 per GiB per hour per machine that a model is deployed. They currently replicate the model to memory in 9 machines for low latency serving purposes, so there is a 9x multiplier applied to this cost.
**Batch prediction costs**: Batch prediction using the model costs $1.16 per hour of computing resources used.
**Online prediction costs**: Online predictions using the model cost $0.21 per hour of compute resources used.
## Conclusion
In this post, we talked about the services provided by Google's AutoML, including Cloud AutoML Natural Language, Cloud AutoML Translation, and Cloud AutoML Tables, their fitness for processing Arabic texts, and their pricing.
## Frequently Asked Questions
### What is Google AutoML?
Google AutoML is a suite of machine learning tools that enables developers with limited ML expertise to train high-quality custom models. It automates model architecture search, hyperparameter tuning, and training, making it accessible to build models for text classification, image recognition, and translation.
### Who should use AutoML?
AutoML is ideal for teams that need custom ML models but lack deep machine learning expertise. It is particularly useful for domain-specific tasks like classifying Arabic text, detecting product defects in images, or building custom entity extractors where pre-trained models fall short.
### How does AutoML compare to manual model training?
AutoML typically achieves competitive performance with significantly less engineering effort. While expert-tuned models may outperform AutoML on specific tasks, AutoML provides a strong baseline quickly and is often sufficient for production use cases where rapid iteration matters more than marginal accuracy gains.
### How to Fact-Check Using Natural Language Processing Techniques? A Literature Review
- URL: https://mohammadshaker.com/en/blog/how-to-fact-check-using-natural-language-processing-techniques-a-literature-revi
- Date: 2019-10-08T00:00:00.000Z
- Tags: almeta.io, fact-check, nlp, research, Human-Written
Automated fact-checking with NLP breaks into three sub-tasks: claim detection, evidence retrieval, and verdict prediction. Closed-source tools like FullFact cover well-known claims; open research systems use knowledge graphs, stance classifiers, and claim verification models. We survey the landscape and explain what each approach can and cannot handle.
#### Content
In this article, we present the summary of our research in the field of fact-checking. We categorized them in two categories, first are the closed source published applications and the second are the research projects done in this field.
## Closed Source
### **Snobs**
Their methodology depends on human annotators to fact check a piece of the news and present a detailed report regarding the inaccuracies in the article
### **Reporters’ Lab**
Their methodology depends on human annotators as well, and dataset [can](https://www.politifact.com/texas/) be found in and
### **Fullfact**
Their methodology builds a fully automated fact checker, but no details are provided about the model and the dataset.
## **Research Projects**
### Automatic Identification and Verification of Political Claims
**Methodology**
The model is composed of both convolutional neural networks and support vector machines. In order to get information to support or to refute a claim, they retrieved a number of snippets by querying Google. They did not select keywords but queried the search engine with full texts. The text of the claim of the most similar retrieved supporting texts were then fed into their model.
Another model they mentioned in their paper is the random forest model. In this case, both the Google and the Bing search engines were used to retrieve five snippets for a query consisting of the full claim. For each of the ten retrieved snippets, three features were computed:
1- the similarity between the claim and the snippet, calculated using word2vec embeddings,
2- the similarity between the claim and the snippet, calculated over the tokens, and 3-the Alexa rank of the website. These features were also combined, considering their mean and standard deviation.
The third method also retrieved supporting documents from the Web; in this case, they went further in trying to find the relevant fragments within the retrieved documents. Rather than using all the contents, they first compute the similarity between the claim and each sentence in the document and then they select those that pass a given threshold. The features for the supervised model are aggregations of the ones computed for each claim–sentence pair and include the stance of the sentence with respect to the claim and the degree of contradiction between the claim and the sentence, calculated at the term level.
The final method opted for an attention-based bidirectional long short-term memory network. Different from the previous approaches, in this case, no external information (e.g., no supporting documents) was used at all. Only the embedding representations of the claim itself were considered.
The best accuracy registered in this paper was for the first method.
**Dataset**
They produced the corpus CT-FCC-18 that includes claims from the 2016 US Presidential campaign, political speeches and a number of isolated claims. In order to derive the annotation, they used the publicly-available analysis carried out by FactCheck.org
### **ClaimRank: Detecting Check-Worthy Claims in Arabic and English**
**Methodology**
In order to rank the English claims, they used a neural network with two hidden layers. They provide the features, which give information not only about the claim but also about its context, as an input to the network. The input layer is followed by the first hidden layer, which is composed of two hundred ReLU neurons. The second hidden layer contains fifty neurons with the same ReLU activation function. Finally, there is a sigmoid unit, which classifies the sentence as check-worthy or not. Apart from the class prediction, they also need to rank the claims based on the likelihood of their check-worthiness. For this, they use the probability that the model assigns to a claim to belong to the positive class. They train the model for 100 iterations using Stochastic Gradient Descent.
For Arabic claims, First, they had to add a language detector in order to use the appropriate sentence tokenizer for each language. For English, NLTK’s sent\_tokenize handles splitting the text into sentences. However, for Arabic, it can only split text based on the presence of the period (.) character. Next comes tokenization. For English, they used NLTK’s tokenizer (Bird et al., 2009), while for Arabic they used Farasa’s segmenter. They further needed a part-of-speech (POS) tagger for Arabic, for which they used Farasa, while they used NLTK’s POS tagger for English.
**Dataset**
The run-time model is trained on seven English political debates and on the Arabic translations of two of the English debates. For evaluation purposes, they needed to reserve some data for testing, and thus the model is trained on five English debates and tested on the other two (either original English or their Arabic translations.)
##
**References**
1- [Snobs fact checker](https://www.snopes.com/fact-check/)
2- [Reporters lab fact checker](https://reporterslab.org/fact-checking/)
3- [Full Fact fact checker](https://fullfact.org/)
4- [Lab on Automatic Identification and Verification of Political Claims](https://arxiv.org/abs/1808.05542)
5- [Detecting Check-Worthy Claims in Arabic and English](https://arxiv.org/abs/1804.07587)
## Related Questions
### Q: How does automated fact-checking work with NLP?
Automated fact-checking uses NLP to verify claims by retrieving supporting or refuting evidence from the web, computing text similarity between claims and retrieved documents, and classifying the stance of evidence relative to the claim. Methods range from neural networks analyzing claim text alone to systems that query search engines for corroborating snippets.
### Q: What is claim detection in fact-checking?
Claim detection is the task of identifying which statements in a text are check-worthy—that is, which sentences contain verifiable factual claims worth investigating. Systems like ClaimRank use neural networks trained on political debate transcripts to rank sentences by their likelihood of being check-worthy.
### Q: Can NLP fact-check claims in Arabic?
Yes, researchers have adapted fact-checking pipelines for Arabic by integrating Arabic-specific NLP tools such as Farasa for tokenization and POS tagging. Systems like ClaimRank support both English and Arabic claim detection, though Arabic fact-checking datasets remain scarce compared to English resources.
### Event Detection in Media Using NLP and AI
- URL: https://mohammadshaker.com/en/blog/event-detection-in-media-using-nlp-and-ai
- Date: 2019-09-30T00:00:00.000Z
- Tags: arabic, ml, nlp, Human-Written
Event detection in NLP automatically identifies real-world occurrences in news text — who did what, where, and when. Document-level approaches cluster articles by topic; sentence-level approaches extract ACE-style event triggers and arguments. Both are foundational for news aggregators, fact-checking systems, and media monitoring pipelines.
#### Content
News stories are created every day at many news agencies. Users may receive news streams from multiple sources. Browsing in large-scale information spaces without guidance is not effective.
Suppose, for example, a person who has returned from a long vacation and wants to find out what happened during the period. It is impossible to read the whole news collection and it is unrealistic to generate specific queries about unknown facts. As a result, it is difficult to retrieve or to check all the potentially relevant stories.
Thus, it is useful to have an intelligent agent to automatically locate related stories in the continuous stream of news articles.
This is our first article on the event detection task, in the rest of these articles we will discuss:
- [How can we process large streams of data using sequential clustering? and how can it be used to handle the task of event detection?](/blog/an-implementation-of-a-news-stream-sequence-clustering-algorithm)
- [How can we extract detailed events from text?](/blog/an-overview-of-the-event-extraction-task-in-nlp)
## What is Event Detection?
Event detection is the task of detecting related stories from a continuous stream of news. Where the event is identified as something happening in a certain place at a certain time.
## News Event Analysis
Let's mention some of the news events properties that will help us in modeling the problem:
1. News stories discussing the same event tend to be temporally proximate.
2. A time gap between bursts of topically similar stories is often an indication of different events.
3. A significant vocabulary shift and rapid changes in term frequency distribution are typical of stories reporting a new event.
4. Events are typically reported in a relatively brief time window (e.g. 1-4 weeks) and contain fewer reports than broader topics.
## Modeling The Problem
In order to compare news articles and get to know which of them are more similar to each other, the first obvious step is to define a way to measure the distance between the articles.
*What do we need to measure a distance?*
- Representing the articles as vectors.
- A function that takes two of those vectors and outputs a real non-negative number that reflects how close those vectors are to each other.
### How to Represent News Articles for Distance Measurement?
We can't define a distance between the news articles in their typical unstructured form. Let's explore the ways used for transforming a news article into the vector space.
#### Traditional Representation
This representation involves the term vectors or bag of words, whose entries are nonzero if the corresponding terms appear in the document. Each term in the vector is typically weighted using the classical term frequency-inverse document frequency ([TF-IDF](https://en.wikipedia.org/wiki/Tf%E2%80%93idf)) approach.
However, according to the third news event property that we mentioned in the previous section, a modification was made to the standard TF-IDF term weighting when applied in this domain, which is using an adaptive IDF instead of the static one.
#### 5Ws1H Representation
This approach tries to take advantage of the structure of a news article. 5W1H is a technique employed in journalism to gather all information about a story, to turn it into a news article. It consists of six questions: *What*, *Who*, *Where*, *When*, *Why*, and *How*. An article is not considered complete until all of these six questions are answered.
It seems that little agreement could be reached on what to consider part of *What* or *Why*, or even how to represent *How*, but *Where*, *Who* and *When* are more concrete and could give better results.
**Who and Where**can be represented as Named Entities (Proper Nouns, Organizations, Locations, etc.) that appears in the text.
**Time Information (When)** represented by the publish date of the article.
**What** was represented by some works as topical information, namly the output vector of the [LDA](https://en.wikipedia.org/wiki/Latent_Dirichlet_allocation) algorithm.
### How to Measure the Distance Between The Articles Vectors?
Typically, the distance between the vectorized information is measured using traditional metrics such as:
- [The Euclidean distance](https://en.wikipedia.org/wiki/Euclidean_distance)
- [Pearson’s correlation coefficient](https://en.wikipedia.org/wiki/Pearson_correlation_coefficient)
- [Cosine similarity](https://en.wikipedia.org/wiki/Cosine_similarity)
- [Hellinger distance](https://en.wikipedia.org/wiki/Hellinger_distance)
Another option that can be considered is to measure the intersection between entities and/or the time zone between the publish dates separately.
## How to "Detect The Events"?
We differentiate between different approaches in terms of supervised and unsupervised learning approach.
### Unsupervised Learning Approach
The most common approach in event detection, which doesn't need annotated datasets.
#### Document-Pivot Approach
Typically, the event detection problem is treated as a stream clustering problem that aims to recognize patterns in an unordered, infinite and evolving stream of observations i.e. news stories.
As documents arrive the system must provide decisions (new or old event). Therefore, the employed clustering approaches are typically based on incremental (greedy) algorithms that process the input streams sequentially and merge an event with the most similar ones to it or create a new cluster if the similarity exceeds a predefined threshold.
#### Feature-Pivot Approach
This approach relies on keywords clustering instead of document clustering. And was found to be more effective in the social media domain.
This method is based on the assumption that documents describing the same event will contain similar sets of keywords. The solutions steps are as follows:
1. Build a network of keywords based on their co-occurrence in documents.
2. Use community detection methods analogous to those used for social network analysis to discover and describe events.
3. Constellations of keywords describing an event can be used to find related articles.
### Supervised Learning Approach
Treating the problem of event detection as a supervised learning problem which needs annotated datasets is less common. However, we're discussing it as a solution in this section.
#### Data Annotation
For a given time window (e.g. a particular period) collect the news articles from specified web news portals. Each article should be assigned to a group of articles reporting the same event. If there are no previously annotated articles that are reporting the same event a new group should be created for the article.
#### Auto Tagging Methods
**Threshold-based**: Calculate the similarity between articles. Similar articles are those whose similarities exceeded a predefined threshold.
**Pair-classification**: Use a binary classifier that given two articles decides if they are similar or not.
## How to Solve Event Detection Across Different Languages?
A few works have been considered the problem of language-independent event detection. For instance, in [1] at first, they clustered the articles in each language separately.
To perform the online clustering, they measured the similarities between the articles using their entities and non-entities weighted by TF-IDF, along with their publication dates.
They were interested in determining where the event took place exactly as an important entity to answer *where*. To achieve this they considered the problem as a word-level classification problem, where each location mentioned in the article is a candidate. An SVM classifier was used to perform the classification.
After all, they tried to identify the clusters in different languages that are discussing the same event. To perform the task they represent it as a learning problem. From the two clusters under consideration, they extracted a set of learning features that can be used for training a pair-classification model. These features included:
- Identification of concepts done by wikification, which is a process of entity linking that uses Wikipedia as the knowledge base. Each mentioned concept is annotated with a URI that is the link to the corresponding Wikipedia page.
- Whether the event locations found for the two clusters are the same or not.
- The absolute difference in hours between the events in the two clusters.
- The similarity of the dates that are being mentioned in the articles in the two clusters.
- The categories of the news articles in the two clusters based on their content, they categorized the news articles into a [DMOZ](https://en.wikipedia.org/wiki/DMOZ) taxonomy.
Another work to consider in this domain is [2] they developed four NLP pipelines for event detection in English, Spanish, Dutch and Italian. The pipelines aimed to identify who did what, when and where by adopting a common semantic representation.
Semantic interoperability across the four languages was achieved by projecting entities, event predicates and roles, time expressions and concepts to language-neutral semantic resources. In order to achieve semantic interoperability, event information from multilingual sources, entity and event mentions were projected onto language-independent knowledge representations.
- named entities were linked to English DBpedia entity identifiers through cross-lingual links existing to the Spanish, Italian, and Dutch DBpedia.
- Nominal and verbal event mentions were aligned to abstract representations through the Predicate Matrix [3].
- Time expressions were all normalized to the ISO time format.
## Conclusion
Event detection is the task of detecting related stories from a continuous stream of news. The main key to solve this problem is to represent each news story with features that describe the reported events effectively. The provided solutions include answering the questions about who, where, when, what, and how from the story. Machine Learning methods including supervised and unsupervised approaches were applied to solve this problem.
This was our first article on the event detection task, in the rest of these articles we will discuss:
- [How can we process large streams of data using sequential clustering? and how can it be used to handle the task of event detection?](/blog/an-implementation-of-a-news-stream-sequence-clustering-algorithm)
- [How can we extract detailed events from text?](/blog/an-overview-of-the-event-extraction-task-in-nlp)
## References
[1] Leban, Gregor, Blaz Fortuna, and Marko Grobelnik. "Using News Articles for Real-time Cross-Lingual Event Detection and Filtering." *NewsIR@ ECIR*. 2016.
[2] Agerri, Rodrigo, et al. "Multilingual event detection using the NewsReader pipelines." *de Castilho RE, Ananiadou S, Margoni T, Peters W, Piperidis S, editors. LREC 2016 Workshop. Cross-Platform Text Mining and Natural Language Processing Interoperability; 2016 May 23; Portoroz, Slovenia.[place unknown]: LREC; 2016. p. 42-6.*. International Conference on Language Resources and Evaluation (LREC), 2016.
[3] De Lacalle, Maddalen Lopez, Egoitz Laparra, and German Rigau. "Predicate Matrix: extending SemLink through WordNet mappings." *LREC*. 2014.
## Further Reading
[1] Yang, Yiming, et al. "Learning approaches for detecting and tracking news events." *IEEE Intelligent Systems and their Applications* 14.4 (1999): 32-43.
[2] Atefeh, Farzindar, and Wael Khreich. "A survey of techniques for event detection in twitter." *Computational Intelligence* 31.1 (2015): 132-164.
[3] Parafita Martínez, Álvaro. "News similarity with natural language processing." (2016).
[4] Sayyadi, Hassan, Matthew Hurst, and Alexey Maykov. "Event detection and tracking in social streams." *Third International AAAI Conference on Weblogs and Social Media*. 2009.
[5] Kumaran, Giridhar, and James Allan. "Text classification and named entities for new event detection." *Proceedings of the 27th annual international ACM SIGIR conference on Research and development in information retrieval*. ACM, 2004.
[6] Kumaran, Giridhar, and James Allan. "Using names and topics for new event detection." *Proceedings of the conference on Human Language Technology and Empirical Methods in Natural Language Processing*. Association for Computational Linguistics, 2005.
[7] Hogenboom, Frederik, et al. "An Overview of Event Extraction from Text." *DeRiVE@ ISWC*. 2011.
[8] Li, Zhiwei, et al. "A probabilistic model for retrospective news event detection." *Proceedings of the 28th annual international ACM SIGIR conference on Research and development in information retrieval*. ACM, 2005.
[9] Lam, Wai, et al. "Using contextual analysis for news event detection." *International Journal of Intelligent Systems* 16.4 (2001): 525-546.
[10] Edouard, Amosse. *Event detection and analysis on short text messages*. Diss. 2017.
[11] Brants, Thorsten and Francine Chen. “A System for new event detection.” *SIGIR* (2003).
[12] Rafea, Ahmed, and Nada A. GabAllah. "Topic Detection Approaches in Identifying Topics and Events from Arabic Corpora." *Procedia computer science* 142 (2018): 270-277.
## Frequently Asked Questions
### What is event detection in NLP?
Event detection is the task of automatically identifying and classifying occurrences of real-world events from text data. It involves detecting event triggers (words that indicate an event), classifying event types, and extracting event arguments such as participants, time, and location.
### How does news clustering work?
News clustering groups related news articles about the same event together using similarity measures. Algorithms compare article content, entities mentioned, timestamps, and semantic embeddings to determine which articles cover the same story, enabling users to see all coverage of an event in one place.
### What is the difference between event detection and event extraction?
Event detection identifies whether an event occurred and classifies its type, while event extraction goes further by identifying all participants, attributes, and relationships associated with the event. Detection answers "what happened?" while extraction answers "who, what, when, where, and how?"
### Top 3 Exciting Ideas in NLP in 2018
- URL: https://mohammadshaker.com/en/blog/top-3-exciting-ideas-in-nlp-in-2018
- Date: 2019-08-04T00:00:00.000Z
- Tags: ideas, nlp, Human-Written
Three NLP ideas from 2018 reshaped the field: BERT's bidirectional pre-training that reads context from both directions, the SWAG benchmark for commonsense reasoning across 113k question pairs, and LISA — a model that jointly learns syntactic structure and semantic role labeling in a single pass. Each attacked a different gap between machine and human language understanding.
#### Content
“Machine learning and natural language are the foundation to any AI system, just in the ability to communicate with us in a human way and to automate that learning process, what you build on top of that, whether it’s predictive, prescriptive analytics, forecasting, optimization, wherever you want to go, that foundation always comes back to these technologies that have been around for decades. ” this quote is said by [SAS](https://protect-us.mimecast.com/s/QXw8CZ6wWJf5nj0Qxtz2uqc?domain=sas.com) Artificial Intelligence and Language Analytics Strategist Mary Beth Moore. That said, several research breakthroughs in 2018 made astonishing improvement in NLP. In this article we give a glance about the top 3 sophisticated language models and new approaches in NLP.
1. **BERT**
BERT is short for **B**idirectional **E**ncoder **R**epresentations from **T**ransformers, it is a new pre-trained cutting edge NLP model that gave new impressive results in solving NLP tasks such as question answering, named entity recognition and language inference. Unlike the other pre-trained language model like OpenAI GPT and ELMO, BERT is designed with bidirectional Transformer to train on each word from both sides left and right. The following figure shows the difference among the three architectures.

BERT bidirectional model avoids the issue of cycles where words can be repeated because it is trained by randomly masking a percentage of input tokens, it can also understand the relationships between sentences by pre-training a sentence relationship model. BERT is considered to be a new era in NLP and it can be used in applications chatbots and customer reviews analysis.
2. **SWAG:**
When a person reads “He started his car” he or she is able to anticipate the rest of the sentence which might be “and he drove away”. Unlike humans, machines are not able to continue this obvious and easy sentence as it requires reasoning and commonsense. SWANG is short for Situations With Adversarial Generations and its goal is to enhance the research field of Natural Language Inference (NLI). SWANG is introduced as a large scale dataset that contains 113k questions about a wide range of commonsense reasoning situations and collected using video captions. The dataset is built by following these steps:
1. extracting a sentence from a video caption.
2. Extracting the correct answer from the next video caption.
3. Generating wrong answers by generating a huge set of wrong answers, picking the most related one statistically and finally filtering the endings that looks like it is generated by the computer and replace those endings with more human like endings, this way of generating answers is called Adversarial Filtering (AF).

The previous figure gives an example of how SWANG works. SWANG accuracy is relatively high as it scored an accuracy of 86.2%, while the human accuracy is as high as 88%. This model can improve commonsense reasoning in question and answer systems and chat bots.
3. **LISA**
LISA is short for Linguistically-Informed Self-Attention and it is a neural network model designed to extract the semantic role labeling using deep learning and linguistic formalism. For example, the sentence “Matt gave the instructions to Kim” the model should recognise the verb “gave to” as the predicate, “Matt” as the supervisor or the person who gave the instructions, “the instructions” as the theme and “Kim” as the recipient.
The neural network takes as an input word embeddings in addition to task specific learned parameters and train with multi-head self-attention with multi-task learning. Unlike the previous semantic role labelling models, LISA consume trivial pre-processing as it can add syntax using only raw tokens input, encode the sequence and then perform parsing, predicate detection and role labelling. LISA is used in automatic summarization, machine translation and Q&A systems, and perform very well in analyzing writing styles in newswires, journals and fictional writing.
##
## Refrences:
[1- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding](https://arxiv.org/abs/1810.04805)
[2- SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference](https://arxiv.org/abs/1808.05326)
[3- Linguistically-Informed Self-Attention for Semantic Role Labeling](https://arxiv.org/abs/1804.08199)
[4- 14 NLP Research Breakthroughs](https://www.topbots.com/most-important-ai-nlp-research/)
### Biggest Challenges in Arabic Natural Language Processing
- URL: https://mohammadshaker.com/en/blog/biggest-challenges-in-arabic-natural-language-processing
- Date: 2019-08-01T00:00:00.000Z
- Tags: arabic, nlp, Human-Written
Arabic NLP is harder than most languages for three compounding reasons: rich morphology produces thousands of word forms from a single root, dialectal variation across 26 countries means Modern Standard Arabic models often fail on colloquial text, and annotated training data remains scarce compared to English. Each challenge amplifies the others.
#### Content
We have mentioned in previous blogs the significance of NLP and the wide range of applications where NLP is used. As the basic goal of NLP is to ease and simplify the communication between machines and humans, it is highly crucial to see how it will impact the lives of the people who speak, communicate and work with the 6th most spoken language in the world, the Arabic language. Arabic is a Semitic language that is spoken by approximately 420 million people in the world, in addition to that, Arabic is an official language in 26 countries and it is one of the 6th official languages of the United Nations. Arabic is morphologically rich and has many varieties, for example, there is the classical form of Arabic which is the language of the Quran (the Muslims holy book) and this is considered to be the most perfect form of Arabic, another variety is the modern standard Arabic which is the official language today and used in literature, education, books, media and other formal locations and situations and finally there are the Arabic dialects that are the everyday speech and they are different in each country. After the previous short introduction on the Arabic language, we will discuss in this article 3 of the most major issues in Arabic NLP.
1. Arabic orthography
The Arabic language alphabet consists of 28 letters, only three are long vowels (ا) pronounced (Alef), (و) pronounced as (Waw) and (ي) pronounced as (Ya’a). In addition to other nine vowels represented as characters (َ ُُ ِِ ً ٌ ٍ ّ ْ ). Arabic is also one of the languages where the shape of the letter can change according to how it is connected with the other letters. For example the letter (ت) (the letter ‘T’ in English) has three forms of writing: it is written as (ت) if it is located at the end of the word, (  ) if it is located at the middle of the word and (  ) if it is located at the beginning of the word. Arabic orthography is very important to consider in all NLP tasks and applications such as: tokenization and text to speech.
2. Arabic morphology
All the verbs in Arabic have a root from three or four letters which make Arabic a highly derivational language. Usually there is a template for Verbs derivation we can write that as verb=Root+pattern. The following table shows some examples of verbs in their past, present/future and commanding form derived from three and four letters roots.
| root | pattern | verb | Transliteration | meaning |
| --- | --- | --- | --- | --- |
| كتب | ي | ي+كتب=يكتب | yaktb | Future/present form from write |
| كتب | ا | ا+كتب=اكتب | Ektb | commanding form from write |
It is also very common in arabic to attach prefixes and suffixes to verbs and we can formulate that with the following equation New\_Verb=Prefix(es)+Verb+Suffix(es). The following table shows an example of inflection in Arabic.
| verb | New Verb | meaning |
| --- | --- | --- |
| يكتب | س + يكتب = سيكتب | He will write |
| يكتب | س + يكتب + ه = سيكتبه | He will write it |
Studying the Arabic language morphology is very important for NLP tasks such as morphological analysis and POS tagging.
3. Complex syntax
Arabic language is rich in vocabulary where each word can have several meanings. for example, "البيت كبير" means (the big house) the word كبير which means (big) can give the sentence a different meaning if we said "كبير القوم" which means (the man that is responsible for a group of people). The problem of having multiple word expressions in the Arabic language will influence applications such as text summarization and translation.
## References
[Challenges in Arabic Natural Language Processing](https://www.researchgate.net/publication/327753798_Challenges_in_Arabic_Natural_Language_Processing)
## Frequently Asked Questions
### Why is Arabic NLP harder than English NLP?
Arabic presents unique NLP challenges due to its rich morphology (a single root can produce hundreds of word forms), right-to-left script, optional diacritics that change meaning, and significant variation between Modern Standard Arabic and regional dialects.
### What is the difference between MSA and dialectal Arabic?
Modern Standard Arabic (MSA) is the formal written standard used in media and education, while dialectal Arabic refers to spoken varieties (Egyptian, Levantine, Gulf, etc.) that differ significantly in vocabulary, grammar, and pronunciation. Most NLP tools are trained on MSA and struggle with dialects.
### What tools exist for Arabic NLP?
Key tools include CAMeL Tools (NYU Abu Dhabi), Farasa (QCRI), Stanford Arabic parser, and AraBERT for transformer-based processing. For dialectal Arabic, MADAR corpus and AOC dataset provide dialect-specific resources.
### 4 Biggest Open Problems in NLP
- URL: https://mohammadshaker.com/en/blog/4-biggest-open-problems-in-nlp
- Date: 2019-07-26T00:00:00.000Z
- Tags: nlp, Human-Written
The four biggest open problems in NLP are natural language understanding, ambiguity resolution, training data scarcity, and semantic meaning extraction. Ambiguity alone covers lexical, syntactic, and referential confusion that models still struggle with. Despite LLM advances, these challenges remain unsolved research frontiers with no clean algorithmic fix.
#### Content
When was the last time you asked your Siri or Alexa to do something and they did not understand what you are saying? or they answered with something totally not related? Siri and Alexa are speech bots that rely basically on an artificial intelligence technology called NLP. If you want to find out more about NLP and what it can and can’t do continue reading this article.
NLP stands for Natural language processing which is defined as a branch of computer science and artificial intelligence concerned with assisting the computers to understand the human natural languages by analyzing huge amounts of natural language data. The NLP problems ranges from simple problems such as answering an enquiry on the web to very complex problems that requires terabytes of data for training, but How much can NLP really understand what humans says? And How long will it take until we have a normal conversation with a computer? In this article we are going to discuss 4 of the most challenging NLP problems:
### 1. Natural language ambiguity
In natural language, a word can have different meanings and the meaning of the word can be extracted from the context. For example, the sentence “A piece of cake” might mean that we are talking about a small portion of a birthday cake, on the other hand, it might mean that something is very easy to do. The humans don’t only use their knowledge of a language to decide the meaning of a piece of text but also consider several other factors such as desires, goals and beliefs to understand the text they are reading or listening to. For example, the sentence “I experienced a feeling I have never had before” might mean that the person experienced a very pleasant feeling or a very bad one and the meaning of this sentence depends on the personal emotions at that moment.
### 2. The lack of training data
One of the biggest challenges in NLP is the shortage of training data as each NLP model need to be trained on terabytes of data in order to be able to understand a specific language, model training is a complex topic which will be covered in another separate article. The lack of training data has several reasons: the first reason is that the language is a minority language which means that it is spoken by a minority of population such as Kurdish and Afrikaan. The second reason is the small amount of resources and text available on the web for example the Zulu language. Another reason for the lack of training data is missing the incentive to work on low resources languages either due to not available skills or the difficulty of the language as the case in Arabic language.
### 3. Spelling mistakes and entity extraction
Correcting misspelt words is an essential process in NLP as Misspellings are very frequent in human-computer interactions and it would be very hard to identify a misspelt entity (the noun in the phrase) in a text. For example: if a user wrote on a chatbot “Is it going to rain today in amestedam?”, it would be hard to identify Amsterdam as a location.
### 4. Semantic meanings extraction (this can be part of ambiguity)
The computer should not only understand the vocabulary of the text but it should also understand the semantic of the text. For example: in the sentence “John called his wife, and so did Sam” we don’t know if Sam called john’s wife of his own.
##
## References
1-[What](https://medium.com/datadriveninvestor/what-are-some-of-the-challenges-we-face-in-nlp-today-2e9d94da1f63) [are some of the challenges we face in NLP today?](https://medium.com/datadriveninvestor/what-are-some-of-the-challenges-we-face-in-nlp-today-2e9d94da1f63)
2- [The 4 Biggest Open Problems in NLP](http://ruder.io/4-biggest-open-problems-in-nlp/)
3- [Six challenges in NLP and NLU - and how boost.ai solves them](https://www.boost.ai/articles/six-challenges-in-nlp-and-nlu-and-how-boostai-solves-them)
## Frequently Asked Questions
### What are the main challenges in NLP?
The four biggest open problems in NLP are: natural language understanding beyond surface patterns, handling ambiguity and context in language, achieving robust cross-lingual transfer, and building systems that can reason about common sense knowledge that humans take for granted.
### Why is NLP ambiguity hard to solve?
Language ambiguity is challenging because the same words can have different meanings depending on context, cultural background, and speaker intent. Resolving ambiguity requires world knowledge, pragmatic reasoning, and understanding of conversational context that current models only partially capture.
### Is NLP a solved problem?
No. While large language models have made remarkable progress on benchmarks, fundamental challenges remain in true language understanding, factual reasoning, handling low-resource languages, and robustness to adversarial inputs. NLP continues to be an active area of research.
### Design Specialization from CalArts
- URL: https://mohammadshaker.com/en/blog/design-specialization-from-calarts
- Date: 2019-01-10T00:00:00.000Z
- Tags: personal, startups, Human-Written
Long time, no see. I'm working on something new I won't talk about yet. Though, I recently finished 5-course design program from CalArts on Coursera.
#### Content
Long time, no see. I'm working on something new I won't talk about yet. Though, I recently [finished](https://www.coursera.org/account/accomplishments/specialization/G6BG3CYVRC5P) 5-course design program from CalArts on Coursera. I was pretty determined, getting 100% across all 5 courses. I really had fun. Totally recommended. Here's the capstone project I did; a guideline for a startup: Damagule. [mshaker\_damagule\_brand](/blog-images/2019/01/mshaker_damagule_brand.pdf "mshaker_damagule_brand")
### الأسبوع 07: "لاااا بجد".. إذا لم تقرأ 100 ختم في رمضان فأنت لست بالشخص الجيد
- URL: https://mohammadshaker.com/en/blog/%d8%a7%d9%84%d8%a5%d8%b3%d8%a8%d9%88%d8%b9-07-%d9%84%d8%a7%d8%a7%d8%a7%d8%a7-%d8%a8%d8%ac%d8%af-%d8%a5%d8%b0%d8%a7-%d9%84%d9%85-%d8%aa%d9%82%d8%b1%d8%a3-5-%d8%ae%d8%aa%d9%85-%d9%81%d9%8a-%d8%b1
- Date: 2017-06-24T00:00:00.000Z
- Tags: weekly-review, personal, Human-Written
كان عنوان هذا البوست بداية "كيف تقرأ 140 كتاباً في السنة؟"
#### Content
كان عنوان هذا البوست بداية "**كيف تقرأ 140 كتاباً في السنة؟**"
وعدت قبل أسابيع بمنشور عن القراءة.. لكنني وأنا أسرد في كتابته، رأيت أن العنوان "**إذا لم تقرأ 5 ختم في رمضان فأنت لست بالشخص الجيد**" أفضل.. ستتعرف في أخر المنشور لماذا.. لن أذكر شيء عن إسبوعي الماضي (عكس المعتاد) لأنه كان إسبوعاً بليالي قدر وعبادة وعمل. ولذلك فبوست هذا الإسبوع عن القراءة فقط.. لنبدأ..
على مقولة [goodreads](https://www.goodreads.com/review/stats/11193121) فقد قرأت سنة 2016 الماضية، 141 كتاباً.  ( لتعرف أكتر الكتب إفادة من الكتب التي قرأتها السنة الماضية فهاذا منشور [هنا](https://mohammadshaker.com/2017/01/07/%D9%86%D9%87%D8%A7%D9%8A%D8%A9-%D8%B9%D8%A7%D9%85-%D9%85%D8%B9-%D8%A7%D9%84%D9%82%D8%B1%D8%A7%D8%A1%D8%A9-2016-a-year-of-reading/)) **أول شي: لماذا تقرأ؟** عبارات: "لأتغذى بالمعرفة"، "حب العلم"، "الشغف بالشيء" تبدو دراماتيكية ومكررة، وعلى الرغم من صحتها، فبالنسبة لي لا أظنها تكتنف ما في عقلي للجواب عن ماذا أقرأ. أظن أن الجواب الأقرب لي هو: "الفضول". حب التعرف على أشياء لا أعرفها، حب أن أجعل نفسي إنساناً أفضل حتى لو قرأت رواية خيالية أو كتاب عن الاقتصاد. **"هلق مالو عصر الكتب"** مؤخراً يوجد مناهضة للكتب وخصوصاً في "عصر التكنولوجيا". كالمقولة التالية: "مافي داع للكتب مزال فيك تقرا مقالات عن كلشي مؤخراً." بصراحة إن كنت تريد معرفة أخر شيء عن لغة برمجة ما، أو تريد تعلم "تكنيك" معين فلا داع لأن تقراً كتاباً عن نفس الموضوع من سنتين لأنه سيكون قديم. لكن في نطاق الأفكار فأظن أن الكتب ما زالت أفضل وسيلة للمعرفة. غالباً ما تكون الكتب أيضاً مركزة وعصارة تجارب وليست كـ "بلوغ - Blog" ينشره الكاتب مرة في كل إسبوع أو أكثر ولا يكترث لكثافة وعمق المادة التي يقدمها. برأيي لم تظهر وسيلة حتى الآن تقارع الكتب لتلقي الأفكار الجديدة أو لتغيير أفكارك أو لعصر التجارب في نطاق ضيق كما هي في الكتب. ربما تقول لي وسيلة التعليم كمشاهدة الفيديوهات عن مادة ما، أقول لك نعم، ولكن في النهاية حتى هذه الكورسات تطرح مراجع وكتب لتقرأها لتتعرف أكثر. **كيف تقرأ 140 كتاباً في السنة؟** 140 كتاب ما يعادل كتاب كل يومين ونصف. إتبع الطرق التالية فوراً لسرعة أكبر في القراءة: **أولاً**: يوجد فيديو لـ Tim Ferriss لتعليمك القراءة السريعة في 5 دقائق، شاهده [هنا](https://www.youtube.com/watch?v=jeOHqI9SqOI). **ثانياً**: بالنسبة لي أستخدم التابلت بشكل أساسي للقراءة. أيضاً أستخدم الموبايل في القطارات (إذا بتلاقي الموبايل صغير لينقرا فيه فكبّر الخط وحاجة مياعة.) لكن قبل أن أكمل، في بوست ماض نشرت التالي باللهجة العامية عن كيفية إدخال المعلومات لدماغنا بطريقة أسرع:
> #### **السؤال يلي دائماً بسألو لحالي هوي كيف أدخل معلومات لعقلي بطريقة أسرع. وقت تدرك إنو عقلك بيستوعب وبيعمل بسرعة أكتر بكتير من عيونك أو أذنيك أو فمك، فهاد بيستدعي تستكشف أكتر عن طريق تدخل فيها المعلومات لعقلك بطريقة أسرع. بالنهاية عيونك وأذنيك وفمك هنن وسط Medium فقط. هاد الوسط هوي يلي بيحدنا إنو نفهم ع بعض أسرع أو ننقل المعلومات بين بعض بطريقة أسرع. لو كان فينا نتفاهم فقط مخ لـ مخ (Mind to mind) فلا حدود لأديش ممكن نعرف عن بعض أكتر، نتعلم أكتر، نتحرك أسرع.**
بالاعتماد على ما قرأته الآن هل تخطر في بالك وسيلة لتسرع بها قراءتك من التابلت أو الموبايل؟ بصراحة جاء بحثي عن طريقتي السريعة في القراءة من فكرتي عن الوسط Medium الذي نستخدمه للقراءة. أردت بداية أن أحرز فهماً أكثر لما أقرأه بالإضافة إلى زيادة سرعتي في القراءة. غالباً ما يظن الجميع أن التناسب عكسي بين الاثنين (السرعة على حساب الفهم والعكس.) الشيء الذي أردته أن يقرأ أحدهم الكتاب لي وأن ألاحق بعيوني ما يقرأه لي. ببساطة أحضر النسخة الصوتية للكتاب. ضع سرعة القراءة على الضعف 2x (أو حسب الرغبة) ومن ثم إفتح الكتاب بنسخته الورقية أو الإلكترونية ولاحق بعيونك ما تسمع. ستلاحظ مدى الفرق في زيادة الفهم وفي سرعة القراءة (بتقلي: طيب ما بقدر صدّق إنو إذا زدت السرعة بفهم أكتر. بقلك هاد شي مثبت علمياً وحديث لوقت آخر إلو علاقة بالـ stressors.) جربها ولاحظ أديش لح تكون أسرع وأديش لح تكون فهمان يلي قريتو بشكل أكبر. هل يوجد طريقة أسهل من تشغيل ملف صوتي وملف للقراءة بالطبع. وأنا لا أستخدم الطريقة التي ذكرتها سابقاً بكثرة أبداً لغلاظتها في كل مرة تريد قراءة كتاب (أنا من نمط الأشخاص الذين يقرؤون عدة كتب في نفس الوقت.) ببساطة حسب الجهاز الذي لديك
- iPhone/iPad فالطريقة سهلة جداً. يمكن جعل جهاز الـ Apple الخاص بك يقرأ أي شيء على الشاشة، كتاباً كان أو صفحة ويب أو أي شيء أخر. هذا ما أفعله أنا. الجميل في الأخر إنه يمكنك زيادة سرعة القراءة أو إبطائها كما أردت. يمكنك أيضاً اختيار صوت الشخص الذي تحبه مع اللهجة التي تحبها (بريطانية، أمريكية.. إلخ)
- إذهب إلى Accessibility > Speech > Speak Screen.
- يمكنك الأن قراءة أي شيء على الشاشة في أي وقت بـ Swipe بإصبعين من أعلى الشاشة
- أقرأ بجنون مع فهم!
- الأمور أغلظ قليلاً في الـ Android (جربته منذ سنتين على جهازي المحمول الماضي)
انظر أكثر [هنا](http://www.guidingtech.com/31832/best-apps-voice-reading-text-ios-android/): يمكن أيضأ على الأندرويد أو الأيفون استخدام تطبيق الكتب بتحميل أي كتاب لـ Play Book وبعدها استخدام خاصية Text to Speech لكي تقرأ لك. أخر مرة استخدمتها في فرنسا وكان يجب عليك أن تترك جهازك غير مقفل unlocked والشاشة في وضعية العمل. الصوت كان غليظاً قليلاً ولكنه "مقبول". لا أعرف ما الوضع الحالي له ولكن أتوقع أنه حتماً أفضل.
- إذا كان لديك Amazon Kindle فالأمور تشبه جداً كما في أجهزة Apple حيث يمكنك أيضاً جعل جهازك يقرأ لك نص الكتاب.
**كيف تقرأ مقالات كثيرة في السنة، بسرعة وبفهم؟** المقالات والـ Blogs غالباً ما أقرأها على الحاسب. (أحياناً حتى بعض الكتب على الحاسب.) فكيف يمكنك استخدام نفس الأساليب السابقة (أسلوب السماع والملاحقة بالعينين) في الحاسب؟ إذا كنت تقرأ على الحاسب مقالة يمكنك استخدام [الإضافة Chrome Speak على متصفح الـ Chrome](https://chrome.google.com/webstore/detail/chrome-speak/diagnfimeecdcecjpnkjgbnlelkclcpj).
- فقط افتح أي صفحة
- حدد ما تريد قراءته
- زر يميني بالماوس ومن ثم Chrome Speak وستقرأها لك
ملاحظات ع جنب:
- يمكنك تعديل السرعة من More Tools > Extensions > Chrome Speak
- بكل بساطة وبنفس الأسلوب إذا كنت تريد قراءة **كتاب** على الحاسب، فقط افتح الكتاب عن طريق متصفح الكروم وليس عن طريق الـ Adobe Reader أو غيرها. حدد ما تريد قراءته ضمن المتصفح.. وأكمل الخطوات..
- شاركتني [Naya Hafez](https://www.facebook.com/Qamar.Alzaman.H?fref=ufi) بمعلومة عن متصفح Microsoft edge وهي: أي ملف من نمط epub بيقراه Microsoft edge مع كل ال features اللي بتعملها Kindle متل ال (size- space -font-theme ...etc ) بالاضافة للقراءة و تسريعها.
**من أين جئت بالـ 140 كتاب؟** الكتب التي أقرأها: ١- صوتية فقط. أي أنني أسمع الكتاب فقط بدون قراءة فعلية. حوالي 30% من الـ 140 هي كتب صوتية في السنة الماضية. غالباً الكتب الأخف أسمعها سماعاً ولا أقرأها. غالباً وقت المواصلات أو الانتظار يكون للكتب الصوتية. ستستغرب من حجم الكتب التي تستطيع سماعها فقط عن طريق المواصلات (ضع السرعة x2 وستنتهي من كتاب صغير في يوم واحد!) ٢- قراءة فقط. أي لا يوجد شيء أسمعه إما لعدم وجود التسجيل الصوتي لدي أو لأنني أريد أن أقرأه بسرعة أنا أريدها. (مثل شبك المعلومات مع بعضها البعض وتقليب الصفحات والذي يجعل السماع أمراً عسيراً في كل مرة تريد أن تقف على فقرة لبعض الوقت.) ٣- صوتية وملاحقة عينين (قراءة) كما ذكرت سابقاً. **هل يجب أن تقرأ 140 كتاباً في السنة؟** لا أظن أن العدد 140 يحمل أي قيمة إذا أكملت وقرأت التالي.. **كلمة أخيرة قبل عيد جديد..** لو كان عنوان البوست "كيف تفهم كتب أكثر؟" هل تظن أنك ستقوم بقراءة البوست؟ لاحظ أن التركيز في أغلب البوست هو على العدد وليس الكمية وهذا ما كنت أحاربه في نفسي خلال السنتين الماضيتين: ركز على النوعية وليس على الكمية. ربما ستقول أن هذا شيء مكرر، منعاد ١٠٠ مرة وكلنا منعرفو.. أغلب الـ 140 كتاب التي قرأتها في السنة الماضية 2016 كانت في النصف الأول من السنة. في النصف الآخر قررت أن أعيد بحبشبة بعض الكتب التي قرأتها سابقاً في حياتي ووجدتها رائعة. فقط مرور على الـ Notes التي وضعتها على الكتب التي أعجبتني. دهشت من حجم المعرفة التي نسيتها مع أنني وضعت العديد من الملاحظات عليها. أردت أن أدوّن وأذكّر نفسي بجميع الأمور التي تجعلني إنساناً أفضل أمام ربي، أمتي ونفسي. وهذه كانت بداية [لا تفكر في حليب أحمر](http://mohammadshaker.com/2017/04/22/%d9%83%d8%aa%d8%a7%d8%a8-%d9%84%d8%a7-%d8%aa%d9%81%d9%83%d9%91%d8%b1-%d9%81%d9%8a-%d8%ad%d9%84%d9%8a%d8%a8-%d8%a3%d8%ad%d9%85%d8%b1/) بدون أن أعلم.. في السنتين الماضيتين حاولت بجد أن تكون حياتي مليئة بالنوعية وليس بالكم | العدد. من شراء قطعة ملابس لقراءة الكتب لتناول طعام لعلاقات مع الناس. حتى في الدين والتدين والصلاة والقيام وقراءة القرآن. "تدبر ساعة خير من عبادة سنة". وهذا ما ينقلنا لعنوان المنشور.. ما يدعى إليه الآن من إكثار للختم ليس بالأمر السيء ولكن في حدود المعقول وفي حدود أن تكون القراءة بتفكر وبتدبر. جمل وعبارات من أمثال "إذا لم تقرأ ٥ ختم فأنت لست بالشخص الجيد" هي جنون. نعم جنون. فمن أنت لتحكم على أي شخص بأنه جيد أو لا. لماذا تضع نفسك في موضع الله عز وجل؟ لماذا التّألّه على الله والعياذ بالله؟ إذا كانت الـ ٥ ختم بتدبر وتعقل فيالها من قراءة وعبادة وتفكر. ولكن التركيز فقط على الكم|العدد في الدين بدون النوع هي دائماً مشكلة تقع وتكرر.. الأسوأ هو الحكم على الأشخاص "بعدد" الختم المقروءة! "تدبر ساعة خير من عبادة سنة" "تدبر ساعة خير من عبادة سنة" "تدبر ساعة خير من عبادة سنة" "تدبر ساعة خير من عبادة سنة" لن أطيل أكثر، وبعيداً عن الدين، لا تكون الأمور سهلة دائماً لأن عقلنا دائماً يميل للكم وليس للنوعية في أغلب الحال (ليس من الأفضل أن تصلي 50 ركعة، أقرا 5 ختم للقرآن الكريم فقط عد صفحات.. إلخ). محاربة هذا الشيء يتطلب تذكير نفسك بأن نوعية حياتك هي من نوعية القرارات التي تتخذها.. الـ 140 كتاب لا أعتبرها كمية، لأنني دائماً أذكر نفسي بالنوعية.. وإذا كانت الـ 140 أو الـ 200 أو الـ 20 تلك هي التي سأقرأها هذه السنة وذات نوعية عالية فيا مرحباً بالكم إذا كان ذا نوع راق. ما أركّز عليه هنا هو أن غفلاني بملاحقة قراءة كتب جديدة عوضاً عن ربط كامل للكتب التي قرأتها سابقاً في "الكم" سنة الأخيرة هو ما اعتبره كمي.. دوماً إبحث عن النوع وليس الكم.. كلام لي ولك..
الآن، ومع نهاية شهر رمضان المبارك، يجب أن تكون بداية لشيء جديد كما كانت بداية لا تفكر في حليب أحمر، لذا سأغيب لبعض الوقت..
**كل عام وأنتم بألف خير**
**عيد سعيد جميل مثلكم**
### الإسبوع 06: عودة للشام، توت شامي، ناعم ودبس
- URL: https://mohammadshaker.com/en/blog/%d8%a7%d9%84%d8%a5%d8%b3%d8%a8%d9%88%d8%b9-06-%d8%b9%d9%88%d8%af%d8%a9-%d9%84%d9%84%d8%b4%d8%a7%d9%85%d8%8c-%d8%aa%d9%88%d8%aa-%d8%b4%d8%a7%d9%85%d9%8a%d8%8c-%d9%86%d8%a7%d8%b9%d9%85-%d9%88%d8%af
- Date: 2017-06-19T00:00:00.000Z
- Tags: weekly-review, personal, Human-Written
هذا الأسبوع عدت إلى الشام، سوريا بدون تخطيط. أمر طارئ وعلي أن أكون هنا لبعض الوقت. لحسن حظي أنه أيضاً أجمل وأحلى أيام السنة في دمشق.
#### Content
هذا الأسبوع عدت إلى الشام، سوريا بدون تخطيط. أمر طارئ وعلي أن أكون هنا لبعض الوقت. لحسن حظي أنه أيضاً أجمل وأحلى أيام السنة في دمشق. شهر رمضان والأسبوع الأخير من رمضان وليالي القدر. ربما تكون هذه المنشورات مقتضبة عما سبق. أرى أنه ليس من الفائدة في تكرار كل إسبوع أنني (أتحدى نفسي في المداومة على اليوغا) مثلاً. أظن عدم ذكر مثل هذا التكرار أكثر فائدة لكم ولعدم إضاعة وقتكم ووقتي. يمكنني ذكرها بعد ٣ أشهر من المداومة مثلاً وليس في كل إسبوع. وهكذا بالنسبة للبقية. لنبدأ.
**كتاب أقرأه حالياً**
- [Architect: The Work of the Pritzker Prize Laureates in Their Own Words](https://www.goodreads.com/book/show/8868729-architect): لاهتمامي القديم بالعمارة، أنهيت كتاباً أخر عنها يعرض العديد من أعمال الحائزين على جائزة [Pritzker](https://www.google.com/search?q=Pritzker&oq=Pritzker&aqs=chrome..69i57j69i61l3.260j0j7&sourceid=chrome&ie=UTF-8#q=Pritzker+award) في العمارة ومقابلات معهم. الكتاب مقبول، ضخم وطويل. فيه فائدة حتماً لغير الخبيرين مثلي. ولكن أتمنى أن يعرض عليي أحد مهندسي العمارة كتب Classics في العمارة لكي أستفاد أكثر. شاركونا وأرسلوا لي!
- [Antifragile](https://www.goodreads.com/book/show/13530973-antifragile): عدت لإكمال قراءة عدة صفحات. أذكره هنا لأنه للكاتب العربي اللبناني نسيم طالب (دكتور في جامعة نيويورك.) ربما سمعت عن كتابه الآخر البجعة السوداء [Black Swan](https://www.goodreads.com/book/show/242472.The_Black_Swan). كتبه (طويلة) وأفكاره مغايرة للبقية بطريقة تغير من طريقة نظرتك للأمور.
**اقتباس أطبقه | أفكر به** فقط تأمل بالآية الكريمة: "**لا ملجأ من الله الا اليه**" لن أستفيض الآن أكثر. **شيء أداوم على فعله** Duolingo كل يوم لتعلم لغة جديدة (أحافظ على الألمانية والفرنسية). وصلت إلى اليوم 85 يوم متتالي. أظن أن آخر حد وصلت إليه سابقاً كان 90 يوماً ولذا إن أتممت هذا الإسبوع.. **شيء جديد أجرّبه** ليس شيئأ جديداً ولكن في كل مرة أزور الشام ينبع "التخبيص في الأكل" بشدة. إحم إحم. **أفضل ما شاهدته**
[The Aviator](https://www.google.com/search?q=The+Aviator&oq=The+Aviator&aqs=chrome..69i57j69i60l3j69i64.144j0j7&sourceid=chrome&ie=UTF-8): بعد تزكية أختي [Noora](http://lynura.com/) من شي عشرين سنة قررت أن أرى هذه الفلم. لمفاجئتي كان عن القصة الحقيقية لـ [Howard Hughes](https://www.google.com/search?q=Howard+Hughes%3A+His+Life+and+Madness&oq=Howard+Hughes&aqs=chrome.1.69i57j69i59l2.3016j0j7&sourceid=chrome&ie=UTF-8) والذي أثرت [قصة حياته (كتاب)](https://www.google.com/search?q=Howard+Hughes%3A+His+Life+and+Madness&oq=Howard+Hughes&aqs=chrome.1.69i57j69i59l2.3016j0j7&sourceid=chrome&ie=UTF-8) على Elon Musk. بطولة ليوناردو ديكابريو من 2004. فلم رائع. **أفضل ما سمعته** فكرة ذكرتها كثيراً في [لا تفكر في حليب أحمر](http://mohammadshaker.com/2017/04/22/%d9%83%d8%aa%d8%a7%d8%a8-%d9%84%d8%a7-%d8%aa%d9%81%d9%83%d9%91%d8%b1-%d9%81%d9%8a-%d8%ad%d9%84%d9%8a%d8%a8-%d8%a3%d8%ad%d9%85%d8%b1/) وهي كيف تتخلص من مشكلة إختيار "قراراتك". يمكنك سماع الـ ١٥ دقيقة جيدة جداً [هنا](http://hwcdn.libsyn.com/p/f/4/1/f414924ce9344385/Short_-_Decision_Fatigue.mp3?c_id=7881402&destination_id=189941&expiration=1497693763&hwt=e6e4135ec2318cb4e6a3f5619ef83a91) لتيم فيريس. **أفضل شيء اشتريته مؤخراً** أضحك وأنا أكتبها الآن. ولكن البارحة وبما أني في دمشق فاشتريت "دلواً" من التوت الشامي. بصراحة وبحق كانت أفضل شيء اشتريته مؤخراً.. الشروة الأخرى كانت للناعم، وما ألذذذذذذذذذه. لمن لا يعرفه هذا هو الناعم مع الدبس، أكلة شامية مختصة برمضان المبارك. لا أعلم إن كانت موجودة في غير بلاد عربية؟ 
## **شيء خجلت به من نفسي**
معنى كلمة صلاة الجمعة. لأنها من يوم الجمعة. ولكن معنى الجمعة تأتي من الجماعة. يوم يجتمع فيه الناس. أظن أنني أعرف المعنى في عقلي الباطن ولكن عقلي لم يفهمها بشكل صريح واضح في الإسلام وعلى الأهمية التي أحلها الإسلام بها بتخصيص يوم باسمها. **شيء أحن إليه** منذ لحظة وصولي قبل أربعة أيام وأنا أذهب بعد المغرب لأصلي جماعة العشاء وما بعدها من صلوات حتى التهجد. شعور رائع، غريب، مريح، طيب يعتمرك عندما تكون في بلادك وتصلي هذه الصلوات التي اعتدت عليها قبل السفر، والتي انقطعت عنها في السفر. منذ ٥ سنوات هي آخر مرة صليت فيها العشاء جماعة في رمضان. **أكثر منشور محبوب على الـ instagram** [ركضة الساعة ١١ بالليل قبل النزلة عالشام](https://www.instagram.com/p/BVUfcLZl4qX/?taken-by=zgtrshaker) 🌸
ليلة الـ ٢٥ من رمضان المبارك
ليلة قدر مباركة
وليال مباركة عامرة بالعبادة والرضا والاستجابة والعفو والمغفرة
سأدعي لنا جميعاً
ادعولنا
### الإسبوع 04-05: أحبّك حبّين
- URL: https://mohammadshaker.com/en/blog/%d8%a7%d9%84%d8%a5%d8%b3%d8%a8%d9%88%d8%b9-04-05-%d8%a3%d8%ad%d8%a8%d9%91%d9%83-%d8%ad%d8%a8%d9%91%d9%8a%d9%86
- Date: 2017-06-12T00:00:00.000Z
- Tags: weekly-review, personal, Human-Written
السلام عليكم ورحمة الله وبركاته ليش متأخر؟ عفواً، ما عت عيدها، عفواً. الإسبوع الماضي (الإسبوع 04)، وأنا أكتب منتصف هذا المنشور، وبعد الإفطار ذهبت لأركض.
#### Content
السلام عليكم ورحمة الله وبركاته
**ليش متأخر؟** عفواً، ما عت عيدها، عفواً. الإسبوع الماضي (الإسبوع 04)، وأنا أكتب منتصف هذا المنشور، وبعد الإفطار ذهبت لأركض. المشكلة أنني تعثرت وجرحت يدي اليمنى بشكل كبير ولهذا لم أنشر شيئاً. ولكنني هنا الآن! لنبداً. **قبل البداية** لهذا الإسبوع، وقدماً، فسأكتب بالعربية الفصحى. لاحظت في الأسابيع الماضية أن العديد من زوار الموقع من خارج سوريا أيضاً كالجزائر، فلسطين، مصر والمغرب العربي. بعضهم أرسل لي خصيصاً لأن أكتب بالفصحى التي هي لغتنا الأم. كتابتي بالعامية الشامية سابقاً كان هدفها أن تكون المنشورات نابعة من القلب ولا تعطي أي إيحاء بالزخرفة أو "عرض عضلات" أو فلسفة ليس لها داع. مرحباً بالجميع من المحيط إلى الخليج. أغلب أصدقائي العرب (والرسائل التي وصلتني) أشارت إلى أنهم يحبون اللهجة الشامية ولكن لن تكون مفهومة لهم دائماً ولهذا السبب.. إلى الفصحى ولكن ليس دوماً. فليسامحني البعض إن كتبت بعضاً من "الغلاظات" بالشامية لأنّ للهجة الشامية أمثالاً لزيزة "مو طبيعية" :D أظنكم ستحبونها على كل حال. لنبدأ. ورمضان مبارك عامر بإذن الله.
## **ماذا أقرأ حالياً**
_Zorba the Greek_ بدأت هذه القصة بعد تزكيات كثيرة. حتى الآن لم تعجبني كثيراً. _رحلتي مع غاندي لأحمد الشقيري._ كتاب قصير، وجيد. ليس بخارق. أحببت أن أقرأ كتاباً للشقيري غير برامجه. على الرغم من اختلافي معه في العديد من وجهات النظر (مثل توجيه اللوم على الحكومات وعدم التركيز على الأفراد مع أن التغيير يبداً بالعكس) فهو شخص يحب التغيير ويقوم بالتغيير.
### **اقتباس أطبقه | أفكّر فيه**
ليس اقتباس وإنما الآية الكريمة: "فامشوا في مناكبها". دلالة من الله عز وجل كي نكون استباقيين، عاملين، جادين، Proactive في أي شيء نفعله. أن نذهب ونسافر ونعود ونعمر وننشأ ونعمل. ألّا نجعل الأمور تحدث لنا وإنما نحن نحدث الأمور. يوجد مقولة لستيف جوبز مغزاها:
> أن الحياة التي تراها الآن ما هي إلا نتيجة أعمال أشخاص مثلنا تماماً. حالما تدرك هذه الحقيقة فستدرك أيضاً أنه يمكنك تغيير Mold جميع الأشياء أمامك. يمكنك وخز الحياة والأشياء لترى ما سيظهر في الطرف المقابل. العبرة: هي أن لا تقف على الجانب. لا تكن عاجزاً. كن دائماً محدثا للأمور، لا أن تحدث لك الأمور.
#### **قيد أتحدى به نفسي**
20 hours of coding. لعدة أيام.
#### **شيء خجلت من نفسي به**
عدم إدراكي لمعنى كلمة إله. كلمة دلالة عن ربي الذي أعبده وأصلي له خمس مرات في اليوم. كلمة إله تأتي من الوله. من اللهيان. من شدة الحب والتعلق، وليس فقط ذلك "الذي يستحق أن يكون معبوداً." أظن أن الكثير من الخطاب الديني يركز كل التركيز على الحقوق والواجبات على العبد. نسينا الجانب العاطفي، الحساس، المحب، الوله،الرقيق الذي يربط العبد بربه، والمتضرع بالمعبود والولهان والعاشق بمحبوبه ربه. نسينا كل شيء عن تلك الغرسة التي غرسها الله فينا لكي نعود بشراً.
#### **أفضل ما سمعته**
[مقابلة تيم فيريس مع كيفن كيلي](http://tim.blog/2014/08/29/kevin-kelly/) على ثلاث حلقات. أحد مؤسسين Wired.com. بالكلام عن معنى كلمة إله فأيضاً تعرفت منذ فترة على فرقة إنشاد صوفية رائعة. [اسمع لرابعة العدوية "أحبك حبين" في عشقها لله](https://www.youtube.com/watch?v=n-Nx5JQkYH0)
#### **أفضل ما شاهدته**
حلقة الدكتور عمر خالد عن [وفاة أم المؤمنين سيدتنا خديجة، والأذى الذي تعرض النبي له في الطائف من أجلنا](https://www.youtube.com/watch?v=JpuRjdBbz_g)
#### **أفضل ما اشتريته**
التطبيق الذي تكلمت عنه الإسبوع الماضي [Freedom App](http://freedom.to) بدون إعطاء رأي. الآن يمكنني القول بأنه تطبيق جيد جداً. التطبيق يمنع Block Lists عن مواقع أو تطبيقات محددة مثل مثلاً تطبيقات الـ Social Media لمدة ٤ أو ٥ ساعات مثلاً. ستحصل عندها على ٤ أو ٥ ساعات من التركيز والنقاء وصفاء الذهن دون أي مقاطعات. التطبيق لا يتيح إمكانية إزالة المنع Unblock وبالتالي "ستعلق" في ٤أو ٥ ساعات من الراحة لا محالة. متاح للـ Desktop وللموبايل. وهو مدفوع. يوجد تطبيق ممثال كنت استخدمه سابقاً على الـ Windows تذكرته واسمه [Rescue Time](https://www.rescuetime.com/) يعطيك جداول وبيانات رائعة أيضأً عن كيفية "حرقك" لوقتك خلال الإسبوع ويمكنك عندها التحسين أكثر.
**أكتر منشورات محبوبة على الإنسغرام** بوست [الفوز بأول هاكاثون في Squla](https://www.instagram.com/p/BU2dQ1pliNN/)، بوست [شمس كالقمر](https://www.instagram.com/p/BU2I0LtlMfV/)، وبوست [ابنة أختي](https://www.instagram.com/p/BUvAGSHFD8c/)! **لغة أتعلمها** مازلت أداوم على تعلّم الألمانية ع Duolingo وعدت إلى تعلم الفرنسية أيضاً. إذا كنت تحب اللغات فأضفني [هنا](https://www.duolingo.com/Mohammad.Shaker)! **مذيبات الوقت لهذا الإسبوع** على مدار الساعة للمشروع الذي يمكن أن تروا شيئاً منه قريباً جداً. سيكون الأكبر تأثيراً لي لأمتي وللعرب للمستقبل. تحمّممممس! :D
**شكر خاص**  الإسبوعين الماضيين عاد كتابي [لا تفكر في حليب أحمر](http://mohammadshaker.com/2017/04/22/%d9%83%d8%aa%d8%a7%d8%a8-%d9%84%d8%a7-%d8%aa%d9%81%d9%83%d9%91%d8%b1-%d9%81%d9%8a-%d8%ad%d9%84%d9%8a%d8%a8-%d8%a3%d8%ad%d9%85%d8%b1/) ليظهر بقوة على وسائل التواصل الاجتماعي ووصلتني العديد من الرسائل بخصوصه. لا أعلم تماماً سبب الظهور مع أنني لم أنشر شيئاً بخصوصه منذ فترة. هدفي هو شكر كل من كتب لي وانتظرني طويلاً كي أجيبه بسبب جرح يدي. شكراً لكم جميعاً! وأعاننا الله على فعل كل خير! سيكون هناك نظرة خلف كواليس لا تفكر في حليب أحمر قريباً عند بزوغ "شقفة" وقت لدي. أرجو أن تكون قراءة قصيرة مفيدة ممتعة. أراكم! _بالمعنى الحرفي هذه المرة_ قريباً جداً!
ليالي القدر قادمة. وما أحلاها من ليال!
تقبل الله منا جميعاً
### الإسبوع 03: بوظة، نظارات ٢ يورو، بيتر، وماذا تعلمنا من قراءة ٥ مليون كتاب؟
- URL: https://mohammadshaker.com/en/blog/%d8%a7%d9%84%d8%a5%d8%b3%d8%a8%d9%88%d8%b9-03-%d9%86%d8%b8%d8%a7%d8%b1%d8%a7%d8%aa-%d9%a2-%d9%8a%d9%88%d8%b1%d9%88-%d9%85%d8%b9-%d8%a8%d9%8a%d8%aa%d8%b1-%d9%88%d8%a8%d9%8a%d8%aa%d8%b1-%d9%88%d8%a8
- Date: 2017-05-28T00:00:00.000Z
- Tags: weekly-review, personal, AI, startups, Human-Written
شهر رمضان مبارك طيب كريم! أفضل وأجمل وأبهى أشهر السنة! أتمنى أن يعيده الله علينا جميعاً بالأمن والأمان والعفو والمغفرة والجد والعمل بالعودة لمنشوراتي
#### Content
**شهر رمضان مبارك طيب كريم!**
**أفضل وأجمل وأبهى أشهر السنة! أتمنى أن يعيده الله علينا جميعاً بالأمن والأمان والعفو والمغفرة والجد والعمل**
بالعودة لمنشوراتي [الإسبوعية](https://mohammadshaker.com/category/weekly-assessment/)، هاد هوي الإسبوع الثالث. المفترض إنو يكون ثيم هذا الإسبوع هو عن القراءة ولكن لتعسر الوقت لح يكون للإسبوع القادم. ما بدي استعجل بكتابة بوست لح يكون غريب ومالو اعتيادي من هالشكل. وبالتالي لح يكون ثيم هذا الإسبوع هو سلطة (أو بوظة متل ما بتحب!)

**شو عم إقرأ حالياً**
- Things Hidden Since the Foundation of The World للكاتب الفرنسي رينيه جيراد. كتاب فيلسوفي قح رائع لدرجة الغرابة. بيتكلم بشكل أساسي عن شي اسمو Mimetic behaviour أو السلوك المحاكي | المقلّد. قريتو بعد تزكية كبيرة من [Peter Thiel](https://en.wikipedia.org/wiki/Peter_Thiel).
- بوست [Contrarian Strategy | الاستراتيجية المضادة](http://fortune.com/2014/09/04/peter-thiels-contrarian-strategy) مرة تانية لأفكار [Peter Thiel](https://en.wikipedia.org/wiki/Peter_Thiel) من الـ Paypal Mafia ومؤلف كتاب [Zero to One](https://www.goodreads.com/book/show/18050143-zero-to-one). بيتر وأفكاروا يمكن من أكتر الأفكار يلي أثرت عليي بالسنتين الماضيين من ناحية شغلي لقدام ع Startup / Business. هاد بوست بيحكي عن بعض الأفكار إلو. رائع بحق. وإذا عندك ٢-٣ ساعات فيك تقرا كتاب [Zero to One](https://www.goodreads.com/book/show/18050143-zero-to-one) لأنو كتاب بيخليك تفكر بطريقة مختلفة عن الـ Startups/ Business. هاد دائماً بيكون من أول الكتب يلي بقولها لأصدقائي وقت حدا بيسألني عن كتب للـ Startup/ Business.
**اقتباس عم طبقوا | فكّر فيه** "It's not about you." فكّر فيها قليلاً. بضل بفكّر بهي المقولة. عميقة جداً من إنو حياتنا هيي مو بس إلنا. ومو بس مشان نعمل الشي يلي نحنا بدنا ياه. **قيد عم إتحدى نفسي فيه** ضل مداوم ع رياضة كل يوم من الإسبوع. صرلي تقريباً فترة ال٣ السنين الماضية بشكل شبه منتظم. هلق عم حاول يكون دائماً بيومي فيه رياضة لو شو ما صار وشو ما كان يومي. شهر رمضان مبارك مالو حجة أبداً للقطع. في رياضة مساء بعد الفطور فيك تعملها أو بالبيت أو يوغا أو أو. في Apps عن هالشي بحكي عنون بإسبوع قادم. **شيء انبهرت فيه** [Webstorm](https://www.jetbrains.com/webstorm/specials/webstorm/webstorm.html?&gclid=CjsKDwjw6qnJBRDpoonDwLSeZhIkAIpTR8KAJM2tzgQhlmsn2KCZ-r-5n56ElUqucoU73UZRRZpOGgJUqfD_BwE&gclsrc=aw.ds.ds&dclid=CMCxkPvQk9QCFVSMdwody50Lvw) إذا كنت مطور ويب فبجد Webstorm (وخصوصاً للـ JS جيد جداً) بالعادة بشتغل ع PyCharm بس للـ JS الـ Webstorm جربتو كتير جيد. مالو مجاني. لهلق لساتني نسخة تجريبية عليه (هل حدا عندو رأي آخر ع IDE أفضل أو مجاني؟) **شيء جديد عم جربه** [Freedom App](http://freedom.to) مالح إحكي عنو أكتر لحتى يطلع معي رأي مناسب. أول شي كتير كرهت طريقتو. بس هلق رد عجبني. هو تطبيق بيمنع Block Lists إنت بتحددون وبالتالي مثلاً فيك تقلّوا مناع الـ Social Media Websites لمدة ٤ ساعات. ما فيك تغير. متاح للابتوب وللموبايل. مدفوع. بظن في بدائل. في تطبيق مستخدمو من قبل عالويندوز مالي متذكروا (إذا بتعرف متل هيك تطبيق شاركني :) ) **شيء ندمت إنو جربته 😂** كان لازمني نظارات شمسية. استراتيجتي الحالية بالشراء: يا اشتري غالي كتير منيح يا أما لا تشتري أبداً. يعني لا تشتري رخيص وبعدين تكب. لا تشتري نص نص. اشتري شغلة منيحة وقيمها من بالك. مشان عدم تضويع وقتك أكثر، المهم، فتت ع محل Primark ولقيت نظارات بـ ٢ يورو. سعرون بجد كتير غريب. اشتريتون بشكل معاكس للاستراتيجية المتبعة 😒 ما بعرف ليش. المهم طلعوا بيخلوك تشوف الدنيا متل عصر التمانينات وكلو ع أصفر. ما بعرف إذا هي Feature لما Bug 😂. بصراحة إجتني أفكار شو إعمل فيون لح تشوفوها إذا طبقتا. **شو عم إسمع** عم عيد كتير [سورة الرعد للشيخ عبد الباسط عبد الصمد](https://www.youtube.com/watch?v=lY7AVtw1ViU) بالفترة الماضية. شخص بصوت غير عادي وبنفس غير عادي. النفس يلي بيقدر ياخدوا رائع. جرّب طريقة تجويده و"شوف|اتنفس" بنفسك. **أحسن شي شفتو**
- [الحلقة الثانية](https://www.youtube.com/watch?v=oN2LuIdQJi4&t) من قمرة 2. حلقة عن إيواء المشردين ونظرتنا للمشردين وشو ممكن تكون حياة المشردين قبل التشرد ونحنا بس عم ننظر للحالة الحالية إلون. قصة حقيقة بتمس المشاعر.
- تيد توك عن "ماذا تعلمنا من ٥ مليون كتاب؟" رائع بحق لغوغل بيحكي عن شو الاستنتاجات يلي فيك تعملها إذا كان في عندك ذكاء AI بيقدر يقرا ٥ مليون كتاب. رائع.
[youtube.com](https://www.youtube.com/watch?v=50MpgVJ8xAE)
**عشو عم حط وقتي** مشروع X يلي مافيني قولوا ليطلع. عم ياخد كل أسابيعي. **أكتر بوست محبوب [عالإنستغرام](https://www.instagram.com/zgtrshaker/)** بوست ساعتي! "[هي الساعة صرلها معي ٩ سنين.](https://www.instagram.com/p/BUQWvo7FLde/)" وبوست البوظة! "[لا أعرف من أين أبداً](https://www.instagram.com/p/BUcTO1KFnoj/)"
إن شاء الله تكون هي النسخة مفيدة وسريعة ولزيزة وعجبتك!
صياماً مقبولاً وفطوراً خفيفاً (مشان تلعب رياضة) وسحوراً شهياً وإيماناً عميقاً
أراكم الأسبوع القادم من شهر رمضان المبارك!
ـــــــــ
### الإسبوع 02: أكل بارد وكيف تنهي كورسيرا في ١٠ ساعات
- URL: https://mohammadshaker.com/en/blog/%d8%a7%d9%84%d8%a5%d8%b3%d8%a8%d9%88%d8%b9-%d9%a0%d9%a2-%d8%a3%d9%83%d9%84-%d8%b9%d8%a7%d9%84%d8%a8%d8%a7%d8%b1%d8%af
- Date: 2017-05-21T00:00:00.000Z
- Tags: weekly-review, personal, Human-Written
هاد هوي الإسبوع الثاني من شو بعمل بإسبوعي. الإسبوع الأول حطيت عليه رد بعد ما إجاني فيدباك وأسئلة.
#### Content
هاد هوي الإسبوع الثاني من شو بعمل بإسبوعي. [الإسبوع الأول](https://mohammadshaker.com/2017/05/13/%D8%A8%D8%B4%D9%88-%D8%AA%D8%BA%D9%91%D9%8A%D8%B1%D8%AA/) حطيت عليه [رد](https://mohammadshaker.com/2017/05/21/%D9%83%D9%8A%D9%81-%D8%B9%D9%86%D8%AF%D9%83-%D9%83%D9%84-%D9%87%D8%A7%D9%84%D9%88%D9%82%D8%AA/) بعد ما إجاني فيدباك وأسئلة.
لهاد الإسبوع ، لح إكتب فقط الشي الجديد. في شي عم يضل من الإسبوع الماضي ما لح ضوعلك وقتك فيه.
**شو عم إقرأ حالياً** Tools For Titans. خلصتوا للمرة الثانية. من أفضل الكتب يلي ممكن واحد يقراها بحياتو. إقراه بجد. **شي عم استمتع فيه** [سورة الكهف كل يوم جمعة بصوت القارئ الزهراني](https://www.youtube.com/watch?v=CC6xexlUQ20) رائعة. **قيد عم إتحدى فيه نفسي**
إخلص كورس كامل ع كورسيرا بـ 10 ساعة. (كانت متعبة قليلاً بس هاد جزء من التحدي) أحد الشغلات يلي كنت إعملها من الجامعة لهلق هيي إني شوف الفيديو، الكورس عسرعة x2, x3, etc.
1. إذا متصفح في extension لمتصفح الكروم بستخدما هيي [Video Speed Controller](https://chrome.google.com/webstore/detail/video-speed-controller/nffaoalbilbmmfgbnbgppjihopabppdk?hl=en) بتشتغل عـ ٩٥٪ من المواقع يلي فيها فيديو. إستخداما كتير بسيط عن طريق فقط Shortcuts.
2. إذا عم تشوف فيديو عجهازك ففي VLC فيه خاصية التسريع.
3. إذا Phone/Tablet بستخدم أكتر شي كمان VLC.
السؤال يلي دائماً بسألو لحالي هيي كيف أدخل معلومات لعقلي بطريقة أسرع. وقت تدرك إنو عقلك بيستوعب وبيعمل بسرعة أكتر بكتير من عيونك أو أذنيك أو فمك، فهاد بيستدعي تستكشف أكتر عن طريق تدخل فيها المعلومات لعقلك بطريقة أسرع. بالنهاية عيونك وأذنيك وفمك هنن وسط Medium فقط. هاد الوسط هوي يلي بيحدنا إنو نفهم ع بعض أسرع أو ننقل المعلومات بين بعض بطريقة أسرع. لو كان فينا نتفاهم فقط مخ لـ مخ (Mind to mind) فلا حدود لأديش ممكن نعرف عن بعض أكتر، نتعلم أكتر، نتحرك أسرع. بالعودة للبوست، فأسهل شي عواض ما تسمع أو تشوف بالسرعة العادية (خصوصاً محاضرات أو كورسات) هيي إنو تقرّب أكتر من الحد يلي بيقدر فيه عقلك يستوعب الكلام أو الصور (أي تخليه الكلام أو الفيديو يسرع للحد يلي بيقدر عقلك يستوعبو.) هي وحدة من الطرق إني مثلاً بتعلم شغلات x2-x3 من غيري باليوم الواحد.هاد هوي كمان جزء من بوست قادم عن كيف تقرأ أكثر من ١٠٠ كتاب في السنة. لح إشرح فيه أكتر. **شيء عم جربه**
أكل بارد (لا يحتاج إلى تسخين) لجمعة كاملة. **شي شفتو مفيد** [Creating Extraordinary Life](http://www.oprah.com/own-supersoulsessions/Tony-Robbins-Creating-an-Extraordinary-Quality-of-Life) إذا كنت عم اتدور ع حدا يحكيلك عالـ Why مو عالـ What فـ Tony Robbins من أكتر الأشخاص يلي لح يفيدوك. هاد التوك لزيز بس مو أحسن شي إلو. بس بيضل مفيد. في إلو كتب إذا حابب تعرف أكتر. في إلو سلسلة صوتية إسمها [Get The Edge](https://www.youtube.com/results?search_query=Anthony+Robbins.+Get+the+Edge) سمعتا من عدة سنوات كبرنامج تعلموا لجمعة كاملة وتغير فيه كيفية رؤيتك للأمور. رائعة بحق ومفيدة بحق. **أكتر بوست محبوب عالإنستغرام** [٣٢٨٤ كيلو متر بعيدة عني.](https://www.instagram.com/p/BUP7yQbltLs/?taken-by=zgtrshaker) صورة أخدتها وقت نزلت عالشام قريب لباب توما. هلق صار دورك تقلي شو عملت شي جديد بهالجمعة أو شي تحديت حالك فيه!
### "كيف عندك كل هالوقت"
- URL: https://mohammadshaker.com/en/blog/%d9%83%d9%8a%d9%81-%d8%b9%d9%86%d8%af%d9%83-%d9%83%d9%84-%d9%87%d8%a7%d9%84%d9%88%d9%82%d8%aa
- Date: 2017-05-21T00:00:00.000Z
- Tags: weekly-review, personal, Human-Written
السلام عليكم مرة أخرى،تكملة ورداً عبوستي قبل يومين عن شو بعمل بإسبوعي. لح جاوب عالأسئلة يلي إجتني مشان إفتح نقاش مرة أخرى لكل يلي بيحب.
#### Content
السلام عليكم مرة أخرى،تكملة ورداً عبوستي قبل يومين عن شو بعمل بإسبوعي. لح جاوب عالأسئلة يلي إجتني مشان إفتح نقاش مرة أخرى لكل يلي بيحب. مايلي أبداً ما بيخولني كون Expert بالموضوع. عم إحكي عن حالي فقط. مو شرط شو هوي الصح. الموضوع مفتوح للنقاش أو أي حدا عندو Tricks تانيين كتير لح حب اسمعون! ـــــــــــــــ "مشان الوقت.. كيف عندك كل هالوقت" أولاً ما بعتقد حالي لسا مظبط وقتي تمام. في أشياء بعرفا إنو في وقت بضوع عليها وإني بقدر كون أحسن. عم حاول. بالنهاية البوست غايتو أعرف كمان كيف العالم غيري بتظبط وقتا واقدر شوف شو فيني إعمل أنا كمان. في كتير أشياء ذكرتها بـ "لا تفكر في حليب أحمر" بتساعد كتير عقصص الـ Management تبع الوقت. يمكن في أقسام انحذفت وما طلعت بالنسخة الأخيرة من الكتاب لح اذكر بعضا هون. إذا عم تحكوا ع وقت كنت بالجامعة فكنت مخلوق فضائي بجد بهداك الوقت لأنو ما كان عندي شي غير دراسة وعلم وعلم وعلم وإني كون أحسن شي بكل شي. هاد ما بعتبروا شي غلط أو ندمان عليه. بالعكس بفتخر فيه لأبعد الحدود ولح داوم عليه. بس حالياً عم اخد شوي غير Approach ليومي لأنو في غير هدف وعم خلي حياتي من غير نمط. ـــــــــــــــ "كيف بتنظم الوقت وكيف بتقدر تلتزم اساسا بشي جديد ولفترات طويلة" | "كيف عم تلحق" المشكلة أبداً مو "بتنظيم الوقت." يلي اكتشفتوا إني بكون Productive أكتر بكتير بدون ما إعمل أي شي زيادة وقت يكون في هدف واضح عم اسعى إلو. (متل لنفرض كتابة كتاب بشهرين، إصدار تطبيق بتاريخ محدد، مادة بدي خلّصها، مهارة جديدة لازم إعملها بتاريخ محدد.) لاحظ إنو كلون في إلون ديدلاين دائماً يلي هوي قد أهمية التاسك بنفسها. لاحظت إنو فقط يكون في هدف واضح مع تاريخ واضح عم اشتغل عليه حالياً بيساعد كتير إنو أي شي تاني مالو ضروري ومالو مرتبط بالهدف يصفى ع جنب. إذا بدنا نحكي ع تنظيم الوقت بالزات فهوي أهم شي يمكن هوي الشي يلي قلتو بالبوست "كونك مشغول دائماً يدل على كسلك." كسلنا من إنو ما منحط أولويات وإذا حطينا ما منشتغل عليها. عم إحكي ع حالي كمان. الشي يلي جربتوا إنو اكتب كتابة ع ورق (لاحظت إنو كتير مهم يكون مكتوب مو بس بمخي) وحط Priority index إدام كل بند من يلي لازم إعملوا. بكل بساطة بعدين ببدا من الأهم بغض النظر عن أي شي تاني. فقط إني بعمل هيك بيخفف الضغط ع باقي نهاري جداً. جربتها وظابطة. (بستعمل تطبيق مجاني اسمو Wunderlist ومن قبل كنت تطبيق Reminders ومن قبل كنت عالـ Notes حتى إن كان أندرويد لما أيفون.) شي تاني جربتوا هوي إني حط الشي يلي بدي أعملوا عالـ Calendar. فقط بإني إكتب مثلاً إنو هاد اليوم في عندي ركض وهاد اليوم سباحة وهاد اليوم هوي الديدلاين للكتاب وهاد اليوم ل.. إلخ بيخلي عقلي دائم صاحي وعرفان شو الحد أو شو يلي لازم يعملوا. هي كمان شغلة مثبتة مو بس من عندي. الفكرة أبداً مو إنو عبي الـ Calendar بشي مالو داع. بس أي عادة جديدة بدي اكتسبها ببلشها عالـ Calendar. شغلة أخيرة مشان إني إلزم حالي بشغلات جديدة. هيي ببساطة حب الـ Experiments. يعني مثلاً ما بحب إمشي بالطريق نفسو مرتين. ما بحب جيب نفس اللون من الأواعي. ما بحب جرب نفس النوع من القهوة مرتين. ما بحب جرب نفس المحل مرتين (إلا أفوكادو ونستلة من عند أبو عبدو بالصالحية هي لا تقاوم). ما بحب جرب نفس السندويشة من نفس المحل مرتين (مع إنو أحياناً كتير باكلها بس بضل مبسوط صراحة إني عم جرب - متل قبل يومين وقت طلبت سندويشة جديدة وطلعت X. إحم إحم @Hisham. هل أنا زعلان لأنو طلبتها وطلعت شي مو منيح؟ أبداً وأبداً وأبداً.) ـــــــــــــــ "لاحظت إنو ما كتير بتفوت عالسوشل ميديا. ليش؟" | "هل بتعتبرها منيحة وإلها فائدة؟" | "كم دقيقة باليوم بتقضيها عالفيسبوك, وانستغرام؟" قبل ما جاوب ع هاد السؤال. في فكرة الـ Fear of Missiong Out وهيي إنو منخاف إنو يروح علينا شي شغلة وهيي مصيبة صراحة. مشان هيك وقت نشوف إنو في Notification بيطق مخنا وخلص لازم نكبس مثلاً ع أيقونة الفيسبوك عالموبايل مشان أعرف شو هيي الـ Notification. مع أنو أغلب الأحيان بتطلع لشي لعبة ما عنا أي اهتمام فيها. مشان هيك مثلاً بالنسبة إلي ع تلفوني ع فكرة لاغي كلشي Notifications وخصوصي من كلشي Social Media. يعني ما بيطلعلي شي عالـ Lock screen أبداً. لحتى أنا فوت عالتطبيق بحد ذاتو لحتى يطلعلي شوفي Notifications. حتى عالأيقونة البرانية مافي أي شي. هاد خصوصي مشان إخلص من هالقصة. ع فكرة مفيدة جداً وسيظل تلفوني صامداً وبسيطاً ورايقاً هكذا. منرجع للسؤال. السؤال بيعني ضمنياً إنو أنا بفوت عالفيسبوك والإنستغرام يومياً. وهاد misconception :D. بفترة الـ ٣ سنين الماضية كانو شي سنتين ونص منون الفيسبوك تبعي مسكر إلا في أوقات محددة. المشكلة مو بالفيسبوك أو بأي شي تاني المشكلة بشو الفائدة منو. حالياً الفائدة منو إنو كل كم يوم بفوت وبرد ع رسائل وعأسئلة وهاد بعتبروا شي منيح ومو ضياعة وقت دائماً لأنو "حوائج الناس إليكم من نعم الله عليكم" وخصوصاً للعالم الفهمانة. جانب إنو الفيسبوك هوي لخلق ترابط مع العالم ورفقاتك وهي الـ story يلي منحطا بعقلنا من إنو الفيسبوك شي منيح في جزء منها صحيح بس يلي عم شوفوا مؤخراً إنو الفيسبوك كتير صاير "غليظ." هي وجهة نظري بالنهاية. من ناحيتي بحاول إذا بدي إنشر يكون شي مفيد إقدر أعطي العالم فائدة من وراه أو افتح نقاش من وراه متل هاد البوست. حالياً بفوت عالإنستغرام بعمل Publish للصور يلي بلقطها وبطلع. الفيسبوك في فترة كانو البوست عصفحتي Strong Emotions كلون Automated. يعني كاتبون مسبقاً وبينشتروا بوقت محدد لحالون. هاد مو شي عيب أو جديد أو أو أو. بالعكس كل الشركات بتعملوا وأنا عملتو مشان كون Efficient أكتر لأنو ساعتها بظبطون بوقت واحد وما بضيع وقت كل مرة بالأضافة إنو ما بصير بفكر دائماً إيمت لازم إفتح الفيسبوك مشان إنشر شي عالصفحة (خلقة في فرق توقيت بين هون وأغلب جمهور الصفحة بالعالم العربي حوالي الساعة أو التنتين.) أخيراً نقطة إضافية هيي إنو عم جرب حالياً إنو ما يكون معي لابتوب بس إرجع عالبيت (اللابتوب فقط بالشغل مثلاً.) ممكن تقلي كيف إذا أنا طالب معلوماتية جامعة ويي ويي واللابتوب أهم شي. لح قلك. أغلبكون بظن عنكون تابليت. وقت أنا كنت جامعة كنت قصداً أحياناً إقرا كتب أو محاضرات من التابلت لأنو استخدام النت عليه وتضويع الوقت كان كتير بطيء وبيطلّع الروح وبالتالي (متل إلغاء الخيار الافتراضي) صار ما عندي مجال غير إنو ركّز فقط بالمحاضرة أو بالكتاب وبدون أي ملهيات تانية. هاد يمكن كان يضفلي ساعتين ع يومي (No multi-tasking or task switching يلي بيضوعوا وقت كبير جداً جداً.) ـــــــــــــ "شو طبيعة عملك/دوامك" Software Engineer / Full stack. 8 ساعات كالمعتاد هنا + 0.5 استراحة + مواصلات. يمكن أكتر شوي من الشام بظن. ـــــــــــــ "ئديش الأعمال المنزلية ممكن تاخد من وقتك؟" بما إني عايش لحالي هنا فبحاول هي الأعمال تاخد أقل وقت ممكن. مثلاً طبيعة أكلي وطعامي بعملها خصوصي إنو ما تحتاج وقت أبداً متل إنو يكون اعتمادي أغلب شي عالخضرة يلي ما بدها طبخ حتى أو فواكه أو حليب + شي تاني. هي شغلة. (بس أحياناً بنسى إني آكل أو ما بيزيد وقت - بس مو دائماً إحم إحم. مو إنو شي بيدايق أو يا حرام أو يا لطيف شو فظيع ما بيقدر ياكل من كتر ما عندو شغلات يعملها. نوب. بالعكس بنظرلها نظرة إنو عم يقوا جسمي وخلي حالي Antifragile أكتر. الفكرة رهيبة من نسيم طالب (بروفسور لبناني بجامعة نيويورك) بكتابو بنفس العنوان Antifragile.) وبغض النظر عن هاد فأحياناً بعملها قصداً كجزء من حمية Ketone Diet or slow-carb. (الحمية مو لأنو الواحد صحتو منيحة. الحمية مشان يكون الجسم أقوى.) ـــــــــــــ "عم تقول: 1- "قيد عم إتحدى نفسي فيه": آخد صور بالأسود والأبيض فقط. شو الهدف من هالقيد؟ فقط تحدي لا أكثر؟!!" نوب مو فقط الواحد ينظرلوا إنو تحدي مع إنو هاد أكيد جانب. بس أولاً بحب التصوير وبدي نمي مهارتي بهالقصة. وتانياً إذا بدي خلي حالي آخد أي صورة حلوة فمالح كون مصور منيح لا هلق ولا مستقبلاً لأنو ساعتها فيني كل دقيقتين وقف بالطريق وأخد صورة من زاوية وتطلع الحلوة. الفكرة إنو صعّبها ع حالي مشان يصير مخي يقدر يلقط فقط الشغلات يلي بتطلع حلوة ضمن الخيارات المتاحة إلو. ما بدي ياه يلقط أي صورة حلوة. بدي فقط الصورة يلي بتطلع حلوة بالأبيض والأسود. وبعدها برجع بكبر الخيار عليه. هي فكرة القيود هيي من أهم الأفكار ع فكرة بـ "[لا تفكر في حليب أحمر](http://mohammadshaker.com/2017/04/22/%d9%83%d8%aa%d8%a7%d8%a8-%d9%84%d8%a7-%d8%aa%d9%81%d9%83%d9%91%d8%b1-%d9%81%d9%8a-%d8%ad%d9%84%d9%8a%d8%a8-%d8%a3%d8%ad%d9%85%d8%b1/)." ما بعرف إذا كل العالم قدّرت قيمتها. ـــــــــــــ إذا في شي أسئلة تانية أنا هنا. لح حب أعرف أيضاً إنتو شو في عنكون Tricks بتعملوها مشان إن كان مشان الـ Management عالوقت أو ع حياتكون بشكل عام.
### الإسبوع 01: بشو تغيّرت هالإسبوع؟
- URL: https://mohammadshaker.com/en/blog/%d8%a8%d8%b4%d9%88-%d8%aa%d8%ba%d9%91%d9%8a%d8%b1%d8%aa
- Date: 2017-05-13T00:00:00.000Z
- Tags: weekly-review, personal, Human-Written
بالعادة أنا شخص كتير Private. ولح ضل هيك. بس. بس بما إني لازم جرب شي جديد هي الجمعة فلح يكون هاد النمط من البوستات تجربة لعدة جمع قادمة إذا طلع ظابط.
#### Content
بالعادة أنا شخص كتير Private. ولح ضل هيك. بس. بس بما إني لازم جرب شي جديد هي الجمعة فلح يكون هاد النمط من البوستات تجربة لعدة جمع قادمة إذا طلع ظابط. ببساطة هوي بوست إجابة عن أسئلة "بعضا غريب أو بسيط جداً" عن أشياء عملتها ضمن الـ ٧ أيام السابقة. أي الجمعة السابقة. إذا بتعرف Tim Ferris (مؤلف كتب Tools for Titans, The 4-hour Workweek) وبغض النظر إذا بتتفق معو أو لاء، فهوي بيعمل كل جمعة إجابة ع نمط الأسئلة يلي لح جاوبو حالياً أنا هون. غايتي فقط عواض ما تكون إجاباتي فقط إلي وبحتفظ فيها إلي ع دفتر، هدفي تكون إجابتي مشاركة مع كل حدا لح يقرا البوست بأخر كل جمعة. مشان إقدر إتعلم من غيري أو فيدوا يمكن بطريقة تفكيري وشو بعمل بإسبوعي. مشان هيك بعد ما تقرا البوست هاد لح حب كتير إنو تشاركني برأيك إنت أيضاً وتجاوب عالأسئلة إنت كمان إذا حبيت. \_\_\_\_\_\_\_ شو عم إقرأ حالياً
Tools For Titans. عم إرجع إقرا للمرة الثانية. يمكن أكتر كتاب فادني بحياتي لهلق. رائع. لح إحكي عنو بشي مفصل لاحقاً.
New Retro. كتاب عن نوع الديزاين الرجعي الريترو. جيد.
At the Edge of Art. كتاب مفترض يكون عن الفن. بس ما حبيتو.
اقتباس عم طبقوا | فكر فيه
“Being perpetually busy is a kind of laziness” - Tim Ferris
"كونك مشغول دائماً دلالة على كسلك"
قيد عم إتحدى نفسي فيه
آخد صور بالأسود والأبيض فقط.
فيك تشوفون عـ instagram هنا:
[instagram.com](https://www.instagram.com/zgtrshaker/)
شيء عم داوم عليه
\- صيام ٣ أيام بالإسبوع (غير إنو شهر شعبان - فيك تشوف فوائد ما يطلق عليه حمية الكيتون يلي بتعتمد على الصيام Ketone Diet)
\- ركض كل يوم فردي بالإسبوع مشان ترجع ركبتي ظابطة (قصة طويلة)
\- أكملت الإسبوع الثاني من لعب Yoga كل يوم
شيء جديد عم جربه
بروكلي مع جبنة Brie.
فيك تشوف الفوائد الكتيرة للبروكلي عالإنترنت ومشان هيك بلشت فيه.
الجبنة فقط مشان تظبط الطعمة.
شو عم إسمع
Tim Ferris Show البودكاست الرائعة
(فيك تسمع كلشي حلقات مجاناً)
شو عم شوف
Gifted. فلم عن فتاة موهوبة والصراع بين إنو تكون فتاة طبيعية أم تتبع موهبتا أم التنين
Guardian of the Galaxy 2. فقط للضحك والغلاظة. لزيز.
أحسن شي اشتريتو مؤخراً
تطبيق Yoga Studio عالأيفون. من أفضل التطبيقات يلي مستخدما لحد الآن.
شو عم إلعب رياضة
ركض. ٤ مرات
Yoga Studio مرة أخرى، يوغا، كل يوم، ٣٠د باليوم.
شو عم أكل
بيض + سبانخ
بروكلي، برتقال، تمر
شو عم إشرب
زهورات
حليب + كاكاو
شو عم إلبس
رجعت لنمط الـ minimalistic.
وإنو يكون فقط عندي الشغلات المحتاجها فقط.
حالياً فقط عندي
٣ بنطلونات + ٢ كنزة شتوي + ٢ كنزة صيفي + ٢ قميص بقبة
شو عم إلعب
duolingo! دولينغو مع إنو مو لعبة
بشو عم فكر كتير مؤخراً
ع شو ركّز بالمستقبل القريب.
أكتر بوست محبوب عالفيس
كتابي يلي طلع من إسبوعين:
[لا تفكر في حليب أحمر](http://mohammadshaker.com/2017/04/22/%d9%83%d8%aa%d8%a7%d8%a8-%d9%84%d8%a7-%d8%aa%d9%81%d9%83%d9%91%d8%b1-%d9%81%d9%8a-%d8%ad%d9%84%d9%8a%d8%a8-%d8%a3%d8%ad%d9%85%d8%b1/)
[facebook.com](https://www.facebook.com/StrongEmotionsNow/posts/1844616479194541)
لغة عم إتعلمها
الألمانية ع Duolingo
إذا كنت محب لتعلم اللغات ضيفني ع دولينغو هنا:
[duolingo.com](https://www.duolingo.com/Mohammad.Shaker)
### اليوم الذي أصبحت فيه كاتباً أفضل
- URL: https://mohammadshaker.com/en/blog/%d8%a7%d9%84%d9%8a%d9%88%d9%85-%d8%a7%d9%84%d8%b0%d9%8a-%d8%a3%d8%b5%d8%a8%d8%ad%d8%aa-%d9%81%d9%8a%d9%87-%d9%83%d8%a7%d8%aa%d8%a8%d8%a7%d9%8b-%d8%a3%d9%81%d8%b6%d9%84
- Date: 2017-05-09T00:00:00.000Z
- Tags: thoughts, personal, Human-Written
هل يمكنك أن تصبح كاتباً أفضل في دقيقتين؟إقرأ أكتر في دقيقتين هنا: http://dilbertblog.typepad.com/the\dilbert\blog/2007/06/the\day\you\bec.html
#### Content
هل يمكنك أن تصبح كاتباً أفضل في دقيقتين؟إقرأ أكتر في دقيقتين هنا: [dilbertblog.typepad.com](http://dilbertblog.typepad.com/the\_dilbert\_blog/2007/06/the\_day\_you\_bec.html)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### سكادوووش! من كونغ فو باندا!
- URL: https://mohammadshaker.com/en/blog/%d8%b3%d9%83%d8%a7%d8%af%d9%88%d9%88%d9%88%d8%b4-%d9%85%d9%86-%d9%83%d9%88%d9%86%d8%ba-%d9%81%d9%88-%d8%a8%d8%a7%d9%86%d8%af%d8%a7
- Date: 2017-05-08T00:00:00.000Z
- Tags: essay, Human-Written
يوجد السؤال التالي عـ Quora: في فيلم كونغ فو باندا، كيف استطاع بو في النهاية تطوير قدرته ليكون مقاتل كونغ فو عظيم؟ الجميل هيي الإجابة يلي عليها أكتر من
#### Content
يوجد السؤال التالي عـ Quora: في فيلم كونغ فو باندا، كيف استطاع بو في النهاية تطوير قدرته ليكون مقاتل كونغ فو عظيم؟ الجميل هيي الإجابة يلي عليها أكتر من 100k مشاهدة من Eric Weinstein يلي عندو أراء كتير عميقة بكل شي غالباً (بالإضافة إنو هوي مدير Thiel Capital التابعة لـ Peter Thiel من PayPal Mafia) قراءة قصيرة ممتعة! [quora.com](https://www.quora.com/In-Kung-Fu-Panda-how-does-Po-end-up-developing-the-capability-to-be-an-awesome-Kung-Fu-fighter)
### أظن أنك سمين! I Think You Are Fat!
- URL: https://mohammadshaker.com/en/blog/%d8%a3%d8%b8%d9%86-%d8%a3%d9%86%d9%83-%d8%b3%d9%85%d9%8a%d9%86-i-think-you-are-fat
- Date: 2017-05-07T00:00:00.000Z
- Tags: personal, Human-Written
هل يجب علينا دوماً عدم الكذب؟ Read on: http://www.esquire.com/news-politics/a26792/honesty0707/
#### Content
هل يجب علينا دوماً عدم الكذب؟ Read on: [esquire.com](http://www.esquire.com/news-politics/a26792/honesty0707/)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### تحتاج فقط إلى ١٠٠٠ مشجع!
- URL: https://mohammadshaker.com/en/blog/%d9%81%d9%82%d8%b7-%d9%a1%d9%a0%d9%a0%d9%a0-%d9%85%d8%b4%d8%ac%d8%b9
- Date: 2017-05-07T00:00:00.000Z
- Tags: personal, Human-Written
Read more in this piece: http://kk.org/thetechnium/1000-true-fans/
#### Content
Read more in this piece: [kk.org](http://kk.org/thetechnium/1000-true-fans/)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### The World's Healthiest 75-Year-Old Man
- URL: https://mohammadshaker.com/en/blog/the-worlds-healthiest-75-year-old-man
- Date: 2017-05-07T00:00:00.000Z
- Tags: personal, Human-Written, Life
Read a post from 2008: http://www.esquire.com/news-politics/a4454/don-wildman-0508/
#### Content
Read a post from 2008: [esquire.com](http://www.esquire.com/news-politics/a4454/don-wildman-0508/)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### كتاب لا تفكّر في حليب أحمر!
- URL: https://mohammadshaker.com/en/blog/%d9%83%d8%aa%d8%a7%d8%a8-%d9%84%d8%a7-%d8%aa%d9%81%d9%83%d9%91%d8%b1-%d9%81%d9%8a-%d8%ad%d9%84%d9%8a%d8%a8-%d8%a3%d8%ad%d9%85%d8%b1
- Date: 2017-04-22T00:00:00.000Z
- Tags: Art, Behavioral Economics, Behaviour, Brain, Design, drawing, entrepreneurship, Life, Mind, Read, Sketching, فن, قراءة, كتاب, تصميم, تصرف, دماغ, رسم, عقل, Human-Written
هل تبادرت لك صورة الحليب الأحمر عندما قرأت عنوان هذا الكتاب؟ قلت لك ألّا تفكّر بحليب أحمر!
#### Content
**هل تبادرت لك صورة الحليب الأحمر عندما قرأت عنوان هذا الكتاب؟ قلت لك ألّا تفكّر بحليب أحمر!**
هذا كتاب يثير فضولك، يغيّر من نمط تفكيرك، يخيّب أملك بنفسك، ومن ثمّ يعيد ثقتك بها. يحبطك ويضحك عليك لتضحك أنت عليه بعدها. هو كتاب لا أرقام للصفحات فيه. فهرسه غريب ورسومه غريبة وألوانه أغرب.

كتاب يجيب عن أسئلة جديدة، عميقة، غريبة وتستحق منّا التّفكّر. لماذا نتصرّف كما نتصرّف. كيف يكون تفكيرنا غير منطقيّ دون أن نعلم. فما علاقة متجر إيكيا IKEA بالدِّين مثلاً؟ لماذا نستخدم نغمة الرّنين الافتراضيّة لهواتفنا؟ هل يمكنك كتابة كتاب بـ 50 كلمة مختلفة فقط؟ هل 90٪ بدون دسم أفضل من ١٠٪ دسم؟ وهل يمكن لمارشميلو أن يتنبّأ بمستقبلك؟

لن أكلّمك عن نفسي أو من أكون أو ماذا أفعل لأنه كتاب لك أولاً قبل كلّ شيء. كتابٌ منّي لأمّتي. كتاب بنمط جديد، مفيد، علميّ، عربيّ، بخطوط جديدة، برسوم غريبة، وألوان أغرب. أردته عربيّاً بفن جديد. أردته عربيّاً لكل عربيّ.

## **إقرأ الكتاب بشكل مجاني لفترة محدودة بالنقر [هنا](https://goo.gl/CQjeQN "MShaker_DontThinkInRedMilk")!** **إذا أحببت الكتاب وأنهيته مباشرة، يمكنك تقييم الكتاب ونشره على الـ goodreads بالنقر [هنا](https://www.goodreads.com/book/show/34936066?ac=1&from_search=true).**
### من مؤسس Duolingo
- URL: https://mohammadshaker.com/en/blog/%d9%85%d9%86-%d9%85%d8%a4%d8%b3%d8%b3-duolingo
- Date: 2017-02-05T00:00:00.000Z
- Tags: thoughts, personal, Human-Written
محاضرة ملهمة من مؤسس تطبيق Duolingo للتعلم اللغوي. This concise summary states the core idea clearly.
#### Content
[youtube.com](https://www.youtube.com/watch?v=-Ht4qiDRZE8)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### نهاية عام مع القراءة - 2016 / A Year of Reading
- URL: https://mohammadshaker.com/en/blog/%d9%86%d9%87%d8%a7%d9%8a%d8%a9-%d8%b9%d8%a7%d9%85-%d9%85%d8%b9-%d8%a7%d9%84%d9%82%d8%b1%d8%a7%d8%a1%d8%a9-2016-a-year-of-reading
- Date: 2017-01-07T00:00:00.000Z
- Tags: Reading, Human-Written
مع أسوء وأفضل كتب قرأتها في 2016، أكثر الكتب تخييباً للآمال (لا تقرأها): Don't Read - Where Good Ideas Come From - The Personal MBA - A Whole New Mind
#### Content
مع أسوء وأفضل كتب قرأتها في 2016، أكثر الكتب تخييباً للآمال (لا تقرأها): Don't Read - Where Good Ideas Come From - The Personal MBA - A Whole New Mind ــــــــــ أكثر الكتب إفادة: Good Reads 1 - The Everything Store: Jeff Bezos and the Age of Amazon 2- الضوء الأزرق، حسين البرغوثي 3 - Thinking Fast and Slow 4 - Influence, Robert Cialdini 5 - Meggs' History of Graphic Design 6 - Pixel Perfect Precision, Ustwo 7 - Jim Henson: The Biography ــــــــــ أكثر كتاب حابو حالياً: Tools for Titans, Tim Ferris, 2017 إذا بتحب القراءة عطيني رأيك، شوف القائمة كاملة و ضيفني هنا: [https://www.goodreads.com/user/year\_in\_books/2016/11193121](https://www.goodreads.com/user/year_in_books/2016/11193121)
كل عام وأنتم بألف ألف خير
### المعاني الرائعة لموسيقا الفصول الأربعة لفيفالدي
- URL: https://mohammadshaker.com/en/blog/%d9%85%d8%b9%d8%a7%d9%86%d9%8a-%d9%85%d9%88%d8%b3%d9%8a%d9%82%d8%a7-%d9%81%d9%8a%d9%81%d8%a7%d9%84%d8%af%d9%8a
- Date: 2016-12-21T00:00:00.000Z
- Tags: Music, Human-Written
استكشاف المعاني العميقة والجمالية في موسيقى فيفالدي الخالدة
#### Content
[youtube.com](https://www.youtube.com/watch?v=Xcpc8VDsv3c)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### بيتر تيل، الاحتكار و صفر-واحد
- URL: https://mohammadshaker.com/en/blog/%d8%a8%d9%8a%d8%aa%d8%b1-%d8%aa%d9%8a%d9%84%d8%8c-%d8%a7%d9%84%d8%a7%d8%ad%d8%aa%d9%83%d8%a7%d8%b1-%d9%88-%d8%b5%d9%81%d8%b1-%d9%88%d8%a7%d8%ad%d8%af
- Date: 2016-10-16T00:00:00.000Z
- Tags: startups, entrepreneurship, Human-Written
محاضرة بيتر تيل حول الاحتكار والابتكار في عالم الشركات الناشئة
#### Content
[youtube.com](https://www.youtube.com/watch?v=6kGND-uZolY)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### عن الإصرار.. الإصرار.. الإصرار..
- URL: https://mohammadshaker.com/en/blog/%d8%b9%d9%86-%d8%a7%d9%84%d8%a5%d8%b5%d8%b1%d8%a7%d8%b1-%d8%a7%d9%84%d8%a5%d8%b5%d8%b1%d8%a7%d8%b1-%d8%a7%d9%84%d8%a5%d8%b5%d8%b1%d8%a7%d8%b1
- Date: 2016-10-15T00:00:00.000Z
- Tags: startups, entrepreneurship, Human-Written
محاضرة ملهمة عن أهمية الإصرار والمثابرة في تحقيق النجاح
#### Content
[youtube.com](https://www.youtube.com/watch?v=KmfuIt96vU4)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### قمة الغلاظة في حدث غوغل البارحة!
- URL: https://mohammadshaker.com/en/blog/galaza-google-hardware-event-2016
- Date: 2016-10-05T00:00:00.000Z
- Tags: thoughts, personal, Human-Written
البارحة قامت غوغل بإقامة حدث غليظ (ومتشدق) للغاية، والذي كان عن العتاد Hardware بشكل رئيسي.
#### Content
البارحة قامت غوغل بإقامة حدث غليظ (ومتشدق) للغاية، والذي كان عن العتاد Hardware بشكل رئيسي. بالمقارنة مع حدث آبل الشهر الماضي فحدث غوغل من حيث العرض يحظى بدرجة تحت الصفر حتماً. من حيث المضمون لا شيء جديد ضمن ساعة ونصف غير تلك التي تحدثت فيه غوغل عن جهازها الجديد لمدة ٢٥ دقيقة. ما تبقى "صف كلام" مشروح سابقاً في مؤتمر Google IO قبل عدة أشهر. المقطع الأغلظ برأيي هو في وسط الحدث عندما تحدثت غوغل عن جهازها الجديد Google Pixel.
> ## الغلاظة الشديدة تكمن في كيفية الضحك على المستخدمين غير المختصين
أمور سأشرحها بالتتالي وأقارن بعضها مع العمة آبل.
- **طريقة العرض منذ البداية تشبه الأكاديميا**
طريقة العرض كالتالي: عرض أفكار في البداية ومن ثم شرحها لاحقاً.لا تشويق ولا هم يحزنون. عرض أكاديمي بامتياز. "إذا غوغل عم تعمل الحدث فهاد ما بيعني بالضرورة إنو الحدث لازم يكون فظيع ولازم نحبو." الذي صمم العرض التقديمي من أسوأ ما يكون. عرض للمواصفات والنقاط مسبقاً ومن ثم شرحها. آبل لا تفعل هذا أبداً في عروضها التقديمية. آبل لا تعرض النقاط مسبقاً أبداً لأنها تريد الاستحواذ على انتباه الجمهور لكامل العرض. تعرض كل نقطة وراء الأخرى لتبقي الجمهور متحمساً ومتشوقاً لما ستأتي به خلال المؤتمر تباعاً. (غوغل: إذا ما بتعرفي تعملي عروض اقري هاد الكتاب [هنا](https://www.google.com/webhp?sourceid=chrome-instant&ion=1&espv=2&ie=UTF-8#q=Resonate+book)) في هذه النقطة بالذات عند هذه الصورة في مؤتمر غوغل فقدت أي اهتمام لما سيأتي (هذا إلى جانب أن النقاط المعروضة نقاط ولاد صغار وكل الجمهور يعرفها منذ مؤتمر Google IO.) الأكثر غلاظة برأيي هي طريقة غوغل في ضحكها على جمهور غير خبير بالتكنولوجيا عندما عرضت نقاط ستشترك فيها أجهزة الأندرويد جميعاً وليس جهاز Pixel بالذات. (سأشرح أكثر مع النقاط القادمة.)
- 
- **من الخلف، فالجهاز الجديد نسخة من جهاز iPhone 6 من سنتين**
لا إبداع أبداً في مجال التصميم. هذا ما اعتدنا عليه من أجهزة غوغل ومن أجهزة الـ Android عموماً. لا إبداع ولا تصميم. فقط عرض عضلات فيما يتعلق بمحتوى العتاد وهذا ينقلنا إلى النقطة التالية. 
- **عضلات بدون مخ**
"عضلات بدون مخ".. هذه المقولة الشائعة عن أي شخص "ذو عضل" والتي تنطبق تماماً على جهاز غوغل الجديد. "تشاطرت" غوغل وعرضت فقط أنها الأفضل على مستوى الكميرا (أكثر ب ٣ نقاط من آبل ضمن اختبار اختارته غوغل بنفسها.) الفكرة أن غوغل لم تعرض أبداً مقارنات أخرى تخسر فيها دائماً أمام أجهزة آبل (دائماً تخسر في مجال أداء المعالج وأداء المعالج الرسومي.) بقا.. حاج غلاظة وضحك عالبشرية.. 
- **مساعد غوغل!**
على الرغم من أن الذكاء الصناعي الذي أطلقت عليه غوغل اسم "مساعد غوغل" Google Assistant هو عمل مذهل، فإن القول بأن الجهاز الجديد فيه مواصفة رهيبة بأنه يحتوي هذا المساعد هو ضرب من الغباء ومرة أخرى ضحك على الجمهور. آجهزة الأندرويد القادمة ستكون مثل جهاز غوغل تماماً فيها مساعد غوغل وبالتالي ليس هناك أي ميزة إضافية هنا!
- **غلاظة تقديم**
انظر إلى الصورتين التاليتين ولاحظ الغلاظة..  ومرة آخرى هنا..  من المفترض أن هذا الشخص يقدم عرض تقديمي ومن المفترض أنه الآن يرسل رسالة لزوجته عن طريق مساعد غوغل. الغلاظة تكمن في أنه مفتعل جداً جداً. إنه يقرأ من نص مكتوب كل نقطة من النقاط التي يجب عليه أن يقوم بها. هذه العملية موجودة دائماً في العروض التقديمية ولكن ليس بهذا الشكل الغليظ جداً.
> ## أضف إلى ذلك أنه يريد أن يرسل رسالة لزوجته، هذا يدل على أنه يوجد ترابط عاطفي بينهما! وبالتالي لن يسميها باسمها الكامل! (يعني بهاد المقطع بالذات كان لازم يكون حافظ شو لازم يساوي ومو عم يكتب كلام لزوجتو من نص مكتوب إدامو!)
- **مكرر.. مكرر**
دائماً الصورة الذهبية عند الكلام عن الكميرات.. صورة تزلج "وواحد عم ينط.." منتهى الغلاظة وعدم الإبداع غوغل.. منتهى الغلاظة.. 
- **مرة أخرى.. ضحك على البشر**
عرضت غوغل خاصية الاتزان Stabilization في جهازها الجديد.. بس.. بحياتي كلها ما شفت كميرا بترج هيك وقت ما يكون فيها Stabilization. ما بعرف ع شو كانوا مركبين الكميرا. كميرة أول موبايل إلي يلي هوي P990i كانت أحسن من يلي عرضتها غوغل.. 
- **فهيمة نقطة**
في المقطع نفسه، كانت غوغل لأول مرة فهيمة بأنها عرضت شخص يمر أثناء التصوير لتظهر لك ميزة الكميرا بالنسبة لشيء "ثابت ومتحرك" بالنسبة لك (فكرة المراقب الداخلي والخارجي في الفيزياء) 
- **ميزات لا أحد يسأل عنها**
الفكرة في أنها عرضت ميزات مثل حفظ الصور بالدقة الكاملة في تطبيق Google Photos. بصراحة الدقة الحالية لـ Google Photos جيدة جداً ولا أظن أن أحداً من المستخدمين يسأل عن مثل هكذا ميزة. هذه ميزة للقلة القليلة جداً التي لديها كميرات احترافية أو تريد معالجة الدقة الكاملة للصور. غير ذلك فهذه ميزة ...... ما بعرف شو قول.. 
- **قال يعني ذكية**
تبعاً للنقطة الماضية عرضت غوغل الـ pop-up التالية في إشارة إلى أن أجهزة آبل محدودة المساحة. من الغباء الشديد في هذه النقطة أنه يمكنك إضافة تطبيق غوغل نفسه Google Photos على آجهزة آبل وتنتهي المشكلة! إذا بدك كمان نزلو من [هنا](https://itunes.apple.com/us/app/google-photos-free-photo-video/id962194608?mt=8) إذا عندك iPhone! لك حاج غلاظة غوغل بجد! 
- **قال يعني ذكية مرة أخرى**
مرة آخرى قالت غوغل بأن تطبيق Google Duo يأتي Pre-installed أي "منزّل خالص عالجهاز" لا أستطيع إدراك مدى غباء هكذا جملة. يمكنك بكل بساطة حتى لو كنت مستخدماً "عادياً" أن تقوم بتنزيل Google Duo. هذه "الميزة قال يعني" ليست ميزة بالأصل لأنه يمكنك تنزيل التطبيق على أي جهاز أندرويد وليس جهاز Pixel فقط!! 
- **قال يعني ذكية مرة أخرى وأخرى**
مرة أخرى "ميزة" ليست بميزة.. وهي.. التحديثات التلقائية.. هل سمعتها لمليون مرة في حياتك قبلاً؟ 
- **مضحك مبكي لغوغل**
من المضحك جداً أنها صنعت وصلة خصيصاً لنقل محتويات جهازك "الآخر" (الأيفون.) 
- **تصاميم...**
تصاميم لا أعرف ماذا أقول عنها.. تصاميم من مصمم من غير عصر وزمان راسب ٤ سنين قبل ما تخرج بالزور.. 
- **غلاظات وغباءات**
من المثير للضحك أنها عرضت جهازها الجديد على أنه جهاز جديد حتى أنه بدون رقم! تخيل هذه الميزة الرهيبة! جهاز ليس لديه رقم! في هذا إشارة إلى أجهزة الأيفون التي لديها أرقام 6s, 7, etc ولكن المبكي لغوغل أن الأرقام عند آبل تعني التقدم والتطوير. هذا على الأقل ما تقوم به آبل. غوغل في كل سنة تصدر جهاز جديد باسم جديد لا أحد يعلمه قبلاً! على الأقل آبل فخورة باسم الأيفون لتبقيه وتزيد الرقم في كل سنة! هذه ميزة في سلة آبل إذاً وليست غوغل! أظن أن غوغل ليس فخورة أبداً بأجهزتها الماضية! هذا ما تعنيه غوغل بالنسبة لي في هذه الصورة!! 
- **شي بيضحك أخيراً**
مرة أخرى عرضت غوغل ميزة تعتبرها رهيبة.. وهي أنه يوجد مدخل لسماعاتك! يا للهول! بجد غلاظة مو طبيعية.. في هذا إشارة أخرى إلى جهاز iPhone 7 من آبل الذي لا يحتوي مدخل سماعات مخصص. ولكنه قرار من آبل للمضي قدماً نحو عالم خالي من الأسلاك! هذه ميزة إذاً في سلة آبل مرة أخرى!  بالنهاية ترى أن جهاز غوغل الجديد سيكون ليس جديد بالمرة بعد عدة أشهر بالمقارنة مع ما سيأتي من Samsung/ LG/ Sony/ etc.. أغلب المواصفات التي عرضتها غوغل على أنها مواصفات "خطيرة" هي بالنهاية مواصفات ستكون ضمن أجهزة أندرويد أخرى وبالتالي كالعادة.. لا إبداع ولا مواصفات ولا أداء لأجهزة الأندرويد في المستقبل القريب..
- **نقطة لا تفهمها غوغل**
النقطة التي لا تفهمها غوغل هي أن الِناس تُعَظّم أجهزة آبل والـ iPhone لأن آبل صنعت هالة حول نفسها. وصنعت منتجات بطابع خاص وبسعر خاص. ولهذا الجميع يرغب في اقتناء الـ Mac or iPhone or iPad وذلك لأن الانطباع العام لأجهزة أبل أنها أجهزة للأذكياء والغيييييكس. أجهزة غوغل والأندرويد خصوصاً منذ نشأتها كانت الفكرة أنها للعامة.. أنها لجميع الناس.. أنها مفتوحة ويمكن لأي كان أن يطور تطبيقات أو يصدر أجهزة عتادية.. كلها أمور منذ البداية حفرت في عقول الناس.. لا يمكن تغيير هذا بسهولة أبداً.. آبل صنعت لنفسها هالة.. وصنعت لنفسها سوقاً.. وصنعت لنفسها سعراً.. كل هذا جزء من الخطة وليس عبثاً.. عندما تأتي غوغل وتصدر جهاز تقارنه مع جهاز آبل كل لحظة وأخرى خلال الحدث هو شيء لا ينفع غوغل بشيء! لأنها تقارن نفسها مع غير سوق وغير سعر وغير شريحة مستخدمين!
### #هيك #رسمي
- URL: https://mohammadshaker.com/en/blog/me-drawing
- Date: 2016-10-03T00:00:00.000Z
- Tags: personal, Human-Written
فيديو يوثق لحظات الرسم والإبداع الفني لمحمد شاكر. This concise summary states the core idea clearly.
#### Content
[youtube.com](https://www.youtube.com/watch?v=mGqqX97uoik)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 29: دواء الخمس دولارات!
- URL: https://mohammadshaker.com/en/blog/5-dollar-medicine
- Date: 2016-09-29T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة ساخرة عن دواء الخمس دولارات من سلسلة اليوميات
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 30 والأخير: المياعة كنز لا يفنى!
- URL: https://mohammadshaker.com/en/blog/almayaa
- Date: 2016-09-29T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة طريفة عن المياعة كنز لا يفنى من اليوميات البصرية
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 28: وهم الطبيعي المحزن!
- URL: https://mohammadshaker.com/en/blog/sad-normaility-illusion
- Date: 2016-09-28T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
Read more here! Thinking, Fast and Slow, by Daniel Kahneman.
#### Content
 Read more here!
- Thinking, Fast and Slow, by [Daniel Kahneman](https://www.goodreads.com/author/show/72401.Daniel_Kahneman).
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### اليوم 27: عقلك السريع البطيء!
- URL: https://mohammadshaker.com/en/blog/thinking-fast-and-slow
- Date: 2016-09-27T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
Read more here! Thinking, Fast and Slow, by Daniel Kahneman.
#### Content
 Read more here!
- Thinking, Fast and Slow, by [Daniel Kahneman](https://www.goodreads.com/author/show/72401.Daniel_Kahneman).
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### اليوم 26: أعط من أضافك في بيته نقوداً... وشوف شو بيصير!
- URL: https://mohammadshaker.com/en/blog/give-your-host-money
- Date: 2016-09-26T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة طريفة عن إعطاء المضيف نقوداً وما يحدث بعدها
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 25: ظروفك بتنبعت للريح!
- URL: https://mohammadshaker.com/en/blog/your-mails-to-air
- Date: 2016-09-25T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة ساخرة عن الظروف التي تنبعت للريح من سلسلة اليوميات
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 24: ضميرك المخبص! ضمير عند الحاجة!
- URL: https://mohammadshaker.com/en/blog/your-concious
- Date: 2016-09-24T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة عن الضمير المخبص - ضمير عند الحاجة!. This concise summary states the core idea clearly.
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 23: هذا المقال مجاااانيييي!
- URL: https://mohammadshaker.com/en/blog/free-post
- Date: 2016-09-23T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة ساخرة عن عقلية التفكير بالمجان من سلسلة اليوميات
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 22: شوربة أخلاق!
- URL: https://mohammadshaker.com/en/blog/moral-soup
- Date: 2016-09-22T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة ساخرة بعنوان شوربة أخلاق من سلسلة اليوميات
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 21: ما علاقة أيكيا بالدين؟!
- URL: https://mohammadshaker.com/en/blog/ikea-and-religion
- Date: 2016-09-21T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة طريفة عن العلاقة غير المتوقعة بين أيكيا والدين
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 20: صايرة وصايرة!
- URL: https://mohammadshaker.com/en/blog/happening-either-way
- Date: 2016-09-20T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة عن الأمور التي تحدث بطريقة أو بأخرى. This concise summary states the core idea clearly.
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 19: هل نحب الضريبة أم الضريبة لا تحبنا؟
- URL: https://mohammadshaker.com/en/blog/like-bill-no-bill
- Date: 2016-09-19T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
- Predictably Irrational: The Hidden Forces That Shape Our Decisions, by Dan Ariely.
#### Content
 Read more here!
- Predictably Irrational: The Hidden Forces That Shape Our Decisions, by [Dan Ariely](https://www.goodreads.com/author/show/788461.Dan_Ariely).
- Thinking, Fast and Slow, by [Daniel Kahneman](https://www.goodreads.com/author/show/72401.Daniel_Kahneman).
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### اليوم 18: من منظور علمي، بدك حسم أم تجنب ضريبة؟
- URL: https://mohammadshaker.com/en/blog/want-discount-or-no-bill
- Date: 2016-09-18T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة ساخرة عن الحسم مقابل تجنب الضريبة من منظور علمي
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 17: الرّيس والهالة!
- URL: https://mohammadshaker.com/en/blog/president-and-halo
- Date: 2016-09-17T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة ساخرة عن الرّيس والهالة المحيطة به. This concise summary states the core idea clearly.
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 16: حط مصروفك عند أختك!
- URL: https://mohammadshaker.com/en/blog/give-your-sister-your-income
- Date: 2016-09-16T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة ساخرة عن وضع المصروف عند الأخت وما يحدث بعدها
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 15: مواعيد لمّا بلا مواعيد؟
- URL: https://mohammadshaker.com/en/blog/appointments-or-not
- Date: 2016-09-15T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة عن الفرق بين المواعيد والحياة بلا مواعيد
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 14: اختبار مارشميلو يتنبأ بمستقبلك!
- URL: https://mohammadshaker.com/en/blog/marshmallow-test-for-you
- Date: 2016-09-15T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
- The Marshmallow Test: Mastering Self-Control by Walter Mischel.
#### Content
 Read more here!
- [The Marshmallow Test: Mastering Self-Control](https://www.goodreads.com/book/show/20454074-the-marshmallow-test) by [Walter Mischel](https://www.goodreads.com/author/show/724482.Walter_Mischel).
- Predictably Irrational: The Hidden Forces That Shape Our Decisions, by [Dan Ariely](https://www.goodreads.com/author/show/788461.Dan_Ariely).
- Thinking, Fast and Slow, by [Daniel Kahneman](https://www.goodreads.com/author/show/72401.Daniel_Kahneman).
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### اليوم 13: كون زاغاروس وتحدى القيود!
- URL: https://mohammadshaker.com/en/blog/be-zagarous
- Date: 2016-09-13T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة تحفيزية: كون زاغاروس وتحدى القيود!. This concise summary states the core idea clearly.
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 11: فن الامتنان
- URL: https://mohammadshaker.com/en/blog/being-grateful
- Date: 2016-09-13T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة عن أهمية الامتنان والشكر في الحياة اليومية
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم 12: بدون دسم!
- URL: https://mohammadshaker.com/en/blog/no-fat
- Date: 2016-09-13T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة طريفة عن الحياة بدون دسم من سلسلة اليوميات
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم ٠٩: قلل من خيارك!
- URL: https://mohammadshaker.com/en/blog/reduce-your-cucumber
- Date: 2016-09-12T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
- The Paradox of Choice: Why More Is Less by Barry Schwartz.
#### Content
 Read more here!
- [The Paradox of Choice: Why More Is Less](https://www.goodreads.com/book/show/10639.The_Paradox_of_Choice) by [Barry Schwartz](https://www.goodreads.com/author/show/6957.Barry_Schwartz).
- [The Wisdom of Crowds](https://www.goodreads.com/book/show/68143.The_Wisdom_of_Crowds) by [James Surowiecki](https://www.goodreads.com/author/show/38391.James_Surowiecki).
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### اليوم ١٠: لماذا يضحك الجمهور في Friends!
- URL: https://mohammadshaker.com/en/blog/why-people-laugh-friends
- Date: 2016-09-12T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
- Presence: Bringing Your Boldest Self to Your Biggest Challenges, by Amy Cuddy.
#### Content
 Read more here!
- [Presence: Bringing Your Boldest Self to Your Biggest Challenges](https://www.goodreads.com/book/show/25066556-presence), by [Amy Cuddy](https://www.goodreads.com/author/show/6862909.Amy_Cuddy).
- [Steal Like an Artist: 10 Things Nobody Told You About Being Creative](https://www.goodreads.com/book/show/13099738-steal-like-an-artist), by [Austin Kleon](https://www.goodreads.com/author/show/2985039.Austin_Kleon).
- [The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work](https://www.goodreads.com/book/show/9484114-the-happiness-advantage), by [Shawn Achor](https://www.goodreads.com/author/show/4024160.Shawn_Achor).
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### اليوم ٠٨: دخلك لا يعني سعادتك!
- URL: https://mohammadshaker.com/en/blog/income-vs-happiness
- Date: 2016-09-10T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
إقرأ المزيد في الكتب الرائعة! Predictably Irrational: The Hidden Forces That Shape Our Decisions, by Dan Ariely Thinking, Fast and Slow, by Daniel
#### Content

إقرأ المزيد في الكتب الرائعة!
- Predictably Irrational: The Hidden Forces That Shape Our Decisions, by [Dan Ariely](https://www.goodreads.com/author/show/788461.Dan_Ariely)
- Thinking, Fast and Slow, by [Daniel Kahneman](https://www.goodreads.com/author/show/72401.Daniel_Kahneman)
- The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work, by [Shawn Achor](https://www.goodreads.com/author/show/4024160.Shawn_Achor)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### اليوم ٠٧: عهد النوكيا ما زال قائماً! - الجزء الثاني
- URL: https://mohammadshaker.com/en/blog/nokia-part2
- Date: 2016-09-10T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
إقرأ المزيد في الكتاب! The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work, by Shawn Achor
#### Content

إقرأ المزيد في الكتاب!
- The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work, by [Shawn Achor](https://www.goodreads.com/author/show/4024160.Shawn_Achor)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### اليوم ٠٦: عهد النوكيا ما زال قائماً! - الجزء الأول
- URL: https://mohammadshaker.com/en/blog/nokia-part1
- Date: 2016-09-09T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
إقرأ المزيد في الكتاب! The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work, by Shawn Achor
#### Content

إقرأ المزيد في الكتاب!
- The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work, by [Shawn Achor](https://www.goodreads.com/author/show/4024160.Shawn_Achor)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### اليوم ٠٥!
- URL: https://mohammadshaker.com/en/blog/wisdom
- Date: 2016-09-08T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة تحمل حكمة حياتية من سلسلة الرسومات اليومية
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم٠٤: ليست تجربة بافلوف آخرى!
- URL: https://mohammadshaker.com/en/blog/dogs-and-meh
- Date: 2016-09-07T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
مرة أخرى مع الكتاب الرائع: The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work, by Shawn Achor.
#### Content

مرة أخرى مع الكتاب الرائع:
- [The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work](https://www.goodreads.com/book/show/9484114-the-happiness-advantage), by [Shawn Achor](https://www.goodreads.com/author/show/4024160.Shawn_Achor).
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### اليوم٠٣: اشعل نفسك وتحدى العطالة!
- URL: https://mohammadshaker.com/en/blog/activation-energy-01
- Date: 2016-09-06T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
اقرا المزيد في الكتب التالية: What the Dog Saw and Other Adventures by Malcolm Gladwell Flow: The Psychology of Optimal Experience by Mihaly
#### Content

اقرا المزيد في الكتب التالية:
- [What the Dog Saw and Other Adventures](https://www.goodreads.com/book/show/6516450-what-the-dog-saw-and-other-adventures) by [Malcolm Gladwell](https://www.goodreads.com/author/show/1439.Malcolm_Gladwell)
- [Flow: The Psychology of Optimal Experience](https://www.goodreads.com/book/show/66354.Flow) by [Mihaly Csikszentmihalyi](https://www.goodreads.com/author/show/27446.Mihaly_Csikszentmihalyi)
- [The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work](https://www.goodreads.com/book/show/9484114-the-happiness-advantage) by [Shawn Achor](https://www.goodreads.com/author/show/4024160.Shawn_Achor)
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### اليوم ٠٢: أنا لا أعمل، أنا أميرة!
- URL: https://mohammadshaker.com/en/blog/am-a-princess
- Date: 2016-09-05T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
قصة مصورة ساخرة: أنا لا أعمل، أنا أميرة!. This concise summary states the core idea clearly.
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### اليوم ٠١: كيف تستفاد من نظرية النافذة المكسورة!
- URL: https://mohammadshaker.com/en/blog/broken-window-theory
- Date: 2016-09-04T00:00:00.000Z
- Tags: visual-storytelling, art, Human-Written
إذا أعجبتك القصة فيمكنك القراءة أكثر في الكتابين الرائعين:
#### Content

إذا أعجبتك القصة فيمكنك القراءة أكثر في الكتابين الرائعين:
- [The Happiness Advantage: The Seven Principles of Positive Psychology That Fuel Success and Performance at Work](https://www.goodreads.com/book/show/9484114-the-happiness-advantage), by [Shawn Achor](https://www.goodreads.com/author/show/4024160.Shawn_Achor).
- [The Tipping Point: How Little Things Can Make a Big Difference](https://www.goodreads.com/book/show/2612.The_Tipping_Point), by [Malcolm Gladwell](https://www.goodreads.com/author/show/1439.Malcolm_Gladwell)
وعن نظرية النافذة المكسورة هنا:
- [https://en.wikipedia.org/wiki/Broken\_windows\_theory](http://Broken Window Theory)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### جن!
- URL: https://mohammadshaker.com/en/blog/%d8%ac%d9%86
- Date: 2016-09-01T00:00:00.000Z
- Tags: Human-Written
رسمة من مجموعة الأعمال الفنية للفنان محمد شاكر. This concise summary states the core idea clearly.
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### A Short Movie About Biking - With Adaptive Music / فلم قصير عن "البسكلة" مع موسيقا متكيفة!
- URL: https://mohammadshaker.com/en/blog/a-short-movie-about-biking-with-adaptive-music
- Date: 2016-08-24T00:00:00.000Z
- Tags: thoughts, personal, Human-Written
Since I haven't done much with music lately (just my games earlier this year) I have decided to come up with something new. Just trying to brush my skill
#### Content
Since I haven't done much with music lately (just my games earlier this year) I have decided to come up with something new. Just trying to brush my skill in film making and adaptive music with my first ever short movie, time-lapsed. Enjoy and let me know what you think!
[youtube.com](https://www.youtube.com/watch?v=hsL5qyRhbe8)
### رسم الأشخاص في الطائرة / Drawing People on the Plane - Pack 06
- URL: https://mohammadshaker.com/en/blog/%d8%b1%d8%b3%d9%85-%d8%a7%d9%84%d8%a3%d8%b4%d8%ae%d8%a7%d8%b5-%d9%81%d9%8a-%d8%a7%d9%84%d8%b7%d8%a7%d8%a6%d8%b1%d8%a9-drawing-people-on-the-plane-pack-06
- Date: 2016-08-22T00:00:00.000Z
- Tags: drawing, art, Human-Written
رسومات سريعة للأشخاص أثناء السفر بالطائرة - الحزمة السادسة
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### رسم الأشخاص في الطائرة / Drawing People on the Plane - Pack 05
- URL: https://mohammadshaker.com/en/blog/%d8%b1%d8%b3%d9%85-%d8%a7%d9%84%d8%a3%d8%b4%d8%ae%d8%a7%d8%b5-%d9%81%d9%8a-%d8%a7%d9%84%d8%b7%d8%a7%d8%a6%d8%b1%d8%a9-drawing-people-on-the-plane-pack-05
- Date: 2016-08-12T00:00:00.000Z
- Tags: drawing, art, Human-Written
رسومات سريعة للأشخاص أثناء السفر بالطائرة - الحزمة الخامسة
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### خليك ع طبيعتك!
- URL: https://mohammadshaker.com/en/blog/%d8%ae%d9%84%d9%8a%d9%83-%d8%b9-%d8%b7%d8%a8%d9%8a%d8%b9%d8%aa%d9%83
- Date: 2016-08-09T00:00:00.000Z
- Tags: drawing, art, Human-Written
رسمة تعبر عن أهمية البقاء على طبيعتك والأصالة الشخصية
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### رسم الأشخاص في الطائرة / Drawing People on the Plane - Pack 04
- URL: https://mohammadshaker.com/en/blog/%d8%b1%d8%b3%d9%85-%d8%a7%d9%84%d8%a3%d8%b4%d8%ae%d8%a7%d8%b5-%d9%81%d9%8a-%d8%a7%d9%84%d8%b7%d8%a7%d8%a6%d8%b1%d8%a9-drawing-people-on-the-plane-pack-04
- Date: 2016-08-03T00:00:00.000Z
- Tags: personal, Human-Written
رسومات سريعة للأشخاص أثناء السفر بالطائرة - الحزمة الرابعة
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### On Discipline and Focus - Jony Ive
- URL: https://mohammadshaker.com/en/blog/the-discipline-to-turn-your-back-on-something-you-believe-in-passionately-so-you-can-apply-yourself-to-whats-at-hand-is-really-remarkable-jony-ive
- Date: 2016-07-20T00:00:00.000Z
- Tags: personal, Human-Written
"The discipline to turn your back on something you believe in passionately, so you can apply yourself to what's at hand, is really remarkable"
#### Content
"The discipline to turn your back on something you believe in passionately, so you can apply yourself to what's at hand, is really remarkable"
\-Jony Ive
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Another Monsters Pack!
- URL: https://mohammadshaker.com/en/blog/another-monsters-pack
- Date: 2016-07-13T00:00:00.000Z
- Tags: personal, Human-Written
Another collection of creative monster character illustrations
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### رسم الأشخاص في الطائرة / Drawing People on the Plane - Pack 03
- URL: https://mohammadshaker.com/en/blog/%d8%b1%d8%b3%d9%85-%d8%a7%d9%84%d8%a3%d8%b4%d8%ae%d8%a7%d8%b5-%d9%81%d9%8a-%d8%a7%d9%84%d8%b7%d8%a7%d8%a6%d8%b1%d8%a9-drawing-people-on-the-plane-pack-03
- Date: 2016-07-12T00:00:00.000Z
- Tags: drawing, art, Human-Written
رسومات سريعة للأشخاص أثناء السفر بالطائرة - الحزمة الثالثة
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### رسم الأشخاص في الطائرة / Drawing People on the Plane - Pack 02
- URL: https://mohammadshaker.com/en/blog/%d8%b1%d8%b3%d9%85-%d8%a7%d9%84%d8%a3%d8%b4%d8%ae%d8%a7%d8%b5-%d9%81%d9%8a-%d8%a7%d9%84%d8%b7%d8%a7%d8%a6%d8%b1%d8%a9-drawing-people-on-the-plane-pack-02
- Date: 2016-07-02T00:00:00.000Z
- Tags: drawing, art, Human-Written
رسومات سريعة للأشخاص أثناء السفر بالطائرة - الحزمة الثانية
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### Caricatures of My Friends!
- URL: https://mohammadshaker.com/en/blog/caricatures-of-my-friends
- Date: 2016-07-01T00:00:00.000Z
- Tags: drawing, art, Human-Written
Collection of caricature drawings of friends by Mohammad Shaker
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### Monsters Pack
- URL: https://mohammadshaker.com/en/blog/monsters-pack
- Date: 2016-07-01T00:00:00.000Z
- Tags: drawing, art, Human-Written
Collection of creative monster character designs and illustrations
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### نور على نور
- URL: https://mohammadshaker.com/en/blog/%d9%86%d9%88%d8%b1-%d8%b9%d9%84%d9%89-%d9%86%d9%88%d8%b1
- Date: 2016-06-30T00:00:00.000Z
- Tags: drawing, art, Human-Written
رسمة فنية تحمل عنوان نور على نور من أعمال محمد شاكر
#### Content

## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### رسم الأشخاص في الطائرة / Drawing People on the Plane - Pack 01
- URL: https://mohammadshaker.com/en/blog/%d8%b1%d8%b3%d9%85-%d8%a7%d9%84%d8%a3%d8%b4%d8%ae%d8%a7%d8%b5-%d9%81%d9%8a-%d8%a7%d9%84%d8%b7%d8%a7%d8%a6%d8%b1%d8%a9-drawing-people-on-the-plane-pack-01
- Date: 2016-06-26T00:00:00.000Z
- Tags: personal, Human-Written
بما أنني كنت أسافر بكثرة في الفترة الماضية فمن أحد الطرق التي كنت استمتع بها في الطائرة هي أن أرسم الأشخاص أمامي أو الشخصيات في المجلات :P
#### Content
بما أنني كنت أسافر بكثرة في الفترة الماضية فمن أحد الطرق التي كنت استمتع بها في الطائرة هي أن أرسم الأشخاص أمامي أو الشخصيات في المجلات :P
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Creating a Visual Book - A Story of Good and Evil
- URL: https://mohammadshaker.com/en/blog/creating-a-visual-book-a-story-of-good-and-evil
- Date: 2016-06-16T00:00:00.000Z
- Tags: design, Human-Written
For the final project in the Image Making course I did in Coursera, we had to design a visual book (cover, 3 spreads (3x2 pages) and a back cover.) I
#### Content
For the final project in the Image Making course I did in Coursera, we had to design a visual book (cover, 3 spreads (3x2 pages) and a back cover.) I really did enjoy this one. My first "book"! I don't need to talk anymore. Here it is. and the back cover You can see the full "book": [MShaker\_Full\_Narrative](/blog-images/2016/06/mshaker_full_narrative.pdf "MShaker_Full_Narrative").
### DT #3: Small Details Make All the Difference
- URL: https://mohammadshaker.com/en/blog/dt-3-small-details-make-all-the-difference
- Date: 2016-06-16T00:00:00.000Z
- Tags: design, reading, Human-Written
Reading this book: "Web Form Design: Filling in the Blanks" and greeted with this #innovative progress bar design on top. Love it!
#### Content
Reading this book: "Web Form Design: Filling in the Blanks" and greeted with this #innovative progress bar design on top. Love it!

## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Drawing - Guacamelee
- URL: https://mohammadshaker.com/en/blog/drawing-guacamelee
- Date: 2016-06-14T00:00:00.000Z
- Tags: drawing, art, Human-Written
If you like this, you can find more on my instagram.
#### Content
If you like this, you can find more on [my instagram](http://instagram.com/mohammadshakergtr/).
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### DT#2: Great Read on Design and UX
- URL: https://mohammadshaker.com/en/blog/dt2-great-read-on-design-and-ux
- Date: 2016-06-14T00:00:00.000Z
- Tags: design, Human-Written
A great guide on design and UX from ustwo, the creator of Monument Valley. (Download the full pdf here.) !free-design-ebooks-ppp3-620
#### Content
A great guide on design and UX from [ustwo](https://ustwo.com/), the creator of Monument Valley. (Download the full pdf [here.](http://cdn.ustwo.com/PPP/PP3.pdf)) 
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### Drawing - Simpsons Gangsters
- URL: https://mohammadshaker.com/en/blog/drawing-simpsons-gangsters
- Date: 2016-06-10T00:00:00.000Z
- Tags: drawing, art, Human-Written
If you like this, you can find more on my instagram.
#### Content
If you like this, you can find more on [my instagram](http://instagram.com/mohammadshakergtr/).
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Drawing - Dirty Animals
- URL: https://mohammadshaker.com/en/blog/drawing-dirty-animals
- Date: 2016-06-08T00:00:00.000Z
- Tags: drawing, art, Human-Written
If you like this, you can find more on my instagram.
#### Content
If you like this, you can find more on [my instagram](http://instagram.com/mohammadshakergtr/).
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Ropossum V1 Is Now Open-Sourced!
- URL: https://mohammadshaker.com/en/blog/ropossum-v1-is-now-open-sourced
- Date: 2016-06-08T00:00:00.000Z
- Tags: games, game-development, Human-Written
I've taken the decision to make Ropossum, my authoring tool for generating content for physics-based game, open-sourced! It's the version of late
#### Content
 I've taken the decision to make Ropossum, my authoring tool for generating content for physics-based game, open-sourced! It's the version of late 2012-early 2013, not my latest work. But it's a great version nonetheless! You can get the physics-engine, CRUST, now on [github](https://github.com/ZGTR/CRUST-Physics-Engine).
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Drawing - The Cool Guy
- URL: https://mohammadshaker.com/en/blog/drawing-the-cool-guy
- Date: 2016-06-07T00:00:00.000Z
- Tags: drawing, art, Human-Written
If you like this, you can find more on my instagram.
#### Content
 If you like this, you can find more on [my instagram](http://instagram.com/mohammadshakergtr/).
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### Drawings - Ramadan Kareem! رمضان كريم!
- URL: https://mohammadshaker.com/en/blog/drawings-ramadan-kareem-%d8%b1%d9%85%d8%b6%d8%a7%d9%86-%d9%83%d8%b1%d9%8a%d9%85
- Date: 2016-06-07T00:00:00.000Z
- Tags: drawing, art, Human-Written
Ramadan greeting artwork celebrating the holy month with artistic illustrations
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### Another Agent to Solve Cut the Rope - Ropossum Authoring Tool - Projection-Based Approach
- URL: https://mohammadshaker.com/en/blog/another-agent-to-solve-cut-the-rope-ropossum-authoring-tool-projection-based-approach
- Date: 2016-06-06T00:00:00.000Z
- Tags: games, game-development, Human-Written
This video is about my second agent to solve Cut the Rope levels. The Projection-based Approach.
#### Content
This video is about my second agent to solve Cut the Rope levels. The Projection-based Approach. This newer agent takes only 0.1sec to check a level for playability. Read more on this and my other agents on my website: www.mohammadshaker.com/ropossum.html
[youtube.com](https://www.youtube.com/watch?v=mZK9uMprew4)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Drawing - Quick Sketches
- URL: https://mohammadshaker.com/en/blog/drawing-quick-sketches
- Date: 2016-06-05T00:00:00.000Z
- Tags: drawing, art, Human-Written
If you like this, you can find more on my instagram.
#### Content
If you like this, you can find more on [my instagram](http://instagram.com/mohammadshakergtr/).
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Book of the Week
- URL: https://mohammadshaker.com/en/blog/book-of-the-week
- Date: 2016-06-03T00:00:00.000Z
- Tags: books, reading, Human-Written
So, I this is my book of the week: Graphics Design Solution's. It's a truly lovely journey on Graphic Design. Recommended.
#### Content
So, I this is my book of the week: [Graphics Design Solution's](https://www.goodreads.com/book/show/15849406-graphic-design-solutions). It's a truly lovely journey on Graphic Design. Recommended. If you wanna read more about it, you know where to look. [](/blog-images/2016/06/img_0550.png)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Drawing - Tree Man
- URL: https://mohammadshaker.com/en/blog/drawing-tree-man
- Date: 2016-06-03T00:00:00.000Z
- Tags: drawing, art, Human-Written
If you like this, you can find more on my instagram.
#### Content
[](/blog-images/2016/05/untitled-6-01.png) If you like this, you can find more on [my instagram](http://instagram.com/mohammadshakergtr/).
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### اسمع يا رجل: الحياة نهر وكل يغترف منه بحجم فنجانه.. فنجانك صغير
- URL: https://mohammadshaker.com/en/blog/%d8%a7%d8%b3%d9%85%d8%b9-%d9%8a%d8%a7-%d8%b1%d8%ac%d9%84-%d8%a7%d9%84%d8%ad%d9%8a%d8%a7%d8%a9-%d9%86%d9%87%d8%b1-%d9%88%d9%83%d9%84-%d9%8a%d8%ba%d8%aa%d8%b1%d9%81-%d9%85%d9%86%d9%87-%d8%a8%d8%ad
- Date: 2016-06-02T00:00:00.000Z
- Tags: books, reading, Human-Written
> ذهنك سعدان ملدوغ. تشبه هذا الفقير الهندي الذي جاء إلى دير بوذي بحثاً عن إنارة روحه… وقعد يروي للراهب عن ماضيه، وعذابه، وذكرياته، وعن حاجته للتنوير،
#### Content
> ذهنك سعدان ملدوغ. تشبه هذا الفقير الهندي الذي جاء إلى دير بوذي بحثاً عن إنارة روحه… وقعد يروي للراهب عن ماضيه، وعذابه، وذكرياته، وعن حاجته للتنوير، ويروي، ويروي، ويروي، والراهب يصغي ويصبُّ الشاي في فنجان على الطاولة. طفح الفنجان، وسال الشاي على الخشب والأرض، والراهب يصبُّ، والرجل يروي ويروي ويروي، إلى حدِّ الملل، وأخيراً انتبه فقال للراهب: طفح الشاي من الفنجان، لماذا تواصل الصَّب فيه؟ فردَّ الراهب: ( ذهنك يشبه هذا الفنجان، مليء، أفرغها مما فيه، كي أصبَّ لك شاياً جديداً)
من الكتاب الرائع "[الضوء الأزرق](https://www.goodreads.com/book/show/13190747)" لـ [حسين البرغوثي](https://www.goodreads.com/author/show/1382148._)
### Drawing - The Old Guy
- URL: https://mohammadshaker.com/en/blog/drawing-the-old-guy
- Date: 2016-06-01T00:00:00.000Z
- Tags: drawing, art, Human-Written
This one is not drawn on a board. This is just drawn on a white paper with inverted colors :P. If you like this, you can find more on my instagram.
#### Content
 This one is not drawn on a board. This is just drawn on a white paper with inverted colors :P. If you like this, you can find more on [my instagram](http://instagram.com/mohammadshakergtr/).
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Drawing - People Sitting
- URL: https://mohammadshaker.com/en/blog/drawing-people-sitting
- Date: 2016-05-31T00:00:00.000Z
- Tags: drawing, art, Human-Written
If you like this, you can find more on my instagram.
#### Content
If you like this, you can find more on [my instagram](http://instagram.com/mohammadshakergtr/).
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### An Agent to Solve Cut the Rope - First Attempt - Simulation-Based Approach
- URL: https://mohammadshaker.com/en/blog/an-agent-to-solve-cut-the-rope-first-attempt-simulation-based-approach
- Date: 2016-05-30T00:00:00.000Z
- Tags: games, game-development, Human-Written
I know I should've published this while ago, but, anyway. This video is about my first agent to solve Cut the Rope levels in late 2012 - early 2013, the
#### Content
I know I should've published this while ago, but, anyway. This video is about my first agent to solve Cut the Rope levels in late 2012 - early 2013, the Simulation-based approach. My later attempts with newer, more clever agents would minimise the time it takes the agent to solve the level from about 30 sec for the simulation-based approach to 0.1 sec for the projection agent for instance. Read more on my website: www.mohammadshaker.com/ropossum.html
[youtube.com](https://www.youtube.com/watch?v=Ck3TN-Ld9Y0)
### Creativity Is Very Much Like Literacy
- URL: https://mohammadshaker.com/en/blog/creativity-is-very-much-like-literacy
- Date: 2016-05-29T00:00:00.000Z
- Tags: quick-read, Human-Written
\"One myth is that only special people are creative. This is not true. Everyone is born with tremendous capacities for creativity.
#### Content
"One myth is that only special people are creative. This is not true. Everyone is born with tremendous capacities for creativity. The trick is to develop these capacities. Creativity is very much like literacy. We take it for granted that nearly everybody can learn to read and write. If a person can’t read or write, you don’t assume that this person is incapable of it, just that he or she hasn’t learned how to do it. The same is true of creativity. When people say they’re not creative, it’s often because they don’t know what’s involved or how creativity works in practice."
Robinson, Ken. “The Element.”
### Graphic Design Specialisation on Coursera
- URL: https://mohammadshaker.com/en/blog/graphic-design-specialisation-on-coursera
- Date: 2016-05-29T00:00:00.000Z
- Tags: courses, learning, Human-Written
Since I'm crazy about everything about design, I'm just into the fourth course in this specialization on Coursera.
#### Content
Since I'm crazy about everything about design, I'm just into the fourth course in [this specialization on Coursera](https://www.coursera.org/specializations/graphic-design). Just finished the first three with all 100% so I'm pretty excited to continue to the Capstone Project later in June. If you are really into design you should definitely start there. A very good specialisation from CalArts, California. Here's one of the crazy things I did. I'll try to publish the rest of my work work during these courses here, soon. 
### Introduction to Typography - Finished with 100%!
- URL: https://mohammadshaker.com/en/blog/introduction-to-typography-finished
- Date: 2016-05-22T00:00:00.000Z
- Tags: design, Human-Written
I LOVE Typography. I LOVE it. This short course was fun more than informative for me I think.
#### Content
I LOVE Typography. I LOVE it. [This short course](https://www.coursera.org/learn/typography/home/welcome) was fun more than informative for me I think. I've read books on typography before but thought that I should do some kind of systemic work with the pro CALArts. The course is both lightweight and fun. Definitely a recommended.
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Drawings - Piano Love.
- URL: https://mohammadshaker.com/en/blog/drawings-piano-love
- Date: 2016-05-18T00:00:00.000Z
- Tags: drawing, art, Human-Written
Artistic expression of the love and passion for piano music
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### Introduction to Imagemaking - Finished with 100%!
- URL: https://mohammadshaker.com/en/blog/introduction-to-imagemaking-finished
- Date: 2016-05-15T00:00:00.000Z
- Tags: design, Human-Written
Just finished Introduction to Imagemaking on Coursera! Had a great fun doing it!!Screen Shot 2016-06-16 at 14.32.55
#### Content
Just finished [Introduction to Imagemaking](https://www.coursera.org/learn/image-making/home/welcome) on Coursera! Had a great fun doing it!
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### Fundamentals of Graphic Design - Finished with 100%!
- URL: https://mohammadshaker.com/en/blog/fundamentals-of-graphic-design-finished
- Date: 2016-05-08T00:00:00.000Z
- Tags: design, Human-Written
This intro course to the Graphic Design Specialisation on Coursera seems very basic, maybe just to keep things cool for beginners.
#### Content
This [intro course](https://www.coursera.org/learn/fundamentals-of-graphic-design/home/welcome) to the [Graphic Design Specialisation on Coursera](https://www.coursera.org/specializations/graphic-design) seems very basic, maybe just to keep things cool for beginners. Being more advanced on the subject I felt a bit dull. Just want to do this to get my graphic design experience from the ground up, the right way with CALArts.
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Caricatures of Friends!
- URL: https://mohammadshaker.com/en/blog/caricatures-of-friends
- Date: 2016-05-06T00:00:00.000Z
- Tags: drawing, art, Human-Written
Was just drawing caricatures of my friends lately! Check moaaaaar on my instagram account.
#### Content
Was just drawing caricatures of my friends lately! Check moaaaaar on my [instagram](http://instagram.com/mohammadshakergtr/) account.
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Drawings - Don't Know What Happened with the Pants.
- URL: https://mohammadshaker.com/en/blog/drawings-dont-know-what-happened-with-the-pants
- Date: 2016-04-07T00:00:00.000Z
- Tags: drawing, art, Human-Written
Humorous character sketch with an interesting wardrobe malfunction
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### Drawings - Hefty Hair
- URL: https://mohammadshaker.com/en/blog/drawings-hefty-hair
- Date: 2016-04-05T00:00:00.000Z
- Tags: drawing, art, Human-Written
Character illustration featuring bold and voluminous hairstyle
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### My Books of the Year Lists [12-15]
- URL: https://mohammadshaker.com/en/blog/my-books-of-the-year-lists-12-15
- Date: 2016-02-01T00:00:00.000Z
- Tags: books, reading, startups, Human-Written
Here you go, my Books of the Year lists \\[All Books\\] 2015 Elon Musk: Inventing the Future 100 Things Every Designer Needs to Know about People Zero to
#### Content
## Here you go, my Books of the Year lists \[[All Books](https://www.goodreads.com/user/show/11193121-mohammad-shaker)\]
**2015**
1. Elon Musk: Inventing the Future
2. 100 Things Every Designer Needs to Know about People
3. Zero to One: Notes on Startups, or How to Build the Future
**2014**
1. Universal Principles of Design
2. The Visual Display of Quantitative Information
3. Made to Stick: Why Some Ideas Survive and Others Die
4. Gone Girl
**2013**
1. Steve Jobs by Walter Isaacson
2. Rework
3. A Brief History Of Time: From Big Bang To Black Holes
**2012**
1. Design Patterns: Elements of Reusable Object-Oriented Software
2. 1984
3. The Science of Digital Media
### Reactive Music Canvas - إظهار الموسيقا بشكل محسوس
- URL: https://mohammadshaker.com/en/blog/reactive-music-canvas-%d8%a5%d8%b8%d9%87%d8%a7%d8%b1-%d8%a7%d9%84%d9%85%d9%88%d8%b3%d9%8a%d9%82%d8%a7-%d8%a8%d8%b4%d9%83%d9%84-%d9%85%d8%ad%d8%b3%d9%88%d8%b3
- Date: 2015-09-06T00:00:00.000Z
- Tags: demo, Game, Music, PCG, Prototype, Rhythm-based, Human-Written
بدك "تشوف" الموسيقا؟ ما بعرف تماماً شو طلع معي. بس المهم إنو "شي" بيأرجيك الموسيقا.
#### Content
[youtube.com](https://www.youtube.com/watch?v=17Bs4xFZq2c)
بدك "تشوف" الموسيقا؟ ما بعرف تماماً شو طلع معي. بس المهم إنو "شي" بيأرجيك الموسيقا.
[#Music](https://www.facebook.com/hashtag/music?source=feed_text&story_id=894197047284575) [#Visualisation](https://www.facebook.com/hashtag/visualisation?source=feed_text&story_id=894197047284575). A "breathing" concrete. A Reactive Music Canvas. Call it what you want. The nice thing about it is that it's coded in a couple of hours and I love the result and how it "breathes" music. Hope you like it too!
### ريادة أم لا ريادة للعرب؟ تظاهر بالريادة؟ أم "لف القصة على جنب"؟
- URL: https://mohammadshaker.com/en/blog/%d8%b1%d9%8a%d8%a7%d8%af%d8%a9-%d8%a3%d9%85-%d9%84%d8%a7-%d8%b1%d9%8a%d8%a7%d8%af%d8%a9-%d9%84%d9%84%d8%b9%d8%b1%d8%a8%d8%9f
- Date: 2015-09-03T00:00:00.000Z
- Tags: entrepreneurship, idea, Read, Human-Written
> \"المشكلة تأتي من عدة مشاكل متجذرة في ثقافتنا من إنه \"يلي حلال علينا حرام على غيرنا.\" كل هذا ينبع من الجهل المدقع المرير بكل شيء نحن العرب.
#### Content
[](/blog-images/2015/09/idea.png)
> "المشكلة تأتي من عدة مشاكل متجذرة في ثقافتنا من إنه "يلي حلال علينا حرام على غيرنا." كل هذا ينبع من الجهل المدقع المرير بكل شيء نحن العرب. مجتمع يريد أن يعرف كل شيء عن كل شيء بدون أن يتحرك ويعمل شعرة هي المشكلة. يصدق أي شيء يمر على آذانه لأنه لم يتعود على أي محاكمة عقلية بنفسه، بذاته، بكيانه الذاتي الشخصي الواحد. مجتمع يتبع نظرية القطيع. الجري وراء "القوي" بحق أو بغير حق. بغض النظر عن كلمة "القوي". يمكن أن نضع كلمة "الفهيم" عوضاً عن "القوي." ما يقوله "الأفهم مني" هو الصحيح دائماً لأنه: "أفهم مني! شبنا!". أو لأن هذا هو ما جاء به آباؤنا. هذا ما علمني به أمي وأبي كان صحيحاً أم لا. هذا هو العرف. هذا هو السائد بين الناس ولا أستطيع أنا مجاراته لأنني عندها سأكون "خارج المجموعة." هذا الأمر نتبعه في كل شيء. من أبسط أمور حياتنا لأعقدها. في حياتنا الشخصية، الزوجية، الدينية، الدنيوية، كل شيء. تعودنا من صغرنا على هذا التفكير السمج القميء! تعودنا أن لا نتعب أنفسنا بالتفكير. "خلي التفكير ع غيري." اكتشاف الحقيقة لا تأتي بالوراثة. اكتشاف الحقيقة هي واجب على كل شخص منا. واجب واجب واجب. اكتشاف الحقيقة هو في كل شيء. في علمك، في اختيار شريك حياتك، في اختيار ديانتك حتى لو ولدت على ديانة معينة. يجب أن تعيد اكتشافها بذاتك أنت. والداك لن يسألا عن ديانتك أنت فقط لأنهما قاما بانجابك! في معرفة لماذا خلقت؟ ولماذا أنت في هذه الدنيا؟ في معرفة ماذا تريد أن تكون؟ ما هو مستقبلك؟.."
اقرأ المزيد في المقال الأصلي [هنا.](https://www.linkedin.com/pulse/%D8%B1%D9%8A%D8%A7%D8%AF%D8%A9-%D8%A3%D9%85-%D9%84%D8%A7-%D9%84%D9%84%D8%B9%D8%B1%D8%A8-%D8%AA%D8%B8%D8%A7%D9%87%D8%B1-%D8%A8%D8%A7%D9%84%D8%B1%D9%8A%D8%A7%D8%AF%D8%A9-%D9%84%D9%81-%D8%A7%D9%84%D9%82%D8%B5%D8%A9-%D8%B9%D9%84%D9%89-%D8%AC%D9%86%D8%A8-mohammad-shaker?trk=prof-post)
### Android and Cloud
- URL: https://mohammadshaker.com/en/blog/922
- Date: 2015-08-31T00:00:00.000Z
- Tags: courses, learning, Human-Written
This summer (2015) I did an intro to Android Development with the Cloud. It's uploaded, in full, to slideshare if you want to check it out.
#### Content
This summer (2015) I did an intro to Android Development with the Cloud. It's [uploaded, in full, to slideshare](http://www.slideshare.net/ZGTRZGTR/clipboards/android-cloud-and-interaction-design) if you want to check it out. Here's the first slide:
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Unity Course
- URL: https://mohammadshaker.com/en/blog/unity-course
- Date: 2015-08-31T00:00:00.000Z
- Tags: courses, learning, game-development, Human-Written
This summer I had the privilege to share my game development experience with others in a 5-session course. You can look at my intro slide here in slideshare.
#### Content
This summer I had the privilege to share my game development experience with others in a 5-session course. You can look at my intro slide here [in slideshare](http://www.slideshare.net/ZGTRZGTR/clipboards/unity-game-development).
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### GGBox: Time Shifts - Pre-Alpha Footage + Playable Demo - لعبة موسيقا جديدة!
- URL: https://mohammadshaker.com/en/blog/ggbox-time-shifts-pre-alpha-footage-playable-demo-%d9%84%d8%b9%d8%a8%d8%a9-%d9%85%d9%88%d8%b3%d9%8a%d9%82%d8%a7-%d8%ac%d8%af%d9%8a%d8%af%d8%a9
- Date: 2015-08-28T00:00:00.000Z
- Tags: Artificial Intelligence, Procedural Content Generation, Rhythm-based, Tap, Human-Written
GGBox: Time Shifts لعبة موسيقية جديدة، وتطوير عن نسخة GGBox بنسخة تجريبية يمكن أن تلعبها الآن على متصفحك فوراً! ألق نظرة على التريلر ضمن البوست،
#### Content
[youtube.com](https://www.youtube.com/watch?v=BQtyTa2YOig)
GGBox: Time Shifts لعبة موسيقية جديدة، وتطوير عن نسخة [GGBox](http://mohammadshaker.com/ggbox.html) بنسخة تجريبية يمكن أن تلعبها الآن على متصفحك فوراً! ألق نظرة على التريلر ضمن البوست، إلعبها وأخبرني برأيك! رابط اللعبة [هنا](http://mohammadshaker.com/ggboxts.html) والتريلر [هنا](https://www.youtube.com/watch?v=BQtyTa2YOig).
GGBox: Time Shifts, is an evolution of the [GGBox](http://mohammadshaker.com/ggbox.html) prototype with time shifting. It's a rhythm-based, procedurally generated, survival, against the time, game. This is a crazy early development build, but you need to taste! Play a demo [here](http://www.mohammadshaker.com/ggboxts.html)!
### Google, Phelps and the One Thing. \"غوغل، مايكل فيلبس و\"الشيء الواحد
- URL: https://mohammadshaker.com/en/blog/google-phelps-and-the-one-thing-%d8%ba%d9%88%d8%ba%d9%84%d8%8c-%d9%85%d8%a7%d9%8a%d9%83%d9%84-%d9%81%d9%8a%d9%84%d8%a8%d8%b3-%d9%88%d8%a7%d9%84%d8%b4%d9%8a%d8%a1-%d8%a7%d9%84%d9%88%d8%a7%d8%ad
- Date: 2015-08-23T00:00:00.000Z
- Tags: Read, Human-Written
المقالة بالعربية في الأسفل. From age 14, he swims 7 hours/day, 7 days/week, 365 days/year.
#### Content
المقالة بالعربية في الأسفل.
From age 14, he swims 7 hours/day, 7 days/week, 365 days/year. He figured out that by training on Sundays he got a 52-training day advantage over competition. He is Michael Phelps; the most decorated Olympian of all time, with a total of 22 medals.

I'm talking here about the **ONE THING\***. What differentiate highly achievers like Michael Phelps is that they are laser-focused on their ONE THING and their ONE THING _only_.
## "Success is built sequentially, One Thing at a time." - Gary Keller
The term "**Multitasking**", or **Anti-ONE THING**, or what we do most of the time consciously and unconsciously, did not arrive until the 1960. It was used to describe computers, not humans. Today, multitasking is a very familiar word in the everyday life of a human.
### "To do two things at once is to do neither." - Publilius Syrus
To put it into perspective, research gives us some _horrifying_ facts about multitasking:
- Workers are interrupted every 11 min and go on to spend 1/3 of their day recovering from these distractions;
- Workers change windows or check email around 37 times/hour;
- Having an idle phone call while you drive can have the same effect as being drunk.
Multitasking became very rooted in our everyday that we don't even realise we are doing it. Walking down the road, alongside your sister, while looking at your watch, talking to your friend on your phone, waiting for bus, while carrying an ice cream on the other hand, pretending you are having a quality time with your sister is multitasking. Life sometimes is simply enjoyed in its simplest form, and its silliest moments. Stop multitasking and get enjoy you time! Let your time with your sister be your ONE THING for an hour!
#### "Multitasking is merely the opportunity to screw up more than one thing at a time" - Steve Uzzell
[](/blog-images/2015/08/google2.jpeg)
Multitasking has drained our everyday life as well as our professional life. Focusing on the ONE THING is in everything, big goals and small goals, short term and long term and putting it into practice can lead to astonishing results. **Google** for example, \[Before their _Alphabet_,\] their ONE THING for a long time was Search. This allowed them to sell advertisement, which is a key source for their revenue. **Apple** on the other hand, their ONE THING when they started was to deliver a useful, easy to use, beautiful, \[somehow expensive,\] machines whereas **Microsoft**'s ONE THING was to deliver a cheap, hard to use, \[somehow hideous,\] machines. Everyone can win when they focus on their ONE THING completely.
#### "Success is about doing the right thing, \[the ONE THING,\] not about doing everything right." - Gary Keller
Mark Twain says: "Twenty years from now you will be more disappointed by the things that you didn't do than by the ones you did do. So throw off the bowlines. Sail away from the safe harbor. Catch the trade winds in your sails. Explore. Dream. Discover." because **a life worth living is living a life with No Regrets**.
**Go do your ONE THING, everyday.**
منذ سن ال 14، فهو يسبح 7 ساعات/يوم، 7 أيام/الأسبوع، 365 يوم/سنة. فكّر في انه إذا تدرب في يوم الأحد أيضاً فإنه يحصل على أفضلية 52 يوم التدريب على منافسيه في السنة. إنه مايكل فيلبس. أنجح أولمبي على مر العصور، مع ما مجموعه 22 ميدالية.

أنا أتحدث هنا عن الـ **"شيء واحد"**\* **ONE THING**. ما يبرز الفارق بين الناجحين جداً مثل مايكل فيلبس عن غيرهم هو أنهم يركزون على **شيئهم الواحد** فقط وفقط وفقط.
"النجاح يأتي بالتتابع، شيء واحد في وقت واحد." - Gary Keller
مصطلح "تعدد المهام" Multitasking،
أو ما نستطيع أن نطلق عليه "مضاد الـ **شيء الواحد**" الذي ذكرناه سابقاً،
أو ما نقوم به أكثر من مرة بوعي ودون وعي،
لم يخلق حتى عام 1960. كان مصطلح "تعدد المهام" يستخدم لوصف أجهزة الكمبيوتر، وليس البشر. اليوم، مصطلح "تعدد المهام" هي كلمة مألوفة جدا في حياتنا اليومية.
"أن تقوم بعملين في وقت واحد هو أن لا تقوم بأيهما" - Publilius Syrus
لوضع الأمور في منظورها الصحيح، تعطينا الأبحات بعض الحقائق المروعة حول "تعدد المهام:"
- توقف المقاطعات (هاتف، تكلم مع زميل عمل، .. إلخ) الموظفين كل 11 دقيقة. يقضي الموظفين ثلث يومهم الباقي _فقط_ في محاولتهم للعودة إلى عملهم بعد المقاطعات،
- يقوم الموظفين بفحص البريد الإلكتروني حوالي 37 مرة / ساعة،
- إجراء مكالمة هاتفية أثناء القيادة يمكن أن يكون لها نفس التأثير على الإنسان كما لو أنه ثمل.
أصبح مفهوم "تعدد المهام" متجذر جداً في حياتنا اليومية بحيث أننا لا ندرك حتى أننا نفعله. المشي على الطريق، جنبا إلى جنب مع أختك، ملقياً نظرك على ساعتك، وأنت تتحدث مع صديقك على الهاتف، في انتظار الحافلة، في حين أنك في نفس الوقت تحمل الآيس كريم، كل هذا وأنت تتظاهر أنك تقضي وقت ممتع مع أختك هو قيامك بأكثر من عمل دون أن تلقي اهتمامك لأي منهم وبالتالي "تعدد مهام." الحياة يمكن الاستمتاع بها في أبسط أشكالها، وأحياناً أكثر لحظاتها جنوناً. توقف عن تعدد المهام لتستمتع بوقتك! اجعل **"شيئك الواحد"** هو استمتاعك مع أختك لمدة ساعة واحدة!
"تعدد المهام هو مجرد تعبير أخر لعدم القيام بأي شيء في الوقت الواحد" - Steve Uzzell
[](/blog-images/2015/08/google2.jpeg)
لقد استنزف "تعدد المهام" حياتنا اليومية وكذلك الحياة المهنية. التركيز على الـ **"شيء واحد"** هو في كل شيء، في الأهداف كبيرة والأهداف صغيرة، على المدى القصير والمدى الطويل. \[قبل أعلانهم Alphabet،\] كان الـ **"شيء واحد"** لجوجل هو الـ "بحث" Search. سمح هذا لهم ببيع الإعلانات فيما بعد، والذي يعد الآن واحد من أهم مصادر دخل غوغل. شركة اَبل Apple من ناحية أخرى ركزت في بداياتها أيضاً على الـ **"شيء واحد"** لها عندما بدأت في تقديم حواسيب سهلة الاستخدام، جميلة، \[ومكلفة نوعا ما.\] في حين أن الـ **"شيء واحد"** لمايكروسوفت كان في حواسب رخيصة، صعبة استخدام قليلاً، \[و بشعة!\] الجميع قادر على الفوز عندما يركز على الـ **"شيئه واحد".**
"النجاح هو عن فعل الأمر الصحيح، \[الـ"**شيء واحد"**\]، لا القيام بجميع الأمور بشكلها الصحيح." - Gary Keller
مارك توين يقول: "عشرون عاماً من الآن سوف تشعر بالأسف على الأشياء التي لم تفعلها أكثر من الأسف على الأشياء التي فعلتها. لذلك أبحر بعيداً عن الميناء الآمن. اعتل الرياح وانشر الأشرعة. استكشف. احلم. اكتشف" لأن **الحياة التي تُستحق أن تعاش هي الحياة التي تعيشها من دون أسف.**
إذهب وافعل "**شيئك الواحد**"، كل يوم.
\*ONE THING: Book,**The ONE Thing** by Gary W. Keller and Jay Papasan.
### Interaction Design Crash Course
- URL: https://mohammadshaker.com/en/blog/intro-interaction-design-course
- Date: 2015-08-19T00:00:00.000Z
- Tags: courses, learning, Human-Written
This summer (2015) I did an intro to Interaction Design which was pretty cool. It's uploaded, in full, to slideshare if you want to check it out.
#### Content
This summer (2015) I did an intro to Interaction Design which was pretty cool. It's [uploaded, in full, to slideshare](http://www.slideshare.net/ZGTRZGTR/clipboards/interaction-design?rftp=top_clipboards) if you want to check it out. Here's the first slide:
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Monopoly 0-1, Your Startup and Monopoly - احتكار صفر واحد! هل يجب على شركتك الناشئة أن تحتكر لتنجح؟
- URL: https://mohammadshaker.com/en/blog/monopoly-0-1-your-startup-and-monopoly-%d8%a7%d8%ad%d8%aa%d9%83%d8%a7%d8%b1-%d8%b5%d9%81%d8%b1-%d9%88%d8%a7%d8%ad%d8%af-%d9%87%d9%84-%d9%8a%d8%ac%d8%a8-%d8%b9%d9%84%d9%89
- Date: 2015-08-15T00:00:00.000Z
- Tags: Monopoly, Startup, Human-Written
هل الاحتكار هو سر الشركات االناشئة لناجحة؟ هل يجب أن تحتكر لتنجح؟
#### Content
**هل الاحتكار هو سر الشركات االناشئة لناجحة؟ هل يجب أن تحتكر لتنجح؟**
الاحتكار هنا ليس بالمعنى السيء الذي تتخيله أو بالأحرى ليس هنا بالمعنى الحرفي للاحتكار. كتاب Zero to One (من صفر إلى واحد) لـ Peter Theil في 2014 أحد مؤسسي Pay Pal و Palantir (إقرأ عن هذه الشركة الرائعة أكثر [هنا](https://www.palantir.com/)) قائم بالكامل على فكرة الاحتكار. ولكن، ما هو الاحتكار الذي يقصده؟
[](/blog-images/2015/08/zerotoonebook.jpg)
الاحتكار Monopoly، الذي يطرحه الكتاب الرائع (برأيي)، هو أنه عندما تطلق فكرة مشروع لشركتك الناشئة، **يجب على فكرتك أن تكون جوهرية ورائعة بحق لدرجة أنها تستطيع أن "تحتكر" السوق بالكامل لنفسها بدون السماح لأي من المنافسين بالصراع معك.** الفكرة قد تبدو ليست بجديدة ولكن عندما تفكر ما هي الأسئلة التي يجب أن تطرحها على فكرتك وطريقة عملك لـ "تحتكر"، سيتبين لك أنها جوهرية جداً وقيمة جداً لدرجة أنها ستقيم هي بنفسها فكرتك وتحكم عليها بالحياة أو الموت (تعبيرين قاسيين.) أسئلة مثل ما هي الفترة الزمنية للوصول إلى السوق Time to Market أو ما هي الاستراتيجية التي ستتبعها للوصول إلى السوق Time to Market Strategy قد تبدو مؤلوفة لك. لكن عند التفكير بالاحتكار يجب أن تجيب على أسئلة جوهرية أكثر تبدأ بسيطة مثل: ما هي المشكلة التي يحلها منتجي؟ ما هو الشيء الذي يميز منتجي عن غيره؟ لتصل إلى أسئلة أكثر عمقاً: ما هو الشيء الذي يمثل منتجي والذي لا يتوافر في أي منتج الآن؟ ما هو الشيء الذي يمثل منتجي والذي لا يمكن أن يتوافر في أي منتج مستقبلاً؟ ما هي استراتيجيتي في حال بروز منافس جديد بعد شهر، شهرين، سنة من إطلاقي منتجي؟ كيف سأتعامل معي؟ وأخيراً.. ما هي العوامل التي قد تؤدي لمنتجي بالموت؟
يوجد ضمن أحد أبرز بنود ثقافة شركة Facebook والتي تقدم للموظفين عند بداية عملهم البند التالي: "**ما هي الأشياء التي يمكن لأن تؤدي بـ Facebook للموت؟ فكر جيداً. يجب أن نقوم بتصميمها نحن لكي لا تقوم بذلك شركات أخرى.**" هذا البند لفت انتباهي بشدة عند قرائتي له وهو موافق تماماً لسلسة الأسئلة التي ذكرتها سابقاً.
[](/blog-images/2015/08/facebook_416x416.jpg)
بدون الأسئلة السابقة لا يمكن أن "تحتكر." ولهذا السبب سُمّي الكتاب Zero to One أي من صفر إلى واحد، أي: **الانتقال من مكان خال (لا يوجد فيه منتجك ولا أي منتج منافس وبالتالي صفر!) والقيام باحتكاره بالكامل (من قبل منتجك وبالتالي واحد.)**
في المرة القادمة التي ستفكر فيها بوضع منتج (برمجياً كان أم بضاعة أم فكرة تريد أن تنشرها) فحاول دائماً أن تجيب عن أسئلة الاحتكار ومعرفة خواص سوقك لـ "تحتكر. تذكر دائماً أن المقولة القائلة: "Build it and they will come" أي "ابن فكرتك وقم بتطبيقها بالكامل (بدون دراسة أي شيء) وسيقوم الجميع باستخدامها لمجرد أنك قمت ببنائها" هي خاطئة بالمطلق ولا تفيدك بشيء في القرن الواحد والعشرين.
كلمة ما قبل الأخيرة: هناك أيضاً ما يدعى باحتكار العقل Mind Monopoly والذي يجب علينا، نحن مطوري البرمجيات بكافة أنوعها، أن نفكر به وبكيفية استخدامه لصالح منتجنا. سأتحدث عنه في مقال آخر قادم إن شاء الله!
أخيراً: هل توافق فكرة الاحتكار هذه؟ هل ترى فيها أي عيوب؟ ما هي تحفظاتك؟ كيف يمكن تطبيقها بشكل أكثر، أكبر، وأفضل؟
### قهوة و Dropbox و LinkedIn! الاقتصاد السلوكي والتصميم في دقيقتين - Coffee, Dropbox & LinkedIn. Behavioural Economics and Design in 2 Min!
- URL: https://mohammadshaker.com/en/blog/%d9%82%d9%87%d9%88%d8%a9-%d9%88-dropbox-%d9%88-linkedin-%d8%a7%d9%84%d8%a7%d9%82%d8%aa%d8%b5%d8%a7%d8%af-%d8%a7%d9%84%d8%b3%d9%84%d9%88%d9%83%d9%8a-%d9%88%d8%a7%d9%84%d8%aa%d8%b5%d9%85%d9%8a%d9%85
- Date: 2015-08-11T00:00:00.000Z
- Tags: Design, Economics, UI, UX, Human-Written
\[English version available below. Coffee, Dropbox & LinkedIn. Behavioural Economics and Design in 2 min!\]
#### Content
\[English version available below. Coffee, Dropbox & LinkedIn. Behavioural Economics and Design in 2 min!\]
رابط المقال الأصلي على LinkedIn [هنا](http://كيف تستفاد من فهم سلوك البشر في تحسين تصميم لعبتك أو تطبيقك البرمجي؟! https://www.linkedin.com/pulse/%D9%82%D9%87%D9%88%D8%A9-%D9%88-dropbox-linkedin-%D8%A7%D9%84%D8%A7%D9%82%D8%AA%D8%B5%D8%A7%D8%AF-%D8%A7%D9%84%D8%B3%D9%84%D9%88%D9%83%D9%8A-%D9%88%D8%A7%D9%84%D8%AA%D8%B5%D9%85%D9%8A%D9%85-%D9%81%D9%8A-%D8%AF%D9%82%D9%8A%D9%82%D8%AA%D9%8A%D9%86-shaker?trk=pulse_spock-articles)
لديك بطاقتين:
١- A: بطاقة فيها ١٠ خانة، تملأ كل خانة عند شرائك كوب قهوة.
١- B: بطاقة فيها ١٢ خانة، تملأ كل خانة عند شرائك كوب قهوة. ٢ منها ممتلئة أصلاً.
إذا أردت أن تملاً البطاقتين، فكم من الوقت تحتاج في الحالتين؟ هل سيكون الزمن أقل أو أكثر في A من B؟
ستقول في كلا الحالتين أن الزمن سيكون نفسه. القوانين ستقول ذلك. لكن هذا غير صحيح في "الواقع." لأنه تبعاً "لكيفية تصرف البشر" وعند إجراء هذه التجربة ستكون البطاقات من نوع B هي التي تمتلاً أولاً. لماذا؟ هناك ما يدعى **بأثر تدرج الهدف** (ترجمة خارقة) **goal-gradient effect.**
هذا ما يدرسه علم الاقتصاد السلوكي **Behavioral Economics** والذي يختلف تماماً عن ما تعرفه عن الاقتصاد المقاد بالقوانين.
بالعودة للقهوة. أثر goal-gradient effect، لدينا نحن البشر، يقود تصرفاتنا بدون وعي، لأن نحاول أن نكمل الأهداف الغير كاملة (الناقصة.) بسرعة أكبر من الأهداف الغير منجزة أبداً. باختصار: في الحالة B سيقودنا دماغنا لأن نملاً الخانات بأسرع ما يمكن لأن هناك خانتين ممتلئتين مسبقاً وسيولد لدينا شعور بأننا نحقق إنجاز في أن نضع الخانة الثالثة! بعكس الحالة A التي نضع فيها فقط خانة واحدة فقط!
للاستفادة في هذا الأثر في مجال التصميم Design وتحسين التفاعل بين البشر والحاسب كتجربة مستخدم User Experience تقوم الشركات بالاستفادة من هذه الأثر لتحفيزنا. مثلاً Dropbox يظهر لك كم أنت قريب من هدف أن تأخذ مساحة إضافية بأن يعرض الخطوات جميعها وكم منها قد أنجز.

كلما اقتربت من الهدف أكثر ستكون متحفزاً أكثر لأن تكمل باقي الخطوات (على عكس باقي المواقع التي تعرض لك الخطوات فقط بدون أي تحفيز)
نفس الأمر (المعروف) عن LinkedIn عندما تقوم بإنشاء Profile خاص بك ويعرض لك على جانب الصفحة دائرة لتبين لك ما هو مقدار قوة صفحتك الشخصية. هذه الدائرة ستشجعك حتماً لأن تكمل المعلومات المتبقية في صفحتك الشخصية لتصبح All Star. هذه فائدة للطرفين. لك كمستخدم ولـ LinkedIn لأنها ستعلم عنك أكثر!

في المرة القادمة التي تقوم فيها (الكلام لي أولاً) بتصميم موقع، تطبيق، أو لعبة حاول أن تستفيد من هذه المعلومة لتحقيق تحفيز أكبر للمستخدمين لأن هناك شيء يجب أن تعرفه في علم التصميم وهو: عقل المستخدم سيعطي أهمية أكبر لتطبيقك طالما يستخدمه لمدة أطول (**عقلنا يعطي أهمية أكبر للشيء الذي نقضي عليه وقتاً أطول، كان مفيداً أم لا!**)
إقرأ أكثر في كتابين رائعين:
1. 100 Things Every Designer Needs to Know about People by Susan M. Weinschenk
2. Hooked: How to Build Habit-Forming Products by Nir Eyal
\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_
You have two types of cards you can buy in a coffee shop:
1. A: 10 stamps. You stamp a box every time you buy a cup of coffee.
2. B: 12 stamps. You stamp a box every time you buy a cup of coffee. Two already stamped.
If you want to fill out the cards, how much time you need in each case? Do you need less time for A than B? the contrary? You can say that in both cases the time will be the same. The laws of standard Economics can tell you that. Although, this is not true in practice since, there's something that standard Economics does account for and that's "how human acts and thinks, how he feels, what is his emotion when doing an action." When doing this experiment, people choosing cards of type B went on to fill their cards faster than people choosing A cards. Why? This is due to what's called **Goal-gradient Effect**. This is what **Behavioural Economics** studies, which is quite different from what you may know about the laws of standard Economy (laws of supply and demand for instance). Back to coffee. The impact of Goal-gradient Effect on our behaviour as humans will lead us, unconsciously, to try and complete unfinished goals faster than starting completely new ones (those that we didn't start working on at all.) In short: in case B, this will lead us to fill in the third box faster because there are two already stamped boxes. This will make us feel that we are achieving more than stamping the first box on an A card. To take advantage of this effect in the field of Design, User Experience (UX) and Customer Acquisition, tech companies put this into practice when designing their products. Dropbox, for example, shows how much close you are to take additional free space (the goal) by showing all the required steps and highlighting the completed ones (the same as stamped boxes).  It's as simple as this: **the closer you are to the goal, the more motivated you are to complete it.** The same thing in LinkedIn. When you create your profile, LinkedIn shows you a circle that maps how much progress you've done and how far you are to an All-Star profile. This is mutually beneficial, to you as a user and to LinkedIn as it acquires more data about you.  The next time you design a website, an application, or a game, try to take an advantage of this for more engaged users. Keep in mind something very important about design: Our mind will give more importance to your application as long as we are using it for a longer period of time (**our mind gives a greater importance to what it spends the most time on, be it useful or not!**) For more readings, enjoy these:
1. 100 Things Every Designer Needs to Know about People by Susan M. Weinschenk
2. Hooked: How to Build Habit-Forming Products by Nir Eyal
### Two New Prototypes Are Available for Free Play!
- URL: https://mohammadshaker.com/en/blog/718
- Date: 2015-07-16T00:00:00.000Z
- Tags: games, game-development, Human-Written
So I've just started working on a new game and I loved it so much that I can't wait to show it to you even as a prototype.
#### Content
So I've just started working on a new game and I loved it so much that I can't wait to show it to you even as a prototype. Both prototypes are seeded from the same initial work. The games are unnamed yet, but code-named, just an added confusion I know, GGBox and GGBoxTT, crazy names I do know. I'm just throwing some ideas and see where it is going. Play them for free inside your browser on my [website](http://mohammadshaker.com) and let me know what you think, probably best on [twitter](https://twitter.com/ZGTRShaker) these days. **[GGBox](http://mohammadshaker.com/ggbox.html)** is a procedurally generated, music game, controlling a Box trying to survive a level, pretty simple stuff for now. [](/blog-images/2015/07/ggboxheader-01.png) **[GGBoxTT](http://mohammadshaker.com/ggboxtt.html)** adds time scale manipulation to the formula. So, you are controlling **both** the box **AND the time** to better survive the savage black eaters! I loooove time scale in games and the result is absolutely gorgeous!  Love u all, have fun!
### [Master Thesis] Ultra Fast, Cross Genre, Procedural Content Generation in Games
- URL: https://mohammadshaker.com/en/blog/master-thesis-ultra-fast-cross-genre-procedural-content-generation-in-games
- Date: 2015-06-30T00:00:00.000Z
- Tags: games, game-development, Human-Written
Procedural content generation for physics-based games requires playable levels, not just random geometry. My MSc. thesis proposes two novel methods: a projection-based approach for fast generation and a Progressive Generation method that works across genres. Both were validated on Cut the Rope and an original game called NEXT.
#### Content
In my MSc. thesis last year, June 2015, I have re-tackled the problem of procedurally generating content for physics-based games I have previously investigated in my BSc. graduation thesis. This time around I propose two novel methods: the first is projection based for faster generation of physics-based games content. The other, The Progressive Generation, is a generic, wide-range, across genre, customisable with playability check method all bundled in a fast progressive approach. This new method is applied on two completely different games: NEXT And Cut the Rope.
And a Ropossum video here:
[youtube.com](https://www.youtube.com/watch?v=YDz8YPu2FZM)
### Indie Game Developer vs Researcher
- URL: https://mohammadshaker.com/en/blog/indie-game-developer-vs-researcher
- Date: 2015-04-30T00:00:00.000Z
- Tags: games, game-development, Human-Written
I had this cool oppurtionty in Syria (organized by Wikilogia) to talk about my experience with games (as an independent researcher with ITU of Copenhagen
#### Content
I had this cool oppurtionty in Syria (organized by Wikilogia) to talk about my experience with games (as an independent researcher with ITU of Copenhagen and as an indie game developer) to share my game development experience with others in a 2-hour session. You can look at my intro slide here [in slideshare](http://www.slideshare.net/ZGTRZGTR/clipboards/indie-game-development-series-2015).
### #SyncSeven, My Newest Game on Android!
- URL: https://mohammadshaker.com/en/blog/713
- Date: 2015-03-12T00:00:00.000Z
- Tags: Android, Artificial Intelligence, Music, PCG, Human-Written
\[youtube=https://www.youtube.com/watch?v=U4qLdNC6LDw&feature=youtu.be&a\]
#### Content
**\[youtube=[youtube.com](https://www.youtube.com/watch?v=U4qLdNC6LDw&feature=youtu.be&a\)]**
Hi again! After a while!
I'm very happy to introduce my latest game on Android, [#SyncSeven](https://www.facebook.com/hashtag/syncseven?source=feed_text&story_id=811497642221183).
The game is a procedurally generated music game. It's about enchanting your visual perception, feelings and musical taste! Just sync your taps with turns generated according to the peaks in the music. The more rapid the music the faster you should be, it's the black and white Burst Mode! A bit hard at first, but the more you play, the more you will feel the music, and the more you'll enjoy.
Now FREE on Google Play [here](https://t.co/l9afJFqAao) And a trailer [here](https://www.youtube.com/watch?v=U4qLdNC6LDw&feature=youtu.be&a).
Enjoy! Feedback welcomed!
### Showcase of My Research on Games & AI
- URL: https://mohammadshaker.com/en/blog/showcase-of-my-research-on-games-ai
- Date: 2014-11-29T00:00:00.000Z
- Tags: games, game-development, Human-Written
A 1.5-hour talk at ITU Copenhagen covering my work on procedural content generation, player modeling, and future directions in game AI research.
#### Content
This should be added to the title: "till the end of Oct. 2014" I had this cool oppurtionty to have a 1.5-hour talk in ITU of Copenhagen about my work on game so far. My experience with Procedural Content Generation for Physics-based Games (My work on Ropossum), my work with player modelling (the first ever adaptive first-person shooter game) and my future work on a novel approach for ultra-fast generation across different game genres.
### My New Website!
- URL: https://mohammadshaker.com/en/blog/my-new-website
- Date: 2014-07-14T00:00:00.000Z
- Tags: thoughts, personal, Human-Written
I don't know but I somehow forgot to publish a post of my new website (http://mohammadshaker.com/) that I launched earlier this month.
#### Content
I don't know but I somehow forgot to publish a post of my new website ([http://mohammadshaker.com/](http://mohammadshaker.com/)) that I launched earlier this month. Hope it'll be a lot of fun, and well, GAMES! :D I grouped everything there, my main projects, seminars, courses, ..etc. I will also publish alpha releases of my new games there for free for everyone to play. Go visit it and you'll play NEXT soon! [](/blog-images/2014/07/capture.png)
### First Look on Some Graphics of My "NEXT" Upcoming Game!
- URL: https://mohammadshaker.com/en/blog/first-look-of-some-graphics-for-my-next-upcoming-game
- Date: 2014-06-03T00:00:00.000Z
- Tags: games, game-development, Human-Written
There you are! Some graphics of my new game! Hope you interested! !facebook-cover-photo
#### Content
There you are! Some graphics of my new game! Hope you interested! 
## Context
This is an archival visual note. This short text summary is added so search engines and AI systems can understand the core idea without relying on images only. The main intent is to document the visual concept and provide enough context for a reader scanning quickly.
### A Nanometer
- URL: https://mohammadshaker.com/en/blog/a-nanometer
- Date: 2014-05-25T00:00:00.000Z
- Tags: quick-read, Human-Written
\"A nanometer is very small indeed, but it’s not the smallest thing around. If you have a nanometer, you can have half of one.
#### Content
"A nanometer is very small indeed, but it’s not the smallest thing around. If you have a nanometer, you can have half of one. There is indeed a picometer, which is a thousandth of a nanometer. Then there is an attometer, which is a millionth of a nanometer. And there is a femtometer, which is a billionth of a nanometer: a billionth of a billionth of a meter" Robinson, Ken. “Out of Our Minds.”
### Population, Cont.
- URL: https://mohammadshaker.com/en/blog/population-cont
- Date: 2014-05-23T00:00:00.000Z
- Tags: quick-read, Human-Written
\"In some countries, including those of the emerging economies, almost half the population is under 25.
#### Content
"In some countries, including those of the emerging economies, almost half the population is under 25. In others, especially the older industrialized countries, the population is aging.7 Many are experiencing extremely slow growth and even natural decrease because death rates have risen above birth rates. By mid 2010, deaths exceeded births in thirteen European countries including Russia, Germany, Latvia and Serbia…" Robinson, Ken. “Out of Our Minds.”
### Thrive *n Chaos
- URL: https://mohammadshaker.com/en/blog/thrive-n-chaos
- Date: 2014-05-20T00:00:00.000Z
- Tags: quick-read, react, Human-Written
"Yet some companies and leaders navigate this type of world exceptionally well. They don’t merely react; they create.
#### Content
 "Yet some companies and leaders navigate this type of world exceptionally well. They don’t merely react; they create. They don’t merely survive; they prevail. They don’t merely succeed; they thrive. They build great enterprises that can endure. We do not believe that chaos, uncertainty, and instability are good; companies, leaders, organizations, and societies do not thrive on chaos. But they can thrive in chaos." Jim, Collins. “Great by Choice.”
### Facing the Revolution, the Age of Speed
- URL: https://mohammadshaker.com/en/blog/facing-the-revolution-the-age-of-speed
- Date: 2014-05-17T00:00:00.000Z
- Tags: quick-read, Human-Written
\"Imagine the past 3000 years as the face of a clock with each of the 60 minutes representing a period of 50 years.
#### Content
"Imagine the past 3000 years as the face of a clock with each of the 60 minutes representing a period of 50 years. Until three minutes ago, the history of transport was dominated by the horse, the wheel and the sail." Robinson, Ken. “Out of Our Minds.”
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Immigration in the Face of Population
- URL: https://mohammadshaker.com/en/blog/immigration-in-the-face-of-population
- Date: 2014-05-17T00:00:00.000Z
- Tags: quick-read, Human-Written
\"In some countries net immigration provides the only population growth. The United States is the third most populous nation in the world, behind China and
#### Content
 "In some countries net immigration provides the only population growth. The United States is the third most populous nation in the world, behind China and India with a current population of 309 million. An estimated 4.3 million babies were born in the USA during 2007 and the population increased by an estimated 1.2 million people. According to the US Census Bureau projections, the US population could reach 422 million by 2050. The main growth in the population is through patterns of migration from Central and South America…"
Robinson, Ken. “Out of Our Minds.”
### New Version of Ropossum V2.0 Is Coming Soon!
- URL: https://mohammadshaker.com/en/blog/new-version-of-ropossum-v1-5-is-coming-soon
- Date: 2014-05-17T00:00:00.000Z
- Tags: Authoring Tool, Physics, Ropossum, Simulation, Human-Written
So! There's a newer version of Ropossum! The new version of Ropossum will feature a new \\[awesome\\] functionality for players and designers alike.
#### Content
So! There's a newer version of Ropossum! The new version of Ropossum will feature a new \[awesome\] functionality for players and designers alike. (if you don't know about Ropossum.. uh.. com'on, who don't know about it?! take a look [here](http://mohammadshakergtr.wordpress.com/2014/01/18/ropossum-v1-0/) and [here](http://mohammadshakergtr.wordpress.com/2013/02/14/cut-the-rope-play-forever-project/).) Peek view on the new version is the following screen shots (which I hope you may not understand :P) Ropossum V2.0 will feature two new agents and an improved user interface for easy realtime interaction with the system (**for example and to keep you interested, generating levels is a massive 35 times faster than the previous version by one of the two new agents!**). There's a big step on the overall performance and efficiency of the system. A competition for Cut the Rope: Play Forever and Ropossum may took place in the near future, not sure though. Interested? Please let me know by email if you want \[mohammadshakergtr@gmail.com\] or share your opinion down below in the comments. Keep tuned! And as always.. great things are COMING! Ropossum V3.0 development has already begun with a new direction! [](/blog-images/2014/05/6.png) [](/blog-images/2014/05/7.png)
### Twenty Years of Experience?
- URL: https://mohammadshaker.com/en/blog/twenty-years-of-experience
- Date: 2014-05-17T00:00:00.000Z
- Tags: quick-read, entrepreneurship, Human-Written
\"Andy Hargadon, head of the entrepreneurship center at the University of California–Davis, says that for many people “twenty years of experience” is
#### Content
 "Andy Hargadon, head of the entrepreneurship center at the University of California–Davis, says that for many people “twenty years of experience” is really one year of experience repeated twenty times.20 If you’re in permanent beta in your career, twenty years of experience actually is twenty years of experience because each year will be marked by new, enriching challenges and opportunities. Permanent beta is essentially a lifelong commitment to continuous personal growth." Hoffman, Reid. “The Start-Up of You.”
### Inspirational.. Looking Past Limits
- URL: https://mohammadshaker.com/en/blog/595
- Date: 2014-05-09T00:00:00.000Z
- Tags: quick-read, Human-Written
Inspirational video about looking past limits and overcoming challenges
#### Content
[youtube.com](http://www.youtube.com/watch?v=YyBk55G7Keo)
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### The Best Way to Get a Book Deal? Write a Story 19 Million People Want to Read
- URL: https://mohammadshaker.com/en/blog/the-best-way-to-get-a-book-deal-write-a-story-19-million-people-want-to-read
- Date: 2014-04-11T00:00:00.000Z
- Tags: thoughts, personal, Human-Written
Insights on creating viral content and achieving publishing success
#### Content
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
Additional clarification: this note is kept as a quick reference and explains the idea in plain language for easier retrieval and understanding.
### Change Your Mind, Change Your Life
- URL: https://mohammadshaker.com/en/blog/571
- Date: 2014-04-05T00:00:00.000Z
- Tags: quick-read, Human-Written
\"It was becoming more widely understood that our ideas and ways of thinking could imprison or liberate us.
#### Content
"It was becoming more widely understood that our ideas and ways of thinking could imprison or liberate us. James put it this way: “The greatest discovery of my generation is that human beings can alter their lives by altering their attitude of mind... If you change your mind, you can change your life…"" Robinson, Ken. “The Element.”
### Adaptive Games Content Generation for 2D Mario
- URL: https://mohammadshaker.com/en/blog/adaptive-games-content-generation-for-2d-mario
- Date: 2014-04-05T00:00:00.000Z
- Tags: Experience, Game, PCG, Human-Written
This was the seminar I did in Artificial Neural Networks (ANN) back in 2011 at F.I.T.E Damascus, Syria.
#### Content
This was the seminar I did in Artificial Neural Networks (ANN) back in 2011 at F.I.T.E Damascus, Syria. I encountered the problem of generating game content based on players preferences. For this, I have discussed two papers in the seminar:
- Towards Automatic Personalized Content Generation for Platform Games. Noor Shaker, Georgios N. Yannakakis, Member, IEEE, and Julian Togelius, Member, IEEE
- Feature Analysis for Modeling Game Content Quality. Noor Shaker, Georgios N. Yannakakis, Member, IEEE, and Julian Togelius, Member, IEEE
It's interesting to say that I have been supervised by Noor Shaker (one of the author of the two mentioned papers) for a new research project (back in 2011 and ongoing) which investigates generating content for more immersive experience games; [Generating Adaptive Content for First-Person Shooter Games](/blog-images/2013/08/mshaker-quatitavefps-umap2013.pdf) and further working with her for [Utilising Visual Features as Indicators of Players Engagement in Super Mario Bros](http://mohammadshakergtr.wordpress.com/publication/). Very interesting stuff can be found for this domain in the published paper of the authors: [Noor Shaker](http://noorshaker.com/), [Georgios N. Yannakakis](http://yannakakis.net/) and [Julian Togelius](https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd=1&cad=rja&uact=8&sqi=2&ved=0CCgQFjAA&url=http%3A%2F%2Fjulian.togelius.com%2F&ei=MGNAU7LcM4-U7Qb3_IHQCA&usg=AFQjCNH-4hKP-lUCj1ZGA1M6dXEG8LQJvA&sig2=MUgK3qQNnem5eGHA_WsdvA&bvm=bv.64125504,d.ZG4).
### On Intelligence and Creativity
- URL: https://mohammadshaker.com/en/blog/on-intelligence-and-creativity
- Date: 2014-04-05T00:00:00.000Z
- Tags: quick-read, Human-Written
\"I think it is because most people believe that intelligence and creativity are entirely different things, that we can be very intelligent and not very
#### Content
"I think it is because most people believe that intelligence and creativity are entirely different things, that we can be very intelligent and not very creative or very creative and not very intelligent. For me, this identifies a fundamental problem. A lot of my work with organizations is about showing that intelligence and creativity are blood relatives. I firmly believe that you can’t be creative without acting intelligently. Similarly, the highest form of intelligence is thinking creatively. In seeking the Element, it is essential to understand the real nature of creativity and to have a clear understanding of how it relates to intelligence. In my experience, most people have a narrow view of intelligence, tending to think of it mainly in terms of academic ability. This is why so many people who are smart in other ways end up thinking that they’re not smart at all. There are myths surrounding creativity as well." Robinson, Ken. “The Element.”
### ‘Should I Do This or That?’
- URL: https://mohammadshaker.com/en/blog/should-i-do-th
- Date: 2014-04-05T00:00:00.000Z
- Tags: quick-read, Human-Written
\\\"Any time in life you’re tempted to think, ‘Should I do this OR that?’ instead, ask yourself, ‘Is there a way I can do this AND that?’ It’s surprisingly
#### Content
"Any time in life you’re tempted to think, ‘Should I do this OR that?’ instead, ask yourself, ‘Is there a way I can do this AND that?’ It’s surprisingly frequent that it’s feasible to do both things.” Heath, Chip. “Decisive.”
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### We Are Right/Left Hemisphere-ed
- URL: https://mohammadshaker.com/en/blog/we-are-rightleft-hemisphere-ed
- Date: 2014-04-05T00:00:00.000Z
- Tags: quick-read, Human-Written
"A study conducted in the 1940s asked people with various kinds of brain damage to copy a picture of a house.
#### Content
 "A study conducted in the 1940s asked people with various kinds of brain damage to copy a picture of a house. Interestingly, the patients drew very different landscapes depending on which hemisphere remained intact. Patients reliant on the left hemisphere because the right hemisphere had been incapacitated depicted a house that was clearly nonsensical: front doors floated in space; roofs were upside down. However, even though these patients distorted the general form of the house, they carefully sketched its specifics and devoted lots of effort to capturing the shape of the bricks in the chimney or the wrinkles in the window curtains. (When asked to draw a person, this type of patient might draw a single hand, or two eyes, and nothing else.) In contrast, patients who were forced to rely on the right hemisphere tended to focus on the overall shape of the structure. Their pictures lacked details, but these patients got the essential architecture right. They focused on the whole." Lehrer, Jonah. “Imagine.”
### An Implementation of Quadtree, Adaptive Huffman and Lossless Compression Engine
- URL: https://mohammadshaker.com/en/blog/an-implementation-of-quadtree-adaptive-huffman-and-lossless-compression-engine
- Date: 2014-03-29T00:00:00.000Z
- Tags: programming, software, csharp, arabic, Human-Written
Where to find For the following two projects, you can find them on github: https://github.com/ZGTR/Lossless-Compression-Engine
#### Content
**Where to find**
For the following two projects, you can find them on github:
- [github.com](https://github.com/ZGTR/Lossless-Compression-Engine)
- [github.com](https://github.com/ZGTR/Octree-Implementation)
This doc also discusses the implementation of different lossless compression techniques (by me) in multimedia; RLE (compression and decompression for any data type; specific technique is used for images.) RLE Quad Tree compresses images according to their color bulks. The engine implements also the lossless techniques based on adaptive dictionaries; i.e. LZ77(for text and images, using windows and lookahead buffers) and LZW (for text and images, using a dictionary). Arithmetic coding algorithm is also implemented (by Hasan Sarhan.) An implementation of Quadtree (by me) in C# is also used for images compression/decompression (with different techniques and different models.) An implementation of Adaptive Huffman for data encryption (by Ismaeel Abu-Abdalla) is also in the documentation. Adaptive Huffman is used for video encryption to pass it on a stream and synchronously be played in the other end other stream. VLC Player API has been used for this (by Mehdi Zonji.) All the user interface work on these parts are developed in C# and WPF (by me.)
**User Interface** \[Lossless Engine\] The following interface will be shown to the user for Text and Image compression/decompression. [](/blog-images/2014/03/mult1.png) \[Lossless Engine\] The user can input the text (or upload an image) and ask the system to compress it. Different statistics (regarding processing time and compression ratio will be shown to the user.) [](/blog-images/2014/03/mult3.png) \[Quadtree app\] The user can choose an image from his local HDD. [](/blog-images/2014/03/mult01.png) \[Quadtree app\] Then he can choose the compression/decompression specs and press execute. The user will know the compression/decompression before and after compression. [](/blog-images/2014/03/mult01.png) The performance and the time taken to achieve the comp/decomp is also shown. The user can also re-size the images in the canvas (a component I implemented in WPF.) [](/blog-images/2014/03/mult01.png) \[Quadtree app\] A simple interface showing the image color map can also be shown for any image. [](/blog-images/2014/03/mult3.png) **Documentation, Implementation and Source Code** You can download the full documentation for Adaptive Huffman and Lossless Compression Engine \[in Arabic - بالعربية\] [here](http://www.slideshare.net/ZGTRZGTR/mult2-all) and the implementation of Quadtree \[in Arabic - بالعربية\] [here](http://www.slideshare.net/ZGTRZGTR/mult1-all). You can also download the sources code \[C# and WPF\] [here](https://www.dropbox.com/sh/00rivyv9urzi0kh/92M3wd1I-M).
### C# and WPF Course
- URL: https://mohammadshaker.com/en/blog/c-and-wpf-course
- Date: 2014-03-29T00:00:00.000Z
- Tags: courses, learning, csharp, reading, Human-Written
Here’s my full C# and WPF course I did between 2012 and 2013. The course starts for beginners and push on for more advanced topics.
#### Content
Here’s my full C# and WPF course I did between 2012 and 2013. The course starts for beginners and push on for more advanced topics. You can take a look at [my slideshare account](http://www.slideshare.net/ZGTRZGTR) for this (and other) courses. Direct links for each slide is here:
1. [C#+WPF L01 - Introduction](http://www.slideshare.net/ZGTRZGTR/cwpf-l01-intro "C#+WPF L01 - Introduction")
2. [C#+WPF L02 - Classes and Objects](http://www.slideshare.net/ZGTRZGTR/cwpf-l02-classes-and-objects "C#+WPF L02 - Classes and Objects")
3. [C#+WPF L03 - Utility P1](http://www.slideshare.net/ZGTRZGTR/cwpf-l03-utility-p1 "C#+WPF L03 - Utility P1")
4. [C#+WPF L04 - Utility P2](http://www.slideshare.net/ZGTRZGTR/cwpf-l04-utility-p2 "C#+WPF L04 - Utility P2")
5. [C#+WPF L05 - Threading and Event Handling](http://www.slideshare.net/ZGTRZGTR/cwpf-l05-threading-and-event-handling "C#+WPF L05 - Threading and Event Handling")
6. [C#+WPF L06 - Collections](http://www.slideshare.net/ZGTRZGTR/cwpf-l06-collections "C#+WPF L06 - Collections")
7. [C#+WPF L07 - LINQ](http://www.slideshare.net/ZGTRZGTR/cwpf-l07-linq)
8. [C#+WPF L08 – Attributes, Reflection and Objects Cloning](http://www.slideshare.net/ZGTRZGTR/cwpf-l08-attributes-reflection-and-objects-cloning)
9. [C#+WPF L09 - WPF P1](http://www.slideshare.net/ZGTRZGTR/cwpf-l09-wpf-p1 "C#+WPF L09 - WPF P1")
10. [C#+WPF L10 - WPF P2](http://www.slideshare.net/ZGTRZGTR/cwpf-l10-wpf-p2 "C#+WPF L10 - WPF P2")
11. [C#+WPF L11 - WPF P3](http://www.slideshare.net/ZGTRZGTR/cwpf-l11-wpf-p3 "C#+WPF L11 - WPF P3")
### C++.NET Windows Forms Course
- URL: https://mohammadshaker.com/en/blog/c-net-windows-forms-course
- Date: 2014-03-29T00:00:00.000Z
- Tags: courses, learning, Human-Written
Here’s my C++.NET (windows forms) course I did between 2011 and 2013. The course is intended for those who have prior knowledge in C++.
#### Content
Here’s my C++.NET (windows forms) course I did between 2011 and 2013. The course is intended for those who have prior knowledge in C++. This course was given to give a concept on User Interfaces and Event-driven Programming. You can take a look at [my slideshare account](http://www.slideshare.net/ZGTRZGTR) for this (and other) courses. Direct links for each slide is here:
1. [C++ Windows Forms L01 - Intro](http://www.slideshare.net/ZGTRZGTR/c-windows-forms-l01 "C++ Windows Forms L01 - Intro")
2. [C++ Windows Forms L02 - Controls P1](http://www.slideshare.net/ZGTRZGTR/c-windows-forms-l02 "C++ Windows Forms L02 - Controls P1")
3. [C++ Windows Forms L03 - Controls P2](http://www.slideshare.net/ZGTRZGTR/cpp-net-l03-controls-p2 "C++ Windows Forms L03 - Controls P2")
4. [C++ Windows Forms L04 - Controls P3](http://www.slideshare.net/ZGTRZGTR/cpp-net-l04-controls-p3 "C++ Windows Forms L04 - Controls P3")
5. [C++ Windows Forms L05 - Controls P4](http://www.slideshare.net/ZGTRZGTR/cpp-net-l05-controls-p4-30977207 "C++ Windows Forms L05 - Controls P4")
6. [C++ Windows Forms L06 - Utlitity and Strings](http://www.slideshare.net/ZGTRZGTR/c-windows-forms-l06-utlitity-and-strings "C++ Windows Forms L06 - Utlitity and Strings")
7. [C++ Windows Forms L07 - Collections](http://www.slideshare.net/ZGTRZGTR/c-windows-forms-l07-collections "C++ Windows Forms L07 - Collections")
8. [C++ Windows Forms L08 - GDI P1](http://www.slideshare.net/ZGTRZGTR/c-windows-forms-l08-gdi-p1 "C++ Windows Forms L08 - GDI P1 ")
9. [C++ Windows Forms L09 - GDI P2](http://www.slideshare.net/ZGTRZGTR/c-windows-forms-l09-gdi-p2 "C++ Windows Forms L09 - GDI P2")
10. [C++ Windows Forms L10 - Instantiate](http://www.slideshare.net/ZGTRZGTR/cpp-net-l10-instantiate "C++ Windows Forms L10 - Instantiate")
11. [C++ Windows Forms L11 - Inheritance](http://www.slideshare.net/ZGTRZGTR/c-windows-forms-l11-inheritance "C++ Windows Forms L11 - Inheritance ")
\[polldaddy poll=7923818\]
### Intro to Event-Driven Programming and Forms with Delphi
- URL: https://mohammadshaker.com/en/blog/intro-to-event-driven-programming-and-forms-with-delphi
- Date: 2014-03-29T00:00:00.000Z
- Tags: courses, learning, Human-Written
This is the first course I've taught back in 2010! I voluntarily taught this course at the Faculty of Information Technology Engineering of Damascus,
#### Content
This is the first course I've taught back in 2010! I voluntarily taught this course at the Faculty of Information Technology Engineering of Damascus, Syria to the first year students (freshmen) when I was on my second year during the midterm holiday. This was, as said, for freshmen who have a knowledge in Pascal (since it was taught as a preparatory programming language in the first semester of the first year.) You can take a look at [my slideshare account](http://www.slideshare.net/ZGTRZGTR) for this (and other) courses. Direct links for each slide is here:
1. [Delphi L01 Intro](http://www.slideshare.net/ZGTRZGTR/delphi-l01-intro "Delphi L01 Intro")
2. [Delphi L02 Controls P1](http://www.slideshare.net/ZGTRZGTR/delphi-l02-controls-p1-30985763 "Delphi L02 Controls P1")
3. [Delphi L03 Forms and Input](http://www.slideshare.net/ZGTRZGTR/delphi-l03-forms-and-input "Delphi L03 Forms and Input")
4. [Delphi L04 Controls P2](http://www.slideshare.net/ZGTRZGTR/delphi-l04-controls-p2 "Delphi L04 Controls P2")
5. [Delphi L05 Files and Dialogs](http://www.slideshare.net/ZGTRZGTR/delphi-l05-files-and-dialogs "Delphi L05 Files and Dialogs")
6. [Delphi L06 GDI Drawing](http://www.slideshare.net/ZGTRZGTR/delphi-l06-gdi-drawing "Delphi L06 GDI Drawing")
7. [Delphi L07 Controls at Runtime P1](http://www.slideshare.net/ZGTRZGTR/delphi-l07-controls-at-runtime-p1 "Delphi L07 Controls at Runtime P1")
8. [Delphi L08 Controls at Runtime P2](http://www.slideshare.net/ZGTRZGTR/delphi-l08-controls-at-runtime-p2 "Delphi L08 Controls at Runtime P2")
\[polldaddy poll=7923818\]
### OpenGL Crash Course
- URL: https://mohammadshaker.com/en/blog/opengl-crash-course
- Date: 2014-03-29T00:00:00.000Z
- Tags: courses, learning, Human-Written
Here’s a very short introduction on computer graphics and OpenGL. This short course was given in 2012 as an introductory session on computer graphics.
#### Content
Here’s a very short introduction on computer graphics and OpenGL. This short course was given in 2012 as an introductory session on computer graphics. The course is intended for those who have no prior knowledge computer graphics. You can take a look at [my slideshare account](http://www.slideshare.net/ZGTRZGTR) for this (and other) courses. Direct links for each slide is here:
1. [OpenGL Starter L01](http://www.slideshare.net/ZGTRZGTR/open-gl-starter-l01 "OpenGL Starter L01")
2. [OpenGL Starter L02](http://www.slideshare.net/ZGTRZGTR/open-gl-starter-l02 "OpenGL Starter L02")
### Car Dynamics [Physics Simulation]
- URL: https://mohammadshaker.com/en/blog/469
- Date: 2014-03-27T00:00:00.000Z
- Tags: Car, Physics, Simulation, Human-Written
This is my project in my third year of studying in the Faculty of Information Technology Engineering in Damascus, Syria, 2011 with Ismaeel Abo Abdalla,
#### Content
This is my project in my third year of studying in the Faculty of Information Technology Engineering in Damascus, Syria, 2011 with Ismaeel Abo Abdalla, Zaher Wanli and Mhd Noor Alhamwi. The project simulates the physics of the car movement with/without Anti Brake-Lock System (ABS), Electronic Stability Program (ESP) and Global Positioning System (GPS) all in realtime.
[youtube.com](http://www.youtube.com/watch?v=8NtnupALEh4)
You can download the presentation and the doc from my slideshare account. A full documentation is available \[in Arabic - بالعربية\] [here](http://www.slideshare.net/ZGTRZGTR/documentation-of-car-dynamics-with-abs-esp-and-gps-systems).
### XNA Game Development Course
- URL: https://mohammadshaker.com/en/blog/xna-game-development-course
- Date: 2014-03-27T00:00:00.000Z
- Tags: courses, learning, shaders, game-development, Human-Written
Here's my full XNA Graphics Language Course I did between 2012 and 2013. The course is intended for beginner in computer graphics languages.
#### Content
Here's my full XNA Graphics Language Course I did between 2012 and 2013. The course is intended for beginner in computer graphics languages. The course gives a glimpse for more advanced topics (shaders) and efficient computer graphics techniques and implementation. You can take a look at [my slideshare account](http://www.slideshare.net/ZGTRZGTR) for this (and other) courses. Direct links for each slide is here:
1. [XNA Game Development L01 - Introduction](http://www.slideshare.net/ZGTRZGTR/xna-l01-introduction-31096167 "XNA Game Development L01 - Introduction")
2. [XNA Game Development L02 – Primitives, IndexBuffer and Vertex Buffer](http://www.slideshare.net/ZGTRZGTR/xna-l02-primitives-index-buffer-and-vertexbuffer-31096212)
3. [XNA Game Development L03 – Basic Matrices and Transformation](http://www.slideshare.net/ZGTRZGTR/xna-l03-basic-matrices-and-transformations)
4. [XNA Game Development L04 – Models, Basic Effect and Animation](http://www.slideshare.net/ZGTRZGTR/xna-l04-models-basic-effect-and-animation)
5. [XNA Game Development L05 – Input, Audio and Video Playback](http://www.slideshare.net/ZGTRZGTR/xna-l05-input-audio-and-video-playback)
6. [XNA Game Development L06 – 2D Graphics and Particles System](http://www.slideshare.net/ZGTRZGTR/xna-l06-2-d-graphics-and-particle-engines)
7. [XNA Game Development L07 - Texturing](http://www.slideshare.net/ZGTRZGTR/xna-l07-texturing)
8. [XNA Game Development L08 – Skybox and Terrain](http://www.slideshare.net/ZGTRZGTR/xna-l08-skybox-and-terrain "XNA Game Development L08 – Skybox and Terrain")
9. [XNA Game Development L09 – Amazing XNA Utilities](http://www.slideshare.net/ZGTRZGTR/xna-l09-amazing-xna-utilities "XNA Game Development L09 – Amazing XNA Utilities")
10. [XNA Game Development L10 – Shaders Part 1](http://www.slideshare.net/ZGTRZGTR/xna-l10-shaders-part-1)
11. [XNA Game Development L11 – Shaders Part 2](http://www.slideshare.net/ZGTRZGTR/xna-l11-shaders-part-2)
### Crospell Engine - Natural Language Processing Engine
- URL: https://mohammadshaker.com/en/blog/crospell-engine-natural-language-processing-engine
- Date: 2014-03-26T00:00:00.000Z
- Tags: Images, Natural Language Processing, NLP, Text, Human-Written
This slide covers CROSPELL Engine; an engine made with multiple approaches for Natural Language Processing.
#### Content
This slide covers CROSPELL Engine; an engine made with multiple approaches for Natural Language Processing. It covers a wide variety of topics in text and image processing. From spell checking to topics prediction. It's a project made in late 2012 and delivered in early 2013 at the F.I.T.E of Damascus, Syria as the final project in NLP course (with Ola Al Naameh and Mhd Hasan Sarhan.) **System Specification (Implementation details can be found in the doc.)** **1\. Auto-correction** The auto-correction algorithm make sure that the misspelled word is matched with a proper correct word. Many approaches can be implemented for this. The option I opted to is the distance between keys on the keyboard map. [](/blog-images/2014/03/cr2.png) But ones should make sure he got the right algorithm. Keys on the keyboard map are not scattered linearly. [](/blog-images/2014/03/cr3.png) The distance between keys are also not linear. The best thing for this is Gaussian curve to measure the right distance. [](/blog-images/2014/03/cr4.png) The CyperSpell Algorithm maps the (possible) misspelled words with their correct-spelled counterparts (using a dictionary). [](/blog-images/2014/03/cr5.png) The user can, in realtime, write and the system will auto-correct (or suggest) the correct words when the user misspell. The system also knows what words the user has misspelled before and rank their chosen correct words higher in the list of suggestions. **** **2\. Language Identification** The user can input any language and the system can figure out what than language is (as long as the corresponding corpses are provided). [](/blog-images/2014/03/cr6.png) if there are more than one language in the text, the system will list them (rank them) according to their occurrences (frequencies in the text). [](/blog-images/2014/03/cr7.png) **3. Word Prediction** Using bi-grams and tri-grams the system can successfully suggest auto-completion while writing words. [](/blog-images/2014/03/cr8.png) **3\. Topic Prediction** Using bi-grams and tri-grams the system can successfully suggest the best topic that match the paragraph. The system, actually, lists all the possible topics prediction and rank them according to the best match. [](/blog-images/2014/03/cr9.png) **4\. Dictionary** The system also provide and Arabic-English dictionary. [](/blog-images/2014/03/cr10.png) **5\. Image Processing using NLP Approaches** Using Minimum Edit Distance (MED), we can match images with others having similar properties (colors in our case). Though, this approach is shallow since it fail completely when images are re-sized or rotated. Anyway, it's just for fun! [](/blog-images/2014/03/cr11.png) The system can best compare images having similar sized and not-transformed. [](/blog-images/2014/03/cr12.png) [](/blog-images/2014/03/cr12.png) **6\. ISRI and Porter Stemming Algorithms** Both, ISRI and Porter stemming algorithms are implemented in the engine. [](/blog-images/2014/03/cr13.png) [](/blog-images/2014/03/cr14.png) **7. Genome Matching using Minimum Edit Distance** The engine interestingly implement Genome matching using MED. The initial interface is: [](/blog-images/2014/03/cr15.png) The user can input two genomes and the system will find the match between the two. [](/blog-images/2014/03/cr16.png) [](/blog-images/2014/03/cr17.png) **8. Sentiment Analysis** The system implement a light sentiment analyzer. Just write a sentence or a paragraph and the system will provide the corresponding emotion for it. [](/blog-images/2014/03/cr18.png) You can download the full project documentation \[in Arabic - بالعربية\] [here](http://www.slideshare.net/ZGTRZGTR/crospell-all). I would be happy to upload the engine source code along with its interface for anyone to use! but the languages corpus are quite big (the project in 400 MB!) so if anyone is interested don't hesitate to contact me by mail and I'll figure something out!
### Mobile Software Engineering Crash Course
- URL: https://mohammadshaker.com/en/blog/mobile-software-engineering-crash-course-c01-intro
- Date: 2014-03-26T00:00:00.000Z
- Tags: Android, Course, iOS, iPhone, Mobile, WindowsPhone, Human-Written
So here we go! This is my (very) short (crash) course slides on Mobile Software Engineering for Android, iPhone and Windows Phone I did back in August,
#### Content
So here we go! This is my (very) short (crash) course slides on Mobile Software Engineering for Android, iPhone and Windows Phone I did back in August, 2012. This is only the first slide of the course. The slides are uploaded in [my SlideShare account](http://www.slideshare.net/ZGTRZGTR). You can also use the direct links here:
1. [Mobile Software Engineering Crash Course - C01 Intro](http://www.slideshare.net/ZGTRZGTR/c01-intro)
2. [Mobile Software Engineering Crash Course - C02 Java Primer](http://www.slideshare.net/ZGTRZGTR/c02-java-primer)
3. [Mobile Software Engineering Crash Course - C03 Android](http://www.slideshare.net/ZGTRZGTR/c03-android)
4. [Mobile Software Engineering Crash Course - C04 Android Cont.](http://www.slideshare.net/ZGTRZGTR/mobile-software-engineering-crash-course-c04-android-cont)
5. [Mobile Software Engineering Crash Course - C05 iOS Intro](http://www.slideshare.net/ZGTRZGTR/c05-i-os-15001213)
6. [Mobile Software Engineering Crash Course - C06 WindowsPhone](http://www.slideshare.net/ZGTRZGTR/mobile-software-engineering-crash-course-c06-windowsphone)
7. [Mobile Software Engineering Crash Course - C07 Frameworks and Conclusion](http://www.slideshare.net/ZGTRZGTR/c07-frameworks)
Your feed back is (very) welcome!
### Social Relationship and Decision-Making Explained by Fuzzy Logic
- URL: https://mohammadshaker.com/en/blog/social-relationship-and-decision-making-explained-by-fuzzy-logic
- Date: 2014-03-26T00:00:00.000Z
- Tags: Decision Making, Fuzzy Logic, Seminar, Social Relationship, Human-Written
This slide was part of my Fuzzy Logic seminar back in 2013 at the Faculty of Information Technology Engineering of Damascus, Syria in the subject of Fuzzy
#### Content
This slide was part of my Fuzzy Logic seminar back in 2013 at the Faculty of Information Technology Engineering of Damascus, Syria in the subject of Fuzzy Logic. The seminar encounter the principle of decision making from fuzzy logic perspective, giving an insight fuzzy-social decision making in a fuzzy social environment. The study is based on this [research paper](http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=4630664&url=http%3A%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D4630664).
### Cockroaches Motivation
- URL: https://mohammadshaker.com/en/blog/cockroaches-motivation
- Date: 2014-03-14T00:00:00.000Z
- Tags: quick-read, Human-Written
\"One task that the cockroaches performed was relatively easy: the roach had to run down a straight corridor.
#### Content
"One task that the cockroaches performed was relatively easy: the roach had to run down a straight corridor. The other, more difficult task required the roach to navigate a somewhat complex maze. As you might expect (assuming you have expectations about roaches), the insects performed the simpler runway task much more quickly when another roach was observing them. The presence of another roach increased their motivation, and, as a consequence, they did better. However, in the more complex maze task, they struggled to navigate their way in the presence of an audience and did much worse than when they performed the same complex task alone. So much for the benefits of social pressure." Ariely, Dan. “The Upside of Irrationality.”
### Personalizing Player Experience in First-Person Shooter Games, UMAP 2013
- URL: https://mohammadshaker.com/en/blog/personalizing-player-experience-in-first-person-shooter-games-poster-umap-2013
- Date: 2014-03-13T00:00:00.000Z
- Tags: AI, arabic, Human-Written
The presentation and poster was part of my fourth year project in Information Technology Engineering, Artificial Intelligence Department, Damascus, Syria.
#### Content
The presentation and poster was part of my fourth year project in Information Technology Engineering, Artificial Intelligence Department, Damascus, Syria. This research study take games development concept to a new level, especially the so called First Person Shooter (FPS) Games. This study outline three basic models: FPS Game Level Design and Procedural Content Generation for FPS games, Preference Learning and Adaptive Content Generation. The framework has been integrated with CUBE opensource game engine. A conference paper has been published in UMAP 2013 which you can find in the [publication section](http://mohammadshakergtr.wordpress.com/publication/ "Publications") (with Noor Shaker, Mehdi Zonji, Ismaeel Abu Abdalla and Mhd Hasan Sarhan.) In this paper (abstract), we describe a methodology for capturing player experience while interacting with a game and we present a data-driven approach for modelling this interaction. We believe the best way to adapt games to a specific player is to use quantitative models of player experience derived from the in-game interaction. Therefore, we rely on crowd-sourced data collected about game context, players behaviour and players self-reports of different affective states. Based on this information, we construct estimators of player experience using neuroevolutionary preference learning. We present the experimental setup and the results obtained from a recent case study where accurate estimators were constructed based on information collected from players playing a first person shooter game. The framework presented is part of a bigger picture where the generated models are utilized to tailor content generation to particular player's needs and playing characteristics. Authors are: Noor Shaker, Mohammad Shaker, Ismaeel Abuabdallah, Mehdi Zonjy, and Mhd Hasan Sarhan.
The poster we presented in the Extended Proceedings of the 2013 Conference on User Modeling, Adaptation and Persolization (UMAP 2013), 2013.
You can download the full project documentation \[in Arabic - بالعربية\] [here](http://www.slideshare.net/ZGTRZGTR/fps-all).
### Ropossum: Cut the Rope Authoring Tool Poster, AIIDE 2013
- URL: https://mohammadshaker.com/en/blog/ropossum-cut-the-rope-authoring-tool-poster-aiide-2013
- Date: 2014-03-13T00:00:00.000Z
- Tags: AI, physics, Human-Written
This is the poster we presented in the Proceedings of Artificial Intelligence and Interactive Digital Entertainment (AIIDE 13), 2013.
#### Content
This is the poster we presented in the Proceedings of Artificial Intelligence and Interactive Digital Entertainment (AIIDE 13), 2013. You can take a look at the papers in the [publication](http://mohammadshakergtr.wordpress.com/publication/) section. You can also download the two papers from [here](/blog-images/2013/08/ctr-playability-final.pdf) and [here](/blog-images/2013/08/ctr-demonstration-final.pdf). In these two paper we present Ropossum, an authoring tool for the generation and testing of levels of the physics-based game, Cut the Rope. Ropossum integrates many features: (1) automatic design of complete solvable content, (2) incorporation of designer’s input through the creation of complete or partial designs, (3) automatic check for playability and (4) optimization of a given design based on playability. The system includes a physics engine to simulate the game and an evolutionary framework to evolve content as well as an AI reasoning agent to check for playability. The system is optimised to allow on-line feedback and realtime interaction. The authors are Mohammad Shaker \[Me\] with Noor Shaker and Julian Togelius.
### The Truth About Relativity
- URL: https://mohammadshaker.com/en/blog/the-truth-about-relativity
- Date: 2014-03-13T00:00:00.000Z
- Tags: quick-read, Human-Written
"I recently worked on a research project examining how one’s own “attractiveness” affects one’s view of the “attractiveness” of others.
#### Content
 "I recently worked on a research project examining how one’s own “attractiveness” affects one’s view of the “attractiveness” of others. One of his good friends, in fact, is a founder of PayPal and is worth tens of millions. But Hong knows how to make the circles of comparison in his life smaller, not larger. In his case, he started by selling his Porsche Boxster and buying a Toyota Prius in its place. “I don’t want to live the life of a Boxster,” he told the New York Times, “because when you get a Boxster you wish you had a 911, and you know what people who have 911s wish they had? They wish they had a Ferrari.” That’s a lesson we can all learn: the more we have, the more we want. And the only cure is to break the cycle of relativity." Ariely, Dan. “Predictably Irrational.”
### Short-Term over Long-Term Objectives
- URL: https://mohammadshaker.com/en/blog/short-term-over-long-term-objectives
- Date: 2014-03-09T00:00:00.000Z
- Tags: quick-read, Human-Written
\"Sadly, most of us often prefer immediately gratifying short-term experiences over our long-term objectives.
#### Content

"Sadly, most of us often prefer immediately gratifying short-term experiences over our long-term objectives. We routinely behave as if sometime in the future, we will have more time, more money, and feel less tired or stressed. “Later” seems like a rosy time to do all the unpleasant things in life, even if putting them off means eventually having to grapple with a much bigger jungle in our yard, a tax penalty, the inability to retire comfortably, or an unsuccessful medical treatment. In the end, we don’t need to look far beyond our own noses to realize how frequently we fail to make short-term sacrifices for the sake of our long-term goals."
"In a perfectly rational world, procrastination would never be a problem. We would simply compute the values of our long-term objectives, compare them to our short-term enjoyments, and understand that we have more to gain in the long term by suffering a bit in the short term. If we were able to do this, we could keep a firm focus on what really matters to us. We would do our work while keeping in mind the satisfaction we’d feel when we finished our project."
Ariely, Dan. "The Upside of Irrationality."
### Motivation Backfire!
- URL: https://mohammadshaker.com/en/blog/motivation-backfire
- Date: 2014-03-08T00:00:00.000Z
- Tags: quick-read, Human-Written
\"The more cognitive skill involved, we thought, the more likely that very high incentives would backfire.
#### Content
"The more cognitive skill involved, we thought, the more likely that very high incentives would backfire. We also thought that higher rewards would more likely lead to higher performance when it came to noncognitive, mechanical tasks. For example, what if I were to pay you for every time you jump in the next twenty-four hours? Wouldn’t you jump a lot, and wouldn’t you jump more if the payment were higher? Would you reduce your jumping speed or stop while you still had the ability to keep going if the amount were very large? Unlikely. In cases where the tasks are very simple and mechanical, it’s hard to imagine that very high motivation would backfire." Ariely, Dan. "The Upside of Irrationality."
### Styx Foodiac Nutrition System
- URL: https://mohammadshaker.com/en/blog/styx-foodiac-nutrition-system
- Date: 2014-03-05T00:00:00.000Z
- Tags: Artificial Intelligence, Damascus, Diet, Expert Systems, Knowledge Base System, Human-Written
This project was part of KBS (Knowledge Base System) subject in the university of Damascus, Syria - department of AI with Mehdi Zonji, Mhd Hasan Sarhan
#### Content
This project was part of KBS (Knowledge Base System) subject in the university of Damascus, Syria - department of AI with Mehdi Zonji, Mhd Hasan Sarhan and Ismaeel Abu-Abdalla. The project was an evolutionary system dedicated to help both the patient and the doctor to get the best diet possible in light of the doctor's instructions and prescriptions and the patient's preferences and needs. It monitors the patient condition and how well he is doing over time and try to feed forward his dietary process in a manner he likes and wish for constrained by the doctor's instructions. Taking the patient's preferences (constraints in a genetic algorithm on fitness function) and merging them with a rule based system (to determine the right prescription) was a new approach that makes a diet an enjoyable experience. The system can suggest meals for the patient taking into consideration the patient's current weight, weight history, food he like or dislike... etc. The system provides a rich, responsive and beautiful UI (WPF + PRISM) that can be used by the doctor and the patient to track the overall dietary process. This project has been awarded the best project of NLP - FIT Damascus, Syria 2012.
### "Loss Aversion"
- URL: https://mohammadshaker.com/en/blog/loss-aversion
- Date: 2014-03-04T00:00:00.000Z
- Tags: quick-read, Human-Written
\\\"Loss aversion is the simple idea that the misery produced by losing something that we feel is ours—say, money—outweighs the happiness of gaining the
#### Content

"Loss aversion is the simple idea that the misery produced by losing something that we feel is ours—say, money—outweighs the happiness of gaining the same amount of money. For example, think about how happy you would be if one day you discovered that due to a very lucky investment, your portfolio had increased by 5 percent. Contrast that fortunate feeling to the misery that you would feel if, on another day, you discovered that due to a very unlucky investment, your portfolio had decreased by 5 percent. If your unhappiness with the loss would be higher than the happiness with the gain, you are susceptible to loss aversion. (Don’t worry; most of us are.)" Ariely, Dan. "The Upside of Irrationality."
### Depression and Happy-Sad People
- URL: https://mohammadshaker.com/en/blog/depression-with-happy-people
- Date: 2014-03-02T00:00:00.000Z
- Tags: quick-read, Human-Written
\"This is one of those observations that is both obvious and (upon exploration) deeply profound, and it explains all kinds of otherwise puzzling
#### Content
 "This is one of those observations that is both obvious and (upon exploration) deeply profound, and it explains all kinds of otherwise puzzling observations. Which do you think, for example, has a higher suicide rate: countries whose citizens declare themselves to be very happy, such as Switzerland, Denmark, Iceland, the Netherlands, and Canada? or countries like Greece, Italy, Portugal, and Spain, whose citizens describe themselves as not very happy at all? Answer: the so-called happy countries. It’s the same phenomenon as in the Military Police and the Air Corps. If you are depressed in a place where most people are pretty unhappy, you compare yourself to those around you and you don’t feel all that bad. But can you imagine how difficult it must be to be depressed in a country where everyone else has a big smile on their face?" Gladwell, Malcolm. “David and Goliath: Underdogs, Misfits, and the Art of Battling Giants.”
### Prisons?
- URL: https://mohammadshaker.com/en/blog/poisoning
- Date: 2014-03-01T00:00:00.000Z
- Tags: quick-read, Human-Written
\"Prison has a direct effect on crime: it puts a bad person behind bars, where he can’t victimize anyone else.
#### Content
"Prison has a direct effect on crime: it puts a bad person behind bars, where he can’t victimize anyone else. But it also has an indirect effect on crime, in that it affects all the people with whom that criminal comes into contact. A very high number of the men who get sent to prison, for example, are fathers. (One-fourth of juveniles convicted of crimes have children.) And the effect on a child of having a father sent away to prison is devastating. Some criminals are lousy fathers: abusive, volatile, absent. But many are not. Their earnings—both from crime and legal jobs—help support their families. For a child, losing a father to prison is an undesirable difficulty. Having a parent incarcerated increases a child’s chances of juvenile delinquency between 300 and 400 percent; it increases the odds of a serious psychiatric disorder by 250 percent." Gladwell, Malcolm. “David and Goliath: Underdogs, Misfits, and the Art of Battling Giants.”
### Risk and Gain
- URL: https://mohammadshaker.com/en/blog/risk-and-gain
- Date: 2014-02-22T00:00:00.000Z
- Tags: quick-read, reading, Human-Written
\\\"In the book that Pallop was reading by Kahneman and Tversky, for example, there is a description of a simple experiment, where a group of people were
#### Content

"In the book that Pallop was reading by Kahneman and Tversky, for example, there is a description of a simple experiment, where a group of people were told to imagine that they had $300. They were then given a choice between (a) receiving another $100 or (b) tossing a coin, where if they won they got $200 and if they lost they got nothing. Most of us, it turns out, prefer (a) to (b). But then Kahneman and Tversky did a second experiment. They told people to imagine that they had $500 and then asked them if they would rather (c) give up $100 or (d) toss a coin and pay $200 if they lost and nothing at all if they won. Most of us now prefer (d) to (c). What is interesting about those four choices is that, from a probabilistic standpoint, they are identical. Nonetheless, we have strong preferences among them. Why? Because we’re more willing to gamble when it comes to losses, but are risk averse when it comes to our gains. That’s why we like small daily winnings in the stock market, even if that requires that we risk losing everything in a crash."
Gladwell, Malcolm. “What the Dog Saw"
### Startup Weekend Damascus, Weebee: A Game That Absorbs the Child Behavior and Change It
- URL: https://mohammadshaker.com/en/blog/a-game-that-absorbs-the-child-behavior-and-change-it
- Date: 2014-02-19T00:00:00.000Z
- Tags: games, game-development, arabic, startups, Human-Written
Welcome to a new game genre! This is our idea (with Rawan Al-Omari, Zeina Al-Helwani, Walaa Baghdadi, Majd Massijeh and Mohanad Al-Helwany and supervised
#### Content

Welcome to a new game genre! This is our idea (with Rawan Al-Omari, Zeina Al-Helwani, Walaa Baghdadi, Majd Massijeh and Mohanad Al-Helwany and supervised by Dr. Noor Shaker) in Startup Weekend Damascus; #SWDamascus.
Weebee is a game intended (for now) for children. It's both a complete new research study and an implementation of a game that absorbs the child behavior and change it. We would like to have your opinion on this by filling in the following survey [Here](https://docs.google.com/forms/d/1Y5JoShBZM1dwtSJMfaDiSl37ev7gUm6_7rpNwr6Nt9Y/viewform) both in Arabic and English.
Thank you for your participation!
### Drift Your Mind
- URL: https://mohammadshaker.com/en/blog/drift-your-mind
- Date: 2014-02-11T00:00:00.000Z
- Tags: quick-read, Human-Written
\"Every injection day, I would stop at the video store on the way to school and pick up a few films that I wanted to see.
#### Content
"Every injection day, I would stop at the video store on the way to school and pick up a few films that I wanted to see. Throughout the day, I would think about how much I would enjoy watching them later. Once I got home, I would give myself the injection. Then I would immediately jump into my hammock, make myself comfortable, and start my mini film fest. That way, I learned to associate the act of the injection with the rewarding experience of watching a wonderful movie. Eventually, the negative side effects kicked in, and I didn’t have such a positive feeling. Still, planning my evenings that way helped me associate the injection more closely with the fun of watching a movie than with the discomfort of the side effects, and thus I was able to continue the treatment. (I was also fortunate, in this instance, that I have a relatively poor memory, which meant that I could watch some of the same movies over and over again.)" Ariely, Dan. "The Upside of Irrationality."
### Blind by Our Own Eyes
- URL: https://mohammadshaker.com/en/blog/blind-by-our-own-eyes
- Date: 2014-02-04T00:00:00.000Z
- Tags: quick-read, Human-Written
\"When a person looks out at the world, he sees it filtered through a screen of his words, and this process is as invisible to him as water is to fish.
#### Content

"When a person looks out at the world, he sees it filtered through a screen of his words, and this process is as invisible to him as water is to fish. We see the world and our words in one impression, as if we’re looking at a forest through a green filter. We can’t see what’s really green and what’s not. If we were to walk around with the filter in our eye long enough, we’d forget it was there, and life would just be green." Dave, Logan. “Tribal Leadership.”
### Emphasize Choice
- URL: https://mohammadshaker.com/en/blog/emphasize-choice
- Date: 2014-02-03T00:00:00.000Z
- Tags: quick-read, Human-Written
Frank Jordan told us: “I tell people they have a choice, and at first they don’t believe me.
#### Content
 Frank Jordan told us: “I tell people they have a choice, and at first they don’t believe me. But I say, ‘I’m like you; I grew up with one parent and not much to do, and I chose a life of service, and so can you.’” There are two aspects of his pitch that are noteworthy. First, he doesn’t look down on people, no matter their cultural stage, and as a result he slowly builds rapport. Second, he emphasizes the one thing that Stage One doesn’t see: choice. As he told us: “If people can see that they have a choice, they sometimes choose a life better than gangs and drugs.” Dave, Logan. “Tribal Leadership.”
### A Lost Generation
- URL: https://mohammadshaker.com/en/blog/lost-generation
- Date: 2014-02-01T00:00:00.000Z
- Tags: quick-read, Human-Written
\"In August 2010, the International Labor Organization (ILO) published its report on Global Employment Trends for Youth 2010.
#### Content
"In August 2010, the International Labor Organization (ILO) published its report on Global Employment Trends for Youth 2010. The report concludes that there are approximately 620 million economically active young people worldwide. At the end of 2009, 81 million of them were unemployed; the highest number ever, and almost 8 million more than in 2007. The youth unemployment rate increased from 11.9 percent in 2007 to 13.0 percent in 2009. The ILO argues that these trends will have “significant consequences for young people as upcoming cohorts of new entrants join the ranks of the already unemployed” and warns of the “risk of a crisis legacy of a ‘lost generation’ comprised of young people who have dropped out of the labor market, having lost all hope of being able to work for a decent living." Robinson, Ken. “Out of Our Minds.”
### \"Companies Often...
- URL: https://mohammadshaker.com/en/blog/companies-ofte
- Date: 2014-01-28T00:00:00.000Z
- Tags: Education, Human-Written
> \"Companies often divide the workforce into two groups: the ‘creatives’ and the ‘suits’.
#### Content
> "Companies often divide the workforce into two groups: the ‘creatives’ and the ‘suits’. You can normally tell who the creatives are because they don’t wear suits. They wear jeans and they come in late because they have been struggling with an idea." Robinson, Ken. “Out of Our Minds.
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### Specializations on Coursera
- URL: https://mohammadshaker.com/en/blog/specializations-on-coursera
- Date: 2014-01-28T00:00:00.000Z
- Tags: Coursera, Education, Online education, Human-Written
I was wondering today why Coursera hasn't yet got the initiative to implement full programs, that consist of multiple courses, on specific domains.
#### Content
[](/blog-images/2014/01/untitled.png)
I was wondering today why Coursera hasn't yet got the initiative to implement full programs, that consist of multiple courses, on specific domains. It will be, for instance, like _specializing_ in particular topic in universities master programs. That thing is I was very delighted when I searched for it today and there you go! [Here it is!](https://www.coursera.org/specializations?utm_medium=topnav) Coursera has really done it! It's a very big opportunity for us to get on with online education FOR REAL now. I hope that I can enroll in one of the courses.
### Games Are Awesome... Hmmm.. Also Make You Smarter and Healthier
- URL: https://mohammadshaker.com/en/blog/games-are-awesome-hmmm-also-make-you-smarter-and-healthier
- Date: 2014-01-22T00:00:00.000Z
- Tags: Brain, Game, Life, Lifestyle, Human-Written
https://www.youtube.com/watch?feature=player\embedded&v=OOsqkQytHOs Good to see how games can influence our life.
#### Content
[youtube.com](https://www.youtube.com/watch?feature=player\_embedded&v=OOsqkQytHOs)
Good to see how games can influence our life.
## Context
This is an archival short-form note. This summary is added to make the main idea explicit for search engines and answer engines. It clarifies what this note is about, why it was shared, and what the reader should take away from it.
### The Story of "0"
- URL: https://mohammadshaker.com/en/blog/the-story-of-0
- Date: 2014-01-22T00:00:00.000Z
- Tags: quick-read, Human-Written
\"ZERO had a long history. The Babylonians invented the concept of zero; the ancient Greeks debated it in lofty terms (how could something be nothing?);
#### Content
 "ZERO had a long history. The Babylonians invented the concept of zero; the ancient Greeks debated it in lofty terms (how could something be nothing?); the ancient Indian scholar Pingala paired zero with the numeral 1 to get double digits; and both the Mayans and the Romans made zero part of their numeral systems. But zero really found its place about AD 498, when the Indian astronomer Aryabhata sat up in bed one morning and exclaimed, “Sthanam sthanam dasa gunam”—which translates, roughly, as “Place to place in 10 times in value.” With that, the idea of decimal-based place-value notation was born. Now zero was on a roll: It spread to the Arab world, where it flourished; crossed the Iberian Peninsula to Europe (thanks to the Spanish Moors); got some tweaking from the Italians; and eventually sailed the Atlantic to the New World, where zero ultimately found plenty of employment (together with the digit 1) in a place called Silicon Valley." Ariely, Dan. “Predictably Irrational.”
### The Story of "Free!" - Cont.
- URL: https://mohammadshaker.com/en/blog/the-story-of-free-cont
- Date: 2014-01-21T00:00:00.000Z
- Tags: quick-read, react, Human-Written
“Let me tell you a story that describes the real influence of FREE! on our behavior.
#### Content
 “Let me tell you a story that describes the real influence of FREE! on our behavior. A few years ago, Amazon.com started offering free shipping of orders over a certain amount. Someone who purchased a single book for $16.95 might pay an additional $3.95 for shipping, for instance. But if the customer bought another book, for a total of $31.90, they would get their shipping FREE! Some of the purchasers probably didn’t want the second book (and I am talking here from personal experience) but the FREE! shipping was so tempting that to get it, they were willing to pay the cost of the extra book. The people at Amazon were very happy with this offer, but they noticed that in one place—France—there was no increase in sales. Is the French consumer more rational than the rest of us? Unlikely. Rather, it turned out, the French customers were reacting to a different deal. Here’s what happened instead of offering FREE! shipping on orders over a certain amount, the French division priced the shipping for those orders at one franc. Just one franc—about 20 cents. This doesn’t seem very different from FREE! but it was. In fact, when Amazon changed the promotion in France to include free shipping, France joined all the other countries in a dramatic sales increase. In other words, whereas shipping for one franc—a real bargain—was virtually ignored by the French, FREE! shipping caused an enthusiastic response” Ariely, Dan. “Predictably Irrational.”
### The Story of "Free!"
- URL: https://mohammadshaker.com/en/blog/the-story-of-free
- Date: 2014-01-20T00:00:00.000Z
- Tags: quick-read, react, Human-Written
\"Let me tell you a story that describes the real influence of FREE! on our behavior.
#### Content
"Let me tell you a story that describes the real influence of FREE! on our behavior. A few years ago, Amazon.com started offering free shipping of orders over a certain amount. Someone who purchased a single book for $16.95 might pay an additional $3.95 for shipping, for instance. But if the customer bought another book, for a total of $31.90, they would get their shipping FREE! Some of the purchasers probably didn’t want the second book (and I am talking here from personal experience) but the FREE! shipping was so tempting that to get it, they were willing to pay the cost of the extra book. The people at Amazon were very happy with this offer, but they noticed that in one place—France—there was no increase in sales. Is the French consumer more rational than the rest of us? Unlikely. Rather, it turned out, the French customers were reacting to a different deal. Here’s what happened instead of offering FREE! shipping on orders over a certain amount, the French division..." Ariely, Dan. “Predictably Irrational.” to be continued!
### Artificial Intelligence! Now What?
- URL: https://mohammadshaker.com/en/blog/artificial-intelligence-now-what
- Date: 2014-01-18T00:00:00.000Z
- Tags: Artificial Intelligence, Games, Human-Written
Two years ago I have done a brief introduction on Artificial Intelligence for those who just finished high school and want to get in universities (which I
#### Content
Two years ago I have done a brief introduction on Artificial Intelligence for those who just finished high school and want to get in universities (which I wrote about before, [here](http://mohammadshakergtr.wordpress.com/2012/08/30/ai-brief-introduction-freshman-2012/).) If I would do it again now, I'll do it completely different. My grasp of AI technologies and how much they emerged in our life has been widened upon the narrow view of regular AI approaches. From Google glass and wearables to Xbox One and leap motion, the hardware can deliver awesome platforms for us to implement amazing ideas into our everyday life. From start-ups and one-man-team to indie developers, we are not limited, anymore, with resources and old-fashion ideas (I think the ideas of AI presented in the slides are pretty much old-fashioned and they represent what I was learning in university and they don't represent the real potential of AI.) For instance, my work on games for the last two years really opened up my mind on many issues we will face in the upcoming years. I worked on emotions modelling on two separate projects (in 2012 and 2013, see the [publication section](http://mohammadshakergtr.wordpress.com/publication/) for this paper: "A Quantitative Approach for Modeling and Personalizing Player Experience in First-Person Shooter Games" and my upcoming work for UMAP conference 2014 "Towards More Accurate Player Experience Models: An Exploration of Context, Behavioral, Visual and Affect Features".) [](/blog-images/2014/01/slide-101-728.jpg) My fourth year project in 2012, and thanks to Noor Shaker who proposed and supervised the implementation of the idea, was about personalizing content generation in first person shooter games through player modelling (don't be afraid from the hilarious name) in a complete state of the art approach ([here](http://www.slideshare.net/ZGTRZGTR/styx-engine-adaptive-fps-games-content-generation) for more). I do admit that, upon finishing the project successfully, I was amazed of what our game can do using the model we built. The model can intelligently manipulate the emotions of the player by automatically generating specific game content tailored by the player actions and hisher style while playing the game. The model can for example auto-generate a game level that can maximize the player's engagement in the game or even maximize the player's frustration! This is a two-edged sword since this can be implemented to obtain _certain_ goals to the authoring party or used in a wrong way. [](/blog-images/2014/01/ctr_our.png "A screenshot of level designed by Ropossum") What I've been working on now (2013 and ongoing and supervised by Noor Shaker and Julian Togelius from ITU, Denmark) is an authoring tool for Cut the Rope, named **[Ropossum V1.0](http://mohammadshakergtr.wordpress.com/2014/01/18/ropossum-v1-0/)**. My own version of the game is named [**Cut the Rope: Play Forever**](http://mohammadshakergtr.wordpress.com/2013/02/14/cut-the-rope-play-forever-project/). Ropossum is all about letting you design your own levels, check your designed levels for playability on real time. You can also ask it for help to complete your unfinished designs according to your own preferences or just play the game you want _forever_ only by asking Ropossum to generate a never-ending levels for you. A more futuristic vision is to create a social community for gamers, letting the players rate the designed/authored/generated levels and share them with their friends to create a whole new social community for games generating themselves by the effort of the players! What we can see is that the games now still fall short on technologies that are proposed a while back. We should enjoy the technology we have because up to now, we are, sadly, lacking the courage to go through into the next phase. To be continued!
### Predictably Irrational–P2
- URL: https://mohammadshaker.com/en/blog/predictably-irrational-p2
- Date: 2014-01-18T00:00:00.000Z
- Tags: quick-read, Human-Written
Let's carry on from yesterday! Yesterday's unfinished part is colored gray. \"Now you are on your second task: you’re shopping for your suit.
#### Content
Let's carry on from yesterday! Yesterday's unfinished part is colored gray. "Now you are on your second task: you’re shopping for your suit. You find a luxurious gray pinstripe suit for $455 and decide to buy it, but then another customer whispers in your ear that the exact same suit is on sale for only $448 at another store, just 15 minutes away. Do you make this second 15-minute trip? In this case, most people say that they would not. But what is going on here? Is 15 minutes of your time worth $7, or isn’t it? In reality, of course, $7 is $7—no matter how you count it. The only question you should ask yourself in these cases is whether the trip across town, and the 15 extra minutes it would take, is worth the extra $7 you would save. Whether the amount from which this $7 will be saved is $10 or $10,000 should be irrelevant." Ariely, Dan. “Predictably Irrational.” To be continued!
### Predictably Irrational-P1
- URL: https://mohammadshaker.com/en/blog/predictably-irrational-p1
- Date: 2014-01-17T00:00:00.000Z
- Tags: quick-read, Human-Written
\"Let me explain with an example from a study conducted by two brilliant researchers, Amos Tversky and Daniel Kahneman.
#### Content
"Let me explain with an example from a study conducted by two brilliant researchers, Amos Tversky and Daniel Kahneman. Suppose you have two errands to run today. The first is to buy a new pen, and the second is to buy a suit for work. At an office supply store, you find a nice pen for $25. You are set to buy it, when you remember that the same pen is on sale for $18 at another store 15 minutes away. What would you do? Do you decide to take the 15-minute trip to save the $7? Most people faced with this dilemma say that they would take the trip to save the $7. Now you are on your second task: you’re shopping for your suit. You find a luxurious gray pinstripe suit for $455 and decide to buy it, but then another customer whispers in your ear that the exact same suit is on sale for only $448 at another store, just 15 minutes away. Do you make this second 15-minute trip? In this case, most people say that they.." Ariely, Dan. “Predictably Irrational.” To be continued tomorrow!
### 1-Min-Reading, Oh Ya!
- URL: https://mohammadshaker.com/en/blog/readings-oh-ya
- Date: 2014-01-17T00:00:00.000Z
- Tags: quick-read, reading, Human-Written
So, I have this idea recently of posting, regularly I think, some of the paragraphs I find interesting while reading some of the books.
#### Content
Hi all!
So, I have this idea recently of posting, regularly I think, some of the paragraphs I find interesting while reading some of the books. To make the reading experience of those paragraphs more enjoyable for you, I'll try to make it as tight as possible, _**1-Min-Reading**_. I'll try sometimes to post the paragraphs _unfinished_ so you can come on a day later and carry on reading from the point we stopped. The paragraphs won't be on a particular domain, I'll only post from the books I'm reading and really find enjoyable. I won't go through an entire book or tackle each book in order. Therefore, different subjects will be posted sequentially.
I'll begin with **Predictably Irrational** book by Dan Ariely. I should note though that all the materials are not mine and they all belong to the authors.
Hope it'll be a nice experience for me and you! so let's start!
### Ropossum V1.0: A Physics-Based Game Authoring Tool
- URL: https://mohammadshaker.com/en/blog/ropossum-v1-0
- Date: 2014-01-17T00:00:00.000Z
- Tags: AuthoringTool, CutTheRope, Fun, GameDesign, Ropossum, Human-Written
https://www.youtube.com/watch?v=0SJykSZ5cm4&feature=youtu.be This is May 2013 version of Ropossum. i.e.
#### Content
[youtube.com](https://www.youtube.com/watch?v=0SJykSZ5cm4&feature=youtu.be) This is May 2013 version of Ropossum. i.e. Ropossum V1.0 which serves as my graduation project in the Faculty of Information Technology Engineering in Damascus, Syria. Supervised by Noor Shaker, Julian Togelius and Ammar Joukhadar. There're three published papers so far for Ropossum in IEEE CIG 2013 (where it was nominated for best paper award) and AIIDE 2013 (see the [publication section](http://mohammadshakergtr.wordpress.com/publication/) for more, and Facebook page [here](https://www.facebook.com/CutTheRopePlayForever?ref=hl)). Ropossum V2.0 is coming soon in May 2014 and it'll be amazing!
### Cut the Rope Play Forever and Ropossum Authoring Tool
- URL: https://mohammadshaker.com/en/blog/cut-the-rope-play-forever-project
- Date: 2013-02-13T00:00:00.000Z
- Tags: First-order Logic, GAME DESIGNG, GAME ENGINEG, GRAMMATICAL EVOLUTIONm PHYSICS ENGINE, Procedural Content Generation, Human-Written
Cut The Rope Play Forever is a game with endless levels. The designed framework let you play the beloved Cut The Rope game as much as you want and the
#### Content
Cut The Rope Play Forever is a game with endless levels. The designed framework let you play the beloved Cut The Rope game as much as you want and the levels will keep coming. Using Ropossum authoring tool, You can also design your own levels with the first ever evolutionary framework for evolving playable content for physics-based games. With Ropossum you can check your designed levels for playability on real time, ask it to complete your unfinished designs according to your own preferences. It can even suggest endless playable design variations according to your initial level design. Visit [Cut the Rope Play Forever Project](http://noorshaker.com/CutTheRope.html#content-inner-1 "Cut the Rope Play Forever Project.") and the [publication](http://mohammadshakergtr.wordpress.com/publication/ "Publications") section here for more information on Cut The Rope Play Forever. Don't forget to watch the [trailer](http://www.youtube.com/watch?v=FM3v0tbdKrs&list=UUSv0OrQI0ROI8z1d9X1zxeA) and like the [page](https://www.facebook.com/CutTheRopePlayForever?ref=hl)!
### Gaming and Robotics - Virtual Reality Seminar 2012
- URL: https://mohammadshaker.com/en/blog/gaming-and-robotics-virtual-reality-seminar-2012
- Date: 2012-08-31T00:00:00.000Z
- Tags: thoughts, personal, Human-Written
This is part of my seminar I did in April 2012, at the Faculty of Information Technology Engineering in Damascus, Syria. The slide makes a brief picture
#### Content
[](/blog-images/2012/08/dsc_0111.jpg)
This is part of my seminar I did in April 2012, at the Faculty of Information Technology Engineering in Damascus, Syria. The slide makes a brief picture of gaming in the past, present and future and how some games will change our life. There's also a peak view on Robotics and the latest inventions in that field. Gamification is also a new concept that's taking the heat nowadays and it's really flourishing.
Pictures on Flickr [here](http://www.flickr.com/photos/mohammadshaker/sets/72157631334120182/).
### Artificial Intelligence Brief Introduction @Wikilogia - Freshman 2012
- URL: https://mohammadshaker.com/en/blog/ai-brief-introduction-freshman-2012
- Date: 2012-08-30T00:00:00.000Z
- Tags: thoughts, personal, AI, Human-Written
Update: for an update and a self-discussion followed by this post go to this post.
#### Content
_Update:_ for an update and a self-discussion followed by this post go to [this post](http://mohammadshakergtr.wordpress.com/2014/01/18/artificial-intelligence-now-what/). This was a brief introduction on Artificial Intelligence for those who just finished high school and want to get in universities. This took place in IT Plaza in Damascus, Syria on August 26, 2012. The event was one of _Wikilogia_ meetings for those fellows. You can take a look at the lecture following the [link](http://www.slideshare.net/ZGTRZGTR/ai-brief-introduction-freshman-2012 "link") . Event pictures on Flickr [here](http://www.flickr.com/photos/wikilogia/sets/72157631255651866/) and you can learn more about Wikilogia events from [here](https://www.facebook.com/wikilogia).
Video [link](https://vimeo.com/48354050).
[vimeo.com](http://www.vimeo.com/48354050)
## Blog Posts (Arabic)
### منحنى GAIA على شكل S لفعالية الوكلاء
- URL: https://mohammadshaker.com/ar/blog/gaia-s-curve-of-agent-effectiveness
- Date: 2026-02-13T00:00:00.000Z
فعالية الوكلاء على معيار GAIA تتشكّل كمنحنى S واضح: النماذج اللغوية وحدها تصطدم بسقف مبكر، ثم تقفز الدقة حين يُضاف استخدام الأدوات والتخطيط متعدد الخطوات، قبل أن تتباطأ المكاسب عند الهضبة. أفضل الأنظمة اليوم تقترب من الأداء البشري عند 91-92%، والعامل المميز لم يعد الذكاء بل قلة الأخطاء غير المفروضة.
#### Content
يمكننا التفكير في GAIA كاختبار إجهاد لـ "المساعدين العامين" الذين يجب عليهم القيام بما يفعله البشر بشكل عادي طوال اليوم: العثور على المصدر الصحيح، وقراءته بدقة، والجمع بين خطوات عدة، وتقديم إجابة دقيقة. ليس مجرد "التفكير" بشكل مجرد، بل التنفيذ: التصفح، والاستخراج، والتحقق، والإنهاء بنظافة.
عندما نرسم فعالية الوكلاء على GAIA مقابل الزمن والقدرة، يتشكل بشكل طبيعي منحنى على شكل S.
في بداية المنحنى، لا يساعد سلوك النموذج اللغوي العادي كثيرًا. يمكننا كتابة نص معقول، لكن مهام GAIA تعاقب المعقولية. بدون استخدام منضبط للأدوات، لا يستطيع النظام إما الوصول إلى المعلومات المطلوبة أو تجميعها بشكل موثوق. التحسينات في التوجيه والتفكير الأساسي تحرك المؤشر، لكن ليس بشكل كبير، لأن نمط الفشل هو التنفيذ وليس البلاغة.
ثم يصل منتصف المنحنى ويصبح الميل حادًا. هنا يظهر استخدام الأدوات والتنسيق: البحث، والتصفح، والاستخراج المنظم، والتخطيط متعدد الخطوات، وإعادة المحاولة، والفحوصات الذاتية الأساسية. بمجرد أن يتمكن الوكيل من تنفيذ "البحث ← القراءة ← الحساب ← الإجابة" بشكل متسق بدلاً من التخمين، تقفز الدقة بسرعة. هذا هو الجزء من المنحنى S الذي يبدو فيه التقدم وكأنه يتراكم فجأة.
أخيرًا نصل إلى الهضبة. ليس لأننا توقفنا عن التحسين، بل لأن الأخطاء المتبقية تعيش في الذيل الطويل. النقاط المئوية الأخيرة ليست عن تحسين الحالة الشائعة. إنها عن عدم الانهيار على صفحات فوضوية، وعدم اختيار المصدر الخاطئ عندما تكون مصادر متعددة معقولة، وعدم قراءة جدول خاطئ في ملف PDF، وعدم إسقاط قيد في منتصف الطريق، والتعافي عندما تسوء خطوة مبكرة.
على GAIA تحديدًا، الخط الأساسي البشري يقع تقريبًا في أوائل التسعينات. أفضل أنظمة الوكلاء على لوحة المتصدرين العامة أصبحت الآن هناك أيضًا: المتوسط العام في نطاق ~91-92%. بعبارة أخرى، من حيث GAIA، نحن بالفعل في أعلى يمين الرسم البياني: المرحلة الثالثة، بالقرب من الحد المقارب.
ما يتغير بمجرد أن نكون على تلك الهضبة هو طبيعة العمل المهم. تصبح المعايير أقل عن النتيجة المتوسطة وأكثر عن خصائص الموثوقية: التباين، ومخاطر الذيل، و"تكلفة التصحيح". يتوقف السؤال عن أن يكون "هل يمكننا حل هذا النوع من المهام؟" ويصبح "كم مرة نفشل بطرق مزعجة، وكم يكلف الإنسان لاكتشاف ذلك وإصلاحه؟"
هذا يفسر أيضًا لماذا تبدو مهام المستوى الأول "محلولة أساسًا" بينما المستويات الأصعب لا تزال تسرب أخطاء. العامل المحدد ليس الذكاء الخام؛ إنه المتانة في ظل الغموض والمدخلات الفوضوية من العالم الحقيقي. يمكن للنظام أن يكون بارعًا ومع ذلك يختار الصفحة الخاطئة. يمكنه التفكير بشكل صحيح ومع ذلك يستخرج الرقم الخاطئ من جدول. يمكنه اتباع خطة ومع ذلك يسقط قيدًا بصمت.
إذن أين نحن في الرسم البياني؟ نحن بالفعل في الجزء الذي تأتي فيه المكاسب من الهندسة المملة عالية التأثير: حلقات التحقق، وانضباط المصدر، واستراتيجيات احتياطية أفضل، والتعافي المحكم عندما تنحرف المحاولة الأولى. بمجرد أن يكون المتوسط قريبًا من البشر، لا يكون العامل المميز "ذكاءً أكثر." بل أخطاء أقل غير مفروضة.
### على ماذا سأقضي 10,000 ساعة؟
- URL: https://mohammadshaker.com/ar/blog/what-would-i-spend-10000-hours-on
- Date: 2025-08-06T00:00:00.000Z
بناء المنتجات هو ما أريد استثمار وقتي فيه، لكن الرافعة الحقيقية تأتي من بناء الثقة لا من الهندسة وحدها. الناس يشترون ممن يثقون بهم، وكسب هذه الثقة ببطء عبر انتصارات محلية متراكمة هو ما يجعل مشروعاً مموَّلاً ذاتياً قادراً على الصمود والنمو دون الحاجة لرأس المال المغامر.
#### Content
كتبت هذا في 2021. كنت أفكر أين سأستثمر وقتاً وجهداً كبيراً.
لدي ميل قوي لبناء المنتجات. لكن لدي نقطة ضعف — التسويق والمبيعات. كنت أستخف بهما وأعتبرهما أقل من العمل الهندسي.
هذا الاستخفاف ينبع من عقلية هندسية تقلل من قيمة المهارات الإنسانية كرواية القصص وبناء الثقة. لكن الحقيقة هي: **نحن نعيش بالثقة، برواية القصص**. هذه الصفات لا يمكن اختزالها في مقاييس كمية.
نجاح تبني المنتج يتطلب أن يثق الناس بالصانع بما يكفي لشراء ما بناه.
## الثقة تطورية
أجدادنا عاشوا في قرى صغيرة. السلوك البشري تشكّل حول الثقة والعلاقات الشخصية. الناس يصادقون من يثقون بهم، ويشترون من أصدقائهم.
مفهوم كيفن كيلي "1,000 معجب حقيقي" يلخص هذا تماماً. المشاريع الذاتية التمويل تتطلب كسب مجموعات صغيرة من المؤيدين المخلصين — من 10 إلى 100 إلى 1,000 معجب عبر انتصارات محلية تدريجية.
## لا تمويل استثماري
لدي موقف صارم ضد نموذج رأس المال المغامر. أفضل بناء المشاريع برأس مال شخصي، **بنفس الطريقة التي فعلها جدي.**
### الوكلاء الذكية تلتهم طبقة الأعمال.
- URL: https://mohammadshaker.com/ar/blog/agents-are-eating-the-business-layer
- Date: 2025-06-25T00:00:00.000Z
الوكلاء المدعومون بنماذج اللغة الكبيرة يُعيدون رسم معمارية البرمجيات: بدلاً من الفصل التقليدي بين طبقة العرض ومنطق الأعمال وقاعدة البيانات، يدمج الوكيل هذه الطبقات في طبقة استدلال واحدة تتخذ القرارات باستقلالية عبر استدعاء الأدوات والوظائف. هذا التحول يغيّر طريقة بناء المنتجات من جذورها.
#### Content
الوكلاء المدعومون بنماذج اللغة الكبيرة يحولون معمارية البرمجيات جذرياً. إنهم يدمجون التصميم التقليدي ثلاثي الطبقات في نموذج جديد.
## البنية التقليدية مقابل معمارية الوكلاء
تاريخياً، البرمجيات اتبعت هذا الهيكل: طبقة العرض، طبقة منطق الأعمال، وتخزين البيانات. **الوكلاء المدعومون بنماذج اللغة يدمجون الطبقتين 1 و2 (وأحياناً أجزاء من 3) في طبقة استدلال واحدة.**
في هذا النموذج الناشئ، الوكلاء يصبحون صناع القرار، يختارون من الأدوات المتاحة بدلاً من اتباع مسارات كود محددة مسبقاً.
## التحول المعماري
بدلاً من الشروط المبرمجة وخطوط الخدمات، الوكلاء المجهزون بقدرات استدعاء الوظائف والأدوات وروابط MCP يتعاملون مع صنع القرار بشكل مستقل.
## من أين تبدأ
للمهندسين الراغبين في استكشاف هذا التحول:
1. ابدأ بشريحة ضيقة من سير العمل
2. اكشف نقاط نهاية CRUD
3. غلّفها كأدوات للوكيل
4. طبّق مراجعة بشرية في الحلقة
5. كرر على مستويات الاستقلالية
### الانزعاج الطوعي.
- URL: https://mohammadshaker.com/ar/blog/voluntary-discomfort
- Date: 2025-06-16T00:00:00.000Z
النمو الحقيقي لا يأتي من الراحة بل من المشقة المقصودة. سينيكا نصح بتخصيص أيام للعيش بأقل الإمكانيات حتى تزول الهواجس من الحرمان. الصيام المطوّل والتدريب المكثف والتعلم المتسارع ليست عقوبات بل أدوات تقسية تُعدّ الإنسان لما هو أصعب ويُطلق طاقة كانت مكبوتة خلف جدار الراحة.
#### Content
نجوت من برنامج الماجستير في فرنسا وأنا شبه مفلس. آكل بتقشف وأعيش ببساطة. بدلاً من النظر لهذا كمشقة، شعرت بالرضا والتركيز على النتيجة النهائية — الحصول على وظيفة هندسة برمجيات بأجر جيد في Squla بعد التخرج.
## الانزعاج كمحفز
الانزعاج والمحن، وليس الراحة، هي ما يحفز النمو الشخصي ويدفع الإنسان لتجاوز حدوده المتصورة.
أنادي بخلق مشقات محتملة عمداً كممارسة: صيام مطول، تمارين مكثفة، تعلم متسارع.
## سينيكا عرف هذا
نصح سينيكا:
> "خصص عدداً معيناً من الأيام... بأقل الطعام... قائلاً لنفسك... 'أهذا هو الحال الذي كنت أخشاه؟'"
الجنود يتدربون في وقت السلم للاستعداد للصراع الحقيقي. وبالمثل، يجب أن يقسّي الإنسان نفسه استباقياً عبر تحديات مفروضة ذاتياً.
### الحرية صعبة
- URL: https://mohammadshaker.com/ar/blog/freedom-is-hard
- Date: 2025-05-21T00:00:00.000Z
اخترت لندن وطناً لأنها تمثل أفضل بيئة في أوروبا للشركات الناشئة والمخاطرة. وصلت إليها بجواز سفر سوري فقط بعد محطات في فرنسا وأمستردام. تأشيرة الموهبة الاستثنائية البريطانية منحتني حرية العمل دون قيود، وهي الحرية التي أضعتُ سنوات أسعى إليها وأعتبرها اليوم أثمن ما أملكه.
#### Content
اخترت لندن كوطني. دعوني أتتبع رحلتي من دمشق، سوريا عبر فرنسا وأمستردام قبل الاستقرار في المملكة المتحدة عام 2018.
## لماذا لندن؟
اخترت لندن — إلى جانب برلين — لأنها تمثل **أفضل مكان في أوروبا للشركات الناشئة والمخاطرة**.
## الجدول الزمني
- **2013**: أنهيت شهادة خمس سنوات في تكنولوجيا المعلومات والذكاء الاصطناعي في سوريا
- **2014–2015**: حصلت على ماجستير في الحوسبة الشاملة والذكاء الاصطناعي في فرنسا
- **2018**: عملت في شركة Squla الهولندية للتكنولوجيا التعليمية في أمستردام
- **2018–الآن**: انتقلت إلى المملكة المتحدة بجواز سفر سوري فقط
## عامل الحرية
الميزة الحاسمة كانت تأشيرة الموهبة الاستثنائية في المملكة المتحدة (الآن تأشيرة الموهبة العالمية)، التي تمنح إقامة لخمس سنوات دون قيود على العمل.
استفدت من هذه الحرية عبر أدوار متعددة: موظف في Neurofenix، مؤسس Almeta (لاحقاً Alphazed)، رئيس الهندسة في Noon، والمؤسس المشارك التقني لـ SpatialX.
**الحرية ثمينة وصعبة المنال؛ يجب أن تُعتز بها فوق كل شيء آخر.**
### "المتطلب" كلمة فارغة.
- URL: https://mohammadshaker.com/ar/blog/requirement-is-an-empty-word
- Date: 2025-04-17T00:00:00.000Z
مصطلح «المتطلب» في هندسة البرمجيات يحمل يقيناً زائفاً: حين يُعلَن أن شيئاً «متطلب»، تُغلق مساحة الحلول دون نقاش حقيقي. البديل الأجدى هو قصص المستخدم والقيود الواضحة، وهي لغة تبقي الخيارات مفتوحة وتدعو الفريق للتفكير بدلاً من الامتثال.
#### Content
مصطلح "المتطلب" مستخدم بإفراط وغامض في سياقات هندسة البرمجيات.
عندما يعلن شخص ما أن شيئاً "متطلب" لحل مشكلة، فإنه يخلق وهم الوضوح بينما يضيّق مساحة الحلول بلا داعٍ فعلياً.
## مشكلة "المتطلب"
تسمية شيء ما "متطلب" تحمل وزناً ضمنياً. تشير إلى اليقين وعدم القابلية للتفاوض — خاصة عندما ينطقها أعضاء الفريق الكبار، الذين قد يعاملون المتطلبات كأوامر لا ينبغي التشكيك فيها.
## ما البديل
أنادي باستبدال لغة المتطلبات بـ**قصص المستخدم والسرديات**. هذا النهج يبقي مساحة الحلول مفتوحة ومرنة.
فريقي في SpatialX يطبق هذه الممارسة عبر جميع تذاكر المنتج والتقنية في Asana.
## الخلاصة
استخدم "متطلب" بتحفظ وفقط عندما تكون متأكداً حقاً أن شيئاً ما ضروري فعلاً. وإلا، تبنَّ لغة تعاونية أكثر تدعو للنقاش والنهج البديلة.
### ما هي تقنيات معالجة اللغات الطبيعية
- URL: https://mohammadshaker.com/ar/blog/ما-هي-تقنيات-معالجة-اللغات-الطبيعية
- Date: 2019-12-14T00:00:00.000Z
معالجة اللغات الطبيعية هي فرع من الذكاء الاصطناعي يُمكّن الحواسيب من فهم اللغة البشرية وتحليلها. تُستخدم في محركات البحث وروبوتات الدردشة وتصفية البريد المزعج وتحليل المشاعر. ما يجعلها صعبة هو الغموض الطبيعي في اللغة: كلمة واحدة قد تحمل معاني متعارضة تبعاً للسياق.
#### Content
قد لا يكون لديك الاطلاع الكافي على معالجة اللغات الطبيعية لكنك بالطبع تعرف كل من سيري أو أليكسا!
"لم أفهم ما قلته للتو." هذا ما يمكن أن تجيبك به [سيري](https://ar.wikipedia.org/wiki/%D8%B3%D9%8A%D8%B1%D9%8A) أو أليكسا مراراً وتكراراً.
متى كانت آخر مرة طلبت فيها من سيري أو أليكسا أن تفعل شيئاً ولم تفهما ما تقوله؟ أو أجابتا بشيء لا علاقة له على الإطلاق بسؤالك؟
سيري وأليكسا هي روبوتات دردشة التي تعتمد بشكل أساسي على تقنية الذكاء الاصطناعي تسمى تقنيات معالجة اللغات الطبيعية (Natural Language Processing - NLP).
إذا كنت ترغب في معرفة المزيد حول تقنيات معالجة اللغات الطبيعية (NLP) وما الذي يمكن أو لا يمكن تحقيقه بها واصل قراءة هذا المقال.
**ما هي معالجة اللغات الطبيعية ؟**
هي فرع من فروع علوم الكمبيوتر والذكاء الصنعي التي تهتم بمجال فهم أجهزة الكمبيوتر للغات الطبيعية البشرية، وذلك من خلال تحليل كميات هائلة من البيانات المستخرجة من اللغة الطبيعية البشرية. تتراوح المشاكل التي يمكن لمعالجة اللغات الطبيعية حلها من مشاكل بسيطة مثل الإجابة على استفسار على شبكة الإنترنت إلى مشاكل معقدة للغاية تتطلب عدة تيرابايت من البيانات للتدريب.
**أين تستخدم معالجة اللغات الطبيعية ؟**
تستخدم معالجة اللغات الطبيعية في جميع البرامج التي تحتاج إلى تحليل البيانات النصية أو الصوتية، نذكر منها على سبيل المثال:
1. محركات البحث: مثل جوجل و ياهو وغيرها. مثلاً عندما تبحث عن كتاب معين فإن محرك البحث سوف يظهر لك الكتاب بالإضافة إلى كتب أخرى تشبهه بالمحتوى أو العنوان.
2. تطوير الشبكات الاجتماعية: على سبيل المثال، إذا كنت تحب صفحات لها علاقة بتربية الحيوانات، فسيتم عرض الإعلانات والمشاركات ذات الصلة بتربية الحيوانات.
3. روبوتات الدردشة: مثل Apple's Siri التي تسألها دائماً على جهازك المحمول.
4. برامج التدقيق الإملائي.
5. فرز رسائل البريد الإلكتروني المزعجة وغير الآمنة.
**ماهي مشاكل تقنيات معالجة اللغات الطبيعية ؟**
تعتبر معالجة اللغة الطبيعية مشكلة صعبة في علوم الكمبيوتر، الأمر الذي يجعلها صعبة هو طبيعة اللغة البشرية.
فهم القواعد التي تعتمد عليها اللغات ومعاني الكلمات ليس سهلاً على أجهزة الكمبيوتر.
بعض الجمل يمكن أن تكون صعبة ومجردة؛ على سبيل المثال:
عندما يقول شخص "لقد كان شعوراً لم يسبق لي أن شعرت به من قبل" هذا يعني أن الشخص قد عانى من شعور جيد للغاية أو سيء جداً، معنى هذه الجملة يعتمد على عواطف الشخص في تلك اللحظة.
قد يهمك: أكبر أربع مشاكل مفتوحة في معالجة اللغات الطبيعية
من ناحية أخرى، يمكن أن تكون بعض هذه الجمل بسيطة، على سبيل المثال:
إذا كتب أحد المستخدمين على روبوت الدردشة (chatbot) "هل ستمطر اليوم في أمسردام؟" ، فسيكون من الصعب تحديد أمستردام كموقع. ولكن يمكن تصحيح الأخطاء الإملائية في الجملة أولاً ليتمكن الحاسوب من فهم الجملة.
هذا بالإضافة إلى الغموض الموجود في اللغات أي أنه يمكن لكلمة واحدة أن تحتمل عدة معاني.
فمثلاً (صدام العلم والدين.) معنى كلمة "والدين" هنا هو بمعنى الديانة وهو يختلف كلياً عن معناها بجملة (والدَين وفائه صعب) وتأتي هنا بمعنى مقدار من المال.
يتطلب فهم اللغة البشرية بشكل شامل فهم كل من الكلمات وكيفية ارتباط المفاهيم لتقديم الرسالة المقصودة.
في حين أن البشر يمكنهم إتقان اللغة بسهولة، إلا أن الغموض وخصائص اللغات الطبيعية هي التي تجعل معالجة اللغة صعبة على الآلات.
نحن في [الميتا](https://play.google.com/store/apps/details?id=io.almeta.almetanewsapp&hl=ar_AR) نستخدم معالجة اللغات الطبيعية لتحليل النصوص وتلخيصها، وفي نهاية هذا المقال أتمنى أن أكون قد أفدتكم بهذه المقدمة الصغيرة عن معالجة اللغات الطبيعية.
هل تعلم أننا نستخدم تقنيات الذكاء الاصطناعي في تطبيقنا؟ انظر إلى أبرز تقنيات الذكاء الصنعي الآن قيد التنفيذ. جرب تطبيق الميتا للأخبار. يمكنك تنزيله من متجر [Google Play](https://play.google.com/store/apps/details?id=io.almeta.almetanewsapp&hl=ar_AR) أو متجر تطبيقات [Apple](https://apps.apple.com/app/id1497390441).
اقرأ أيضاً: ما هو الذكاء الصنعي والتعلم الآلي وما علاقتهما ببعضهما
### أكبر أربع مشاكل مفتوحة في معالجة اللغات الطبيعية
- URL: https://mohammadshaker.com/ar/blog/أكبر-أربع-مشاكل-مفتوحة-في-معالجة-اللغات-الطبيعية
- Date: 2019-09-06T00:00:00.000Z
أربع مشكلات لا تزال تُعيق أنظمة معالجة اللغات الطبيعية حتى اليوم: الغموض الدلالي حين تحتمل الكلمة معاني متعددة، وشُح بيانات التدريب لكثير من اللغات، وصعوبة تصحيح الأخطاء الإملائية واستخراج الأسماء، وأخيراً استخراج المعاني الضمنية من النصوص. هذه المشاكل هي ما يجعل سيري وأليكسا تفشلان أحياناً في فهمك.
#### Content
قبل أن نتحدث عن مشاكل معالجة اللغات الطبيعية دعونا نبدأ بمثال معروف للجميع.
متى كانت آخر مرة طلبت فيها من سيري أو أليكسا أن تفعل شيئًا ولم تفهما ما تقوله؟ أو أجابتا بشيء لا علاقة له على الإطلاق بسؤالك؟
سيري وأليكسا هي روبوتات الكلام التي تعتمد بشكل أساسي على تقنية الذكاء الاصطناعي تسمى NLP. إذا كنت ترغب في معرفة المزيد حول معالجة اللغات الطبيعية (NLP) وما الذي يمكن أو لا يمكن تحقيقه بها واصل قراءة هذا المقال.
(Natural language processing) NLP تعني معالجة اللغات الطبيعية والتي تُعرف بأنها فرع من فروع علوم الكمبيوتر والذكاء الصنعي التي تهتم بمجال فهم أجهزة الكمبيوتر للغات الطبيعية البشرية وذلك من خلال تحليل كميات هائلة من البيانات المستخرجة من اللغة الطبيعية البشرية.
تتراوح مشاكل معالجة اللغات الطبيعية من مشاكل بسيطة مثل الإجابة على استفسار على شبكة الإنترنت إلى مشاكل معقدة للغاية تتطلب عدة تيرابايت من البيانات للتدريب، ولكن إلى أي مدى يمكن أن باستخدام معالجة اللغات الطبيعية فهم ما يقوله البشر؟ وما المدة التي سنستغرقها بالبحث والتدريب حتى نجري محادثة طبيعية مع جهاز كمبيوتر؟
سنناقش في هذه المقالة أربعة من أكثر مشكلات معالجة اللغات الطبيعية صعوبة.
## 1. الغموض في اللغات الطبيعية
في اللغة الطبيعية ، يمكن أن يكون للكلمة معاني مختلفة ويمكن استخلاص معنى الكلمة من سياق الجملة. على سبيل المثال، قد تعني الجملة "أَحْسِنْ إلى الناس تستعبد قلوبهم" أننا نتحدث عن الاستعباد المأخوذ من العبودية للانسان وهي تعطي معنى سيئ للجملة، ومن ناحية أخرى، قد تأخذ معنى ايجابي وهو أنك اذا عاملت الناس بشكل حسن أحبوك.
لا يستخدم البشر معرفتهم باللغة فقط لتحديد معنى النص، لكنهم يفكرون أيضًا في عدة عوامل أخرى تساعدهم مثل الرغبات والأهداف والمعتقدات لفهم النص الذي يقرؤونه أو الكلام يستمعون إليه. على سبيل المثال، قد تعني الجملة "لقد كان شعوراً لم يسبق لي أن شعرت به من قبل" أن الشخص قد عانى من شعور جيد للغاية أو سيء جدًا، معنى هذه الجملة يعتمد على عواطف الشخص في تلك اللحظة.
## 2. عدم وجود بيانات للتدريب
أحد أكبر التحديات في معالجة اللغات الطبيعية NLP هو نقص بيانات التدريب حيث يجب تدريب كل نموذج من نماذج ال NLP على تيرابايت من البيانات حتى يتمكن النموذج من فهم لغة معينة، التدريب النموذج موضوع معقد سيتم تغطيته في مقال منفصل آخر .
إن نقص البيانات التدريبية له عدة أسباب: السبب الأول هو أن اللغة هي من لغات الأقليات العرقية مما يعني أن عدداً قليلاً من سكان الأرض يتحدث بها مثل الكردية والأفريكانية. السبب الثاني هو قلة الموارد والنصوص المتوفرة على الويب، على سبيل المثال، لغة الزولو.
سبب آخر لعدم وجود بيانات التدريب هو أن الحافز للعمل على اللغة إما بسبب عدم توفر المهارات المناسبة أو صعوبة اللغة كما هو الحال في اللغة العربية.
## 3. الأخطاء الإملائية واستخراج الاسم
يعد تصحيح الكلمات التي بها أخطاء إملائية عملية أساسية في معالجة اللغات الطبيعية NLP ، حيث أن الأخطاء الإملائية شائعة جدًا عند استخدام الإنسان للحاسوب وسيكون من الصعب جدًا تحديد الاسم في الجملة من نص معين. على سبيل المثال: إذا كتب أحد المستخدمين على روبوت الدردشة (chatbot) "هل ستمطر اليوم في أميستدام؟" ، فسيكون من الصعب تحديد أمستردام كموقع.
## 4. استخراج المعاني الدلالية (يمكن أن يعتبرهذا جزءًا من غموض اللغات الطبيعية)
يجب أن لا يفهم الكمبيوتر مفردات النص فحسب، بل يجب أن يفهم أيضًا دلالات النص. على سبيل المثال: في الجملة "اتصل جون بزوجته ، وكذلك فعل سام" ، لا نعرف ما إذا كان سام قد اتصل بزوجته جون أم اتصل بزوجته.
هل تعلم أننا نستخدم كل هذا وتقنيات الذكاء الاصطناعي الأخرى في تطبيقنا؟ انظر إلى ما تقرأه الآن قيد التنفيذ. جرب تطبيق الميتا للأخبار. يمكنك تنزيله من متجر [Google Play](https://play.google.com/store/apps/details?id=io.almeta.almetanewsapp&hl=ar_AR) أو متجر تطبيقات [Apple](https://apps.apple.com/app/id1497390441).
## المراجع:
1-[What](https://medium.com/datadriveninvestor/what-are-some-of-the-challenges-we-face-in-nlp-today-2e9d94da1f63) [are some of the challenges we face in NLP today?](https://medium.com/datadriveninvestor/what-are-some-of-the-challenges-we-face-in-nlp-today-2e9d94da1f63)
2- [The 4 Biggest Open Problems in NLP](http://ruder.io/4-biggest-open-problems-in-nlp/)
3- [Six challenges in NLP and NLU - and how boost.ai solves them](https://www.boost.ai/articles/six-challenges-in-nlp-and-nlu-and-how-boostai-solves-them)
### أكبر التحديات في معالجة اللغة العربية
- URL: https://mohammadshaker.com/ar/blog/أكبر-التحديات-في-معالجة-اللغة-العربية
- Date: 2019-09-05T00:00:00.000Z
معالجة اللغة العربية آليًا أصعب بكثير من معظم اللغات لثلاثة أسباب: تتغير أشكال الحروف بحسب موضعها في الكلمة، والصرف العربي يُنتج مئات الأشكال من جذر واحد، فضلاً عن الهوّة العميقة بين اللهجات المحكية والفصحى المكتوبة. هذه العوامل تُعقّد مهام النماذج من تحليل المشاعر إلى الترجمة الآلية.
#### Content
قبل أن نبدأ بالحديث عن معالجة اللغة العربية ، ذكرنا في المدونات السابقة أهمية معالجة اللغات الطبيعية ومجموعة التطبيقات الواسعة التي يتم فيها استخدام معالجة اللغات الطبيعية.
نظرًا لأن الهدف من معالجة اللغات الطبيعية (NLP) هو تسهيل وتبسيط التواصل بين الآلات والبشر، فمن المهم جداً أن نرى كيف سيؤثر ذلك على حياة الأشخاص الذين يتحدثون ويتواصلون ويعملون مع اللغة التي تأتي بالمرتبة السادسة لأكثر اللغات تحدثًا في العالم، اللغة العربية.
اللغة العربية هي لغة سامية يتحدث بها حوالي 420 مليون شخص في العالم ، بالإضافة إلى ذلك، تعد اللغة العربية هي الغة الرسمية في 26 دولة وهي واحدة من اللغات الرسمية للأمم المتحدة.
اللغة العربية غنية من الناحية المورفولوجية وتتكون من عدة أنواع على سبيل المثال، هناك اللغة العربية الكلاسيكية وهي لغة القرآن الكريم (الكتاب المقدس للمسلمين) والتي تعتبر الشكل الأكثر مثالية للغة العربية، وهناك نوع آخر هو اللغة العربية الفصحى الحديثة.
وهي اللغة الرسمية اليوم والمستخدمة في الأدب والتعليم والكتب ووسائل الإعلام وغيرها من المواقع والمواقف الرسمية وأخيراً هناك اللهجات العربية التي تعتبر اللغة المحكية اليومية وهي مختلفة في كل بلد.
بعد هذه المقدمة القصيرة السابقة عن اللغة العربية، سنناقش في هذه المقالة ثلاثة من أهم القضايا في معالجة اللغة العربية.
## 1. الهجاء العربي (Arabic orthography)
تتكون أبجدية اللغة العربية من 28 حرفًا، وتحتوي فقط على ثلاثة أحرف علَّة (ا)، (و)، (ي). بالإضافة إلى تسعة محارف أخرى وهي التنوين (َ ُُ ِِ ً ٌ ٍ ّ ْ). اللغة العربية هي أيضًا إحدى اللغات التي يمكن أن يتغير شكل الأحرف وفقًا لكيفية ارتباطها بالحروف الأخرى.
على سبيل المثال ، يحتوي حرف التاء (ت) على ثلاثة أشكال من الكتابة: يتم كتابته كـ (ت) إذا كان موجوداً في نهاية الكلمة، (  ) إذا كان موجوداً في منتصف الكلمة و (  ) إذا كان موجودا ً في بداية الكلمة. تهجأة الأحرف في اللغة العربية مهم جدًا في جميع مهام وتطبيقات معالجة اللغات الطبيعية، مثل: تقسيم الكلمات والجمل وتحويل النص إلى كلام.
## 2. مورفولوجيا اللغة العربية
جميع الأفعال في اللغة العربية لها جذر من ثلاثة أو أربعة أحرف مما يجعل اللغة العربية لغة صعبة للغاية. عادةً، هناك قالب لاشتقاق الأفعال ويمكننا معرفة الفعل الجديد وفقاً لل معادلة التالية الفعل= الجذر+ النمط. يعرض الجدول التالي بعض أمثلة الأفعال في ثلاث أزمنة الماضي والحاضر والمستقبل وجذورها مستمدة من أصل ثلاثي أو رباعي.
| الجذر | النمط | الفعل | اللفظ | المعنى |
| --- | --- | --- | --- | --- |
| كتب | ي | ي+كتب=يكتب | yaktb | الزمن الحالي والمضارع من كتب |
| كتب | ا | ا+كتب=اكتب | Ektb | الفعل الأمر من كتب |
| احضر | ي | يحضر | yhder | الزمن الحالي والمضارع من احضر |
من الشائع جدًا أيضًا باللغة العربية إرفاق البادئات واللواحق بالأفعال، ويمكننا صياغة ذلك باستخدام المعادلة التالية الفعل الجديد = السوابق + الفعل + اللواحق. يوضح الجدول التالي مثالاً على التصريف باللغة العربية.
| الفعل | الفعل الجديد | المعنى |
| --- | --- | --- |
| يكتب | س + يكتب = سيكتب | سوف يكتب |
| يكتب | س + يكتب + ه = سيكتبه | هو سوف يكتبه |
دراسة مورفولوجيا اللغة العربية مهمة جداً لمهام معالجة اللغات الطبيعية مثل التحليل الصرفي وتنميط POS (Part Of Speech tagging).
## 3. البناء المعقد للجملة
اللغة العربية غنية بالمفردات حيث يمكن أن يكون لكل كلمة عدة معانِ. على سبيل المثال، "البيت كبير" كلمة "كبير" يمكن أن تعطي الجملة معنى مختلفة في حال قلنا "كبير القوم" مما يعني (الرجل المسؤول عن مجموعة من الأشخاص). سوف تؤثر مشكلة وجود معاني متعددة للكلمات في اللغة العربية على تطبيقات مثل تلخيص النصوص والترجمة.
هل تعلم أننا نستخدم تقنيات الذكاء الاصطناعي في تطبيقنا؟ انظر إلى أبرز تقنيات الذكاء الصنعي الآن قيد التنفيذ. جرب تطبيق الميتا للأخبار. يمكنك تنزيله من متجر [Google Play](https://play.google.com/store/apps/details?id=io.almeta.almetanewsapp&hl=ar_AR) أو متجر تطبيقات [Apple](https://apps.apple.com/app/id1497390441).
قد يهمك أيضاً: ما هو الذكاء الصنعي والتعلم الآلي وما علاقتهما ببعضهما
## المراجع:
[Challenges in Arabic Natural Language Processing](https://www.researchgate.net/publication/327753798_Challenges_in_Arabic_Natural_Language_Processing)
### الأفكار الثلاث الأكثر إثارة في معالجة اللغات الطبيعية (NLP)
- URL: https://mohammadshaker.com/ar/blog/الأفكار-الثلاث-الأكثر-إثارة-في-معالجة-اللغات-الطبيعية-nlp
- Date: 2019-09-03T00:00:00.000Z
ثلاثة اختراقات غيّرت مسار معالجة اللغات الطبيعية عام 2018: BERT الذي درّب النماذج من الاتجاهين بدلاً من اتجاه واحد، ومعيار SWAG الذي يقيس الاستدلال بالمعرفة العامة عبر 113 ألف سؤال، ونموذج LISA الذي يستخرج الأدوار الدلالية دون معالجة مسبقة مكثفة. كل منها يعالج فجوة حقيقية بين فهم الإنسان وفهم الآلة.
#### Content
"إن معالجة اللغات الطبيعية و التعلم الآلي هما الأساس لأي نظام من الذكاء الصنعي، حيث تكمن أهميتهم في القدرة على التواصل معنا بطريقة إنسانية وأتمتة عملية التعلم، بغض النظر عماتريد الوصول اليه سواءً كان تحليلات تنبؤية أو إرشادية، تنبؤ، تحسين النموذج، أينما تذهب، يعود الأساس دائماً إلى هذه التقنيات التي كانت موجودة منذ عقود. "
هذا الاقتباس من قبل ماري بيت مور خبيرة SAS الذكاء الصنعي وتحليل اللغة الاستراتيجي. وبعد ذكر أهمية التعلم الآلي ومعالجة اللغات الطبيعية، يجب أن نذكر أن العديد من الأبحاث الجديدة في عام 2018 قد حققت تحسناً مذهلاً في معالجة اللغات الطبيعية. في هذه المقالة سوف نلقي نظرة على أفضل ثلاث نماذج متطورة لمعالجة اللغات وطرق جديدة في معالجة اللغات الطبيعية.
1. BERT
يعتبر BERT اختصارًا (Bidirectional Encoder Representations from Transformers) ممثلي التشفير ثنائي الاتجاه من المحولات ، وهو نموذج جديد متطور تم تدريبه مسبقًا على معالجة اللغات الطبيعية (NLP) وأعطى نتائج مبهرة جديدة في حل مهام معالجة اللغات الطبيعية مثل الإجابة عن الأسئلة والتعرف على الكيان المسمى والاستدلال على اللغة. على عكس نماذج معالجة اللغات الطبيعية الأخرى المدربة مسبقًا مثل OpenAI GPT و ELMo ، تم تصميم BERT بمحول ثنائي الاتجاه للتدريب على كل كلمة من الجانبين من اليسار واليمين. يوضح الشكل التالي الفرق بين النماذج الثلاثة.

يتجنب نموذج BERT ثنائي الاتجاه مسألة الحلقات التي يمكن فيها تكرار الكلمات لأنه يتم تدريبه عن طريق إخفاء نسبة مئوية من الدلالات المدخلة (input tokens) بشكل عشوائي، كما يمكنه أيضًا فهم العلاقات بين الجمل عن طريق التدريب المسبق لنموذج يحدد العلاقات بين الجمل. يعتبر بيرت حقبة جديدة في معالجة اللغات الطبيعية ويمكن استخدامه في تطبيقات الدردشة وتحليل ملاحظات العملاء.
2. SWANG
عندما يقرأ شخص ما جملة "أحمد ارتدى ملابسه" ، يكون قادرًا على توقع بقية الجملة التي قد تكون "و رحل". بخلاف البشر، لا تستطيع الآلات تكملة هذه الجملة الواضحة والسهلة لأنها تتطلب تفكيراً منطقياً. SWANG اختصار لـ (Situations With Adversarial Generations) هدفه هو تعزيز مجال البحث في استنتاج اللغة الطبيعية (Natural Language Inference)(NLI). تم تقديم SWANG كمجموعة بيانات واسعة النطاق تحتوي على 113 ألف سؤال حول مجموعة واسعة من المواقف التي تحتاج إلى التفكير المنطقي والتي تم جمعها باستخدام التعليقات التوضيحية على مقاطع فيديو متنوعة. وقد تم بناء مجموعة البيانات باتباع الخطوات التالية:
1. استخراج جملة من تعليق توضيحي للفيديو
2. استخراج الإجابة الصحيحة من تعليق الفيديو التالي
3. توليد إجابات خاطئة عن طريق إنشاء مجموعة كبيرة من الإجابات الخاطئة، واختيار الإجابات الأكثر ارتباطًا باستخدام نماذج الاحصائية وأخيراً تصفية النهايات التي تبدو وكأنها قد تم توليدها بواسطة الحاسوب واستبدال تلك النهايات بنهايات مشابهة لما قد يفعله أو يقوله الإنسان، تسمى هذه الطريقة لتوليد الإجابات Adversarial Filtering(AF).

يوضح الشكل السابق مثالاً على كيفية عمل SWANG. دقة SWANG عالية نسبياً حيث بلغت الدقة 86.2٪ ، بينما دقة الإنسان تصل إلى 88٪. يمكن لهذا النموذج أن يحسن التفكير المنطقي في أنظمة السؤال والجواب وروبوتات الدردشة.
3. LISA
تعد LISA اختصارًا Linguistically-Informed Self-Attention، وهي عبارة عن نموذج لشبكة عصبونية مصمم لاستخراج وصف الدور الدلالي (semantic role labeling) باستخدام التعلم العميق (deep learning) والشكليات اللغوية (linguistic formalism). على سبيل المثال، الجملة "أحمد أعطى التعليمات إلى ليلى" يجب على النموذج أن يتعرف على الفعل "أعطى إلى" باعتباره المسند، "أحمد" كمسؤول عن الاعطاء أو الشخص الذي أعطى التعليمات، "التعليمات" كموضوع و "ليلى" كمتلق للإٍعطاء.
تأخذ الشبكة العصبية الكلمات المضمنة (word embeddings) كمدخلات بالإضافة إلى باراميترات مخصصة لهذه المهمة قد تم تعلمها ويتم التدريب باستخدام تقنية multi-head self-attention مع التعلم متعدد المهام (multi-task learning). على عكس النماذج السابقة لاستخراج وصف الأدوار الدلالية، لا يحتاج نموذج LISA الكثير من المعالجة المسبقة حيث يمكنها إضافة كلمات أو رموز بإدخال الرموز الخام (raw tokens) فقط ثم يتم ترميز التسلسل المدخل وبعد ذلك يتم إجراء التحليل واستخراج وصف الأدوار الدلالية. يستخدم LISA في تطبيقات التلخيص، والترجمة الآلية وأنظمة الأسئلة والأجوبة ، ويعطي نتائج جيدة للغاية في تحليل أنماط الكتابة في الصحف والمجلات والكتابات الخيالية(fictional writing).
هل تعلم أننا نستخدم تقنيات الذكاء الاصطناعي في تطبيقنا؟ انظر إلى أبرز تقنيات الذكاء الصنعي الآن قيد التنفيذ. جرب تطبيق الميتا للأخبار. يمكنك تنزيله من متجر [Google Play](https://play.google.com/store/apps/details?id=io.almeta.almetanewsapp&hl=ar_AR) أو متجر تطبيقات [Apple](https://apps.apple.com/app/id1497390441).
المراجع:
[1- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding](https://arxiv.org/abs/1810.04805)
[2- SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference](https://arxiv.org/abs/1808.05326)
[3- Linguistically-Informed Self-Attention for Semantic Role Labeling](https://arxiv.org/abs/1804.08199)
[4- 14 NLP Research Breakthroughs](https://www.topbots.com/most-important-ai-nlp-research/)
---
## Projects
### CRUST: A 2D/3D Physics Engine (2012-2013)
- URL: https://mohammadshaker.com/en/projects/crust-a-2d-and-3d-physics-engine
A 2D and 3D physics engine implemented in C# with XNA. The engine is a heavily modified version of Millington’s engine and is able to provide efficient...
#### Content
A 2D and 3D physics engine implemented in C# with XNA. The engine is a heavily modified version of Millington’s engine and is able to provide efficient handling for physics simulations, implementing an impulse force collision modelling to deal with rigid objects. Other physics-based motions such as springs, ropes, and hard constraints, water, ragdolls can also be simulated in CRUST. The engine is also facilitated with a friendly user interface that allows editing objects and their physical properties at runtime. A version of Crust was used in implementing a 3D augmented reality environment to interact with physics objects in realtime.
### Game Design and Development (2014-2015)
- URL: https://mohammadshaker.com/en/projects/game-development-and-design
Developed 5 Android games and another 10 web games prototypes using Unity3D, solely.
#### Content
Developed 5 Android games and another 10 web games prototypes using Unity3D, solely. As frameworks: I developed the state-of-the-art Ropossum for procedurally generated physics-based games (Cut the Rope) and Selence for rhythm-based games.
### Ropossum Framework (2014-2015)
- URL: https://mohammadshaker.com/en/projects/ropossum-framework
This project investigates the use of procedural content generation techniques and in particular Grammatical Evolution to generate endless content for the...
#### Content
This project investigates the use of procedural content generation techniques and in particular Grammatical Evolution to generate endless content for the popular physics-based puzzle video game Cut the Rope. The project aims at generating infinite interesting yet playable content. The levels structure is defined in a design grammar employed by grammatical evolution while a set of playability rules is specified in a first-order logic format. The project features Ropossum, a mix-initiative design tool in which a partially designed level by a human (designer or player) can be automatically completed and checked for playability in realtime. Read more.
### Selene Framework (2014-2016)
- URL: https://mohammadshaker.com/en/projects/selene-framework
Selene, is a Dependency Injection (DI), Inversion of Control (IoC) framework I built on top of Unity3D.
#### Content
Selene, is a Dependency Injection (DI), Inversion of Control (IoC) framework I built on top of Unity3D.
Selene is built on the premise that music can act as the basic building block of any game. Selene serves as a framework for building rhythm-based, procedurally generated, music games on top of the Unity3D game engine. Selene was born out of SyncSeven: a very light-weight, music-based game released on Google play on March 2015.Selene V2.0, now a full-fledged framework, powers the new games of TheX, Paper Ski, Flopp and the prototypes of Time Shifts, Spectre, GGBox, Reactive Music Canvas andExcavation.
### SpatialX EXPLORE: Multi-Modal Annotation for Cancer Pathology
- URL: https://mohammadshaker.com/en/projects/spatialx-multimodal-annotation
Cancer is personal. Treatment should be too. As co-founder and CTO of SpatialX, I architected EXPLORE — an AI platform fusing spatial biology with digital pathology under HIPAA and ISO 15189 — and led the patented Multi-Modal Annotation System, with a ProCreate-style annotation UX built for pathologists.
#### Content
Cancer is personal. Treatment should be too.
As co-founder and CTO of SpatialX (2024–2025), I architected EXPLORE — an AI platform that fuses spatial biology with digital pathology to make cancer treatment personal. I built it end-to-end, under HIPAA and ISO 15189, hands-on across frontend and backend.
I led the SpatialX patent from concept to delivery: an AI-powered Multi-Modal Annotation System, plus 5 research papers. The pipeline serves whole-slide images at up to 1M+ image-tiles per minute.
Tech: multimodal AI, spatial biology, whole-slide imaging (WSI), tile-streaming pipelines, AWS, Pulumi/IaC, CircleCI.
UX — An Annotation Tool That Feels Like ProCreate on iPad
Pathologists do not think in form fields and dropdowns — they think with their hands. So I built EXPLORE's annotation surface to feel like ProCreate on iPad: direct, low-latency brush strokes on gigapixel slides, pressure-aware marking, and pan/zoom that never fights the user.
The goal was zero friction between an expert's intent and the label that trains the model. Annotating a tumor region should feel like sketching, not data entry. The recording below shows the intuitive annotation tool in action on a whole-slide image.
SpatialX EXPLORE annotation tool — ProCreate-style direct manipulation on gigapixel whole-slide images.
### Startup/Almeta: Advancing the understanding of the Arabic language.
- URL: https://mohammadshaker.com/en/projects/startup-almeta-advancing-the-understanding-of-the-arabic-language
A fresh look into gamified, quality and smart education in the Arab and MENA region. In 2021, I’ve founded Almeta.io (website) .
#### Content
A fresh look into gamified, quality and smart education in the Arab and MENA region.
In 2021, I’ve founded Almeta.io (website).
Almeta is an AI initiative advancing the understanding of the Arabic language. We developed programmable APIs that can measure bias, neutrality, readability, informativity and other metrics for any Arabic text content on the web, tackling false news and fact-checking first.
Released web and app News Platform having these metrics. And published our technical and research work in more than 150+ machine learning research and technical blogs in Natural Language Processing (NLP) for the Arabic language on our blog for the public. All in English for anyone to use and read: https://www.almeta.io/en/blog/
Tech Stack:- Backend: GCP and AWS - all microservice architecture. Cloud Run, SES, SQS, Docker, Python/Flask, AWS, Lambda, SQLAlchemy, Serverless (sls), DynamoDB, ElasticSearch, Redis, Step Functions, Snorkel, Wikifier, TDD, IaC, CircleCI (CI/CD).- Frontend: Flutter, Dart, CodeMagic (CI/CD)- Analysis and Marketing: Segment, Amplitude, Drip for marketing automation.
Team:Built and led a team of 10 members: 8 engineers, 1 marketing, 1 business.
Brand & marketing
The Almeta News app
### Startup/Alphazed: Gamified EdTech
- URL: https://mohammadshaker.com/en/projects/alphazed-gamified-edtech
Developed the first, smart and automated, lip-syncing technology for the Arabic language. Amal (250K+ students) and Thurayya (100K+ families) — two award-winning EdTech apps for Arabic language and Quran learning.
#### Content
Developed the first, smart and automated, lip-syncing technology for the Arabic language.
A fresh look into gamified, quality and smart education in the Arab and MENA region.
Between 2020 and 2022, I’ve founded Alphazed (website): the first AI-led, language-agnostic platform for transforming any curriculum into its gamified digital twin, end-to-end.We focused on the Arab and MENA region as green, unsatisfied market.
We reached 250,000+ students across two products with a 0 marketing budget.
Tech Stack: Python, Flask, AWS, Lambda, SQLAlchemy, Serverless (sls), IaC with serverless and CloudFormation, CircleCI (CI/CD). Flutter, Dart, CodeMagic (CI/CD). Segment, Amplitude, Drip for marketing automation.Team: Hired, built and led a team of 19 members: 12 engineers, 2 marketing, 2 business, 2 content.
Amal — Complete Arabic Language Learning System for Kids
أمل — نظام تعلّم اللغة العربية الكامل للأطفال
Amal is the first Arabic language learning app for children to integrate speech therapy, AR face filters, and AI-powered pronunciation recognition — all in one gamified system. Winner of the Seedstars World Award for Best Arabic Platform.
أمل هو أول تطبيق عربي لتعلّم اللغة يدمج العلاج النطقي وفلاتر الواقع المعزز وتقنية التعرف على النطق بالذكاء الاصطناعي — كل ذلك في نظام تعليمي متكامل وممتع. حصل على جائزة Seedstars العالمية لأفضل منصة عربية.
Thurayya — Full Quran Learning System for Kids
ثريا — نظام تعلّم القرآن الكريم الكامل للأطفال
Thurayya is the first Islamic app with lipsyncing AI technology, purpose-built for teaching children Quran recitation and memorization. Features voice recognition, adjustable recitation speed, full Quran coverage, and a gamified memorization system. 100,000+ happy families since launch.
ثريا هو أول تطبيق إسلامي يستخدم تقنية مزامنة الشفاه بالذكاء الاصطناعي، مصمم لتعليم الأطفال تلاوة القرآن الكريم وحفظه. يتميز بالتعرف على الصوت وضبط سرعة التلاوة وتغذية القرآن الكامل ونظام حفظ ممتع. أكثر من 100,000 عائلة سعيدة منذ الإطلاق.
Raw Gameplay Footage
لقطات اللعب الحقيقية
للقراء العرب
لقطات شاشة التطبيقين باللغة العربية
---
## Games
### Collapse (2014)
- URL: https://mohammadshaker.com/en/games/collapse
Co[l]apse is a very, lightweight, prototype of an abstract swipe game, gray all the way!
#### Content
Co[l]apse is a very, lightweight, prototype of an abstract swipe game, gray all the way!
### Flopp on Android (2015)
- URL: https://mohammadshaker.com/en/games/flopp-on-android
[youtube https://www.youtube.com/watch?v=t0skM_gsyXY&w=730&h=443] Flopp is a game syncs your perception of movement, color, and music with your conceptual...
#### Content
[youtube https://www.youtube.com/watch?v=t0skM_gsyXY&w=730&h=443]
Flopp is a game syncs your perception of movement, color, and music with your conceptual understanding of how marriage work, how people interact and how relationships evolve.This game is a fusion between music, abstract art, and psychology. It's crazy simple though. You only have to tap on the screen once objects pump, move or color. You can start to feel the link between the previous domains once you get the hang of the game. With time.Grab it now for Free!
### GGBox (2015)
- URL: https://mohammadshaker.com/en/games/ggbox
GGBox , is a rhythm-based, procedurally generated game on top of the new Selene 2.0 framework.
#### Content
GGBox, is a rhythm-based, procedurally generated game on top of the new Selene 2.0framework. Everything: enemies, items, blocks, platforms and boxes are generated from the music itself.
[youtube https://www.youtube.com/watch?v=s-m2sNUwqJM&w=555&h=343]
### Inversion (2014)
- URL: https://mohammadshaker.com/en/games/inversion
Like NEXT, developed with Unity3d, iNversion features a 3D orthographic look and is developed for Android, iOS and the web.
#### Content
Like NEXT, developed with Unity3d, iNversion features a 3D orthographic look and is developed for Android, iOS and the web. It features an iNverse gameplay mechanics to NEXT with a fresh look and design making it a whole new gameplay and thinking experience. Like NEXT, iNversion uses JANUS, a level generator, to generate endless playable levels with specified difficulty without any human intervention. To be out soon.
### NEXT (2014)
- URL: https://mohammadshaker.com/en/games/next
Developed with Unity3d, NEXT features a 3D orthographic look and is developed for Android, iOS and the web, NEXT is a unique puzzle game where its game...
#### Content
Developed with Unity3d, NEXT features a 3D orthographic look and is developed for Android, iOS and the web, NEXT is a unique puzzle game where its game generator, JANUS, generates endless playable levels with specified difficulty without any human intervention. Coming on mobile.
### Paper Ski (2015)
- URL: https://mohammadshaker.com/en/games/paper-ski-on-android
[youtube https://www.youtube.com/watch?v=QmtamMOuqak&w=730&h=443] Paper Ski is a rhythm-based game about skiiiing in a crazy environment of rocks and...
#### Content
[youtube https://www.youtube.com/watch?v=QmtamMOuqak&w=730&h=443]
Paper Ski is a rhythm-based game about skiiiing in a crazy environment of rocks and trees and some animals (just two to be exact if the algorithm decided to show them for you.) If you love the experience of skiiiing upside down or at night, [or when it's snowing] or in the sahara, or with hostile rocks AND you love music, then, this game is for you. Grab it now for Free!
### Project Mom (2023)
- URL: https://mohammadshaker.com/en/games/project-mom
A quiet, hand-drawn narrative game about memory, distance and family — designed and coded end-to-end in Unity. A complete prototype, not yet launched.
#### Content
A quiet, hand-drawn narrative game about memory, distance and family.
Project Mom is a personal, monochrome narrative game I designed and built end-to-end in Unity — art direction, writing, interaction and code. It is a slow, memoir-style piece told through line-drawn hands and short fragments of text: Studied. Travelled. Injured my knees. I visited her.
Role: solo — design, narrative, illustration and engineering.Engine: Unity (C#).Status: complete prototype, not yet launched.
### Reactive Music Canvas (2015)
- URL: https://mohammadshaker.com/en/games/reactive-music-canvas
Music Visualisation. A "breathing" concrete. A Reactive Music Canvas. Call it what you want.
#### Content
Music Visualisation. A "breathing" concrete. A Reactive Music Canvas. Call it what you want. The nice thing about it is that it's coded in a couple of hours and I love the result and how it "breathes" music.
[youtube https://www.youtube.com/watch?v=17Bs4xFZq2c&w=555&h=343]
### Spectre (2015)
- URL: https://mohammadshaker.com/en/games/spectre
Spectre , is a rhythm-based, procedurally generated, shoot'em up all game on top of the new Selene 2.0 framework.
#### Content
Spectre, is a rhythm-based, procedurally generated, shoot'em up all game on top of the new Selene 2.0 framework. Enjoy shooting enemies generated from the music in a rhythm-based fashion.
### SyncSeven on Android (2015)
- URL: https://mohammadshaker.com/en/games/syncseven-on-android
Developed with Unity3d and available for free on Android , SyncSeven is a game all about enchanting your visual perception and musical taste! The game is...
#### Content
Developed with Unity3d and available for free on Android, SyncSeven is a game all about enchanting your visual perception and musical taste! The game is procedurally generated, taking music as the main source. A level is generated according to its music genre, type, and how intense or relaxed it is. You only have to sync your taps with turns generated according to the peaks in the music. Try to feel the music to anticipate when the next turn will happen. The more rapid the music the faster you should be. Collect as much correct turns to boost your score.
### TheX on Android (2015)
- URL: https://mohammadshaker.com/en/games/thex-on-android
[youtube https://www.youtube.com/watch?v=P3zUkfEpcCo&w=730&h=443] TheX is a minimalistic, swipe, 2D puzzler for color enthusiasts and brain kings.
#### Content
[youtube https://www.youtube.com/watch?v=P3zUkfEpcCo&w=730&h=443]
TheX is a minimalistic, swipe, 2D puzzler for color enthusiasts and brain kings.Let your brain juice drops in colors! Grab it now!
### Time Shifts (2015)
- URL: https://mohammadshaker.com/en/games/time-shifts
GGBox: Time Shifts , is an evolution of the GGBox prototype with time shifting on top of the new Selene 2.0 framework.
#### Content
GGBox: Time Shifts, is an evolution of the GGBox prototype with time shifting on top of the new Selene 2.0 framework. It's a rhythm-based, procedurally generated, survival, against the time game. Everything: enemies, items, blocks, platforms are generated from the music itself.
[youtube https://www.youtube.com/watch?v=BQtyTa2YOig&w=555&h=343]
---
## Publications
URL: https://mohammadshaker.com/en/publications
#### Content
Master Thesis
Mohammad Shaker. Novel Approaches for Real-time Assessment and Automatic Generation of Game Content. UJF, Grenoble, France, 2015. Related Games on Android: TheX, Flopp, PaperSki and SyncSeven.
Peer-review Journal Papers
2017
Mohammad Shaker, Mehdi Zonjy, Mhd Hasan Sarhan, Ismaeel Abuabdallah and Noor Shaker. Personalizing Content Generation in First Person Shooter Games through Player Modeling. To appear.
Peer-review Conference Papers
2015
Noor Shaker, Mohammad Shaker and Mohamed Abou-Zleikha. Towards Game-Independent Models of Player Experience, in Proceedings of Artificial Intelligence and Interactive Digital Entertainment (AIIDE 15), 2015.
Noor Shaker, Mohamed Abou-Zleikha and Mohammad Shaker. Active Learning for Player Modeling, in Proceedings of the 10th International Conference on Foundations of Digital Games, 2015.
Mohammad Shaker, Noor Shaker, Julian Togelius and Mohamed Abou-Zleikha. A Progressive Approach to Content Generation, in Proceedings of EvoGames: Applications of Evolutionary Computation, Lecture Notes on Computer Science, 2015. Nominated for best paper award.
Mohammad Shaker, Noor Shaker, Mohamed Abou-Zleikha and Julian Togelius. A Projection-Based Approach for Real-time Assessment and Playability Check for Physics-Based Games, in Proceedings of EvoGames: Applications of Evolutionary Computation, Lecture Notes on Computer Science, 2015. Nominated for best paper award.
Walaa Baghdadi, Fawzya Shams Eddin, Rawan Al-Omari, Zeina Alhalawani, Mohammad Shaker and Noor Shaker. A Procedural Method for Automatic Level Generation in Spelunky, in Proceedings of EvoGames: Applications of Evolutionary Computation, Lecture Notes on Computer Science, 2015. Nominated for best paper award.
2014
Noor Shaker and Mohammad Shaker. Towards Understanding the Nonverbal Signatures of Engagement in Super Mario Bros, in Proceedings of the 2014 Conference on User Modeling, Adaptation and Persolization (UMAP 2014), 2014.
2013
Mohammad Shaker, Noor Shaker and Julian Togelius. Evolving Playable Content for Cut the Rope through a Simulation-Based Approach, in Proceedings of Artificial Intelligence and Interactive Digital Entertainment (AIIDE 13), 2013. Poster here.
Mohammad Shaker, Noor Shaker and Julian Togelius. Ropossum: An Authoring Tool for Designing, Optimizing and Solving Cut the Rope Levels, in Proceedings of Artificial Intelligence and Interactive Digital Entertainment (AIIDE 13), 2013.
Mohammad Shaker, Mhd Hasan Sarhan, Ola Al Naameh, Noor Shaker and Julian Togelius. Automatic Generation and Analysis of Physics-Based Puzzle Games, in Proceedings of the 2013 IEEE Conference on Computational Intelligence and Games (CIG 2013), 2013. Nominated for best paper award.
Noor Shaker, Mohammad Shaker, Ismaeel Abuabdallah, Mehdi Zonjy, and Mhd Hasan Sarhan. A Quantitative Approach for Modeling and Personalizing Player Experience in First-Person Shooter Games, in the Extended Proceedings of the 2013 Conference on User Modeling, Adaptation and Persolization (UMAP 2013), 2013. Poster here.
---
## Seminars
URL: https://mohammadshaker.com/en/publications#seminars
#### Content
Here you can find links to talks, seminars, presentations and courses I taught and learnt. All my presentations are uploaded to my slideshare page.
Courses I Teach (Up to 2015)
Game Development with Unity3D
Android Programming and Cloud
Interaction Design Crash Course
Mobile Software Engineering Crash Course
C++ Programming
C++.NET Windows Forms
C# Programming
C# Advanced and .NET Techniques
Windows Presentation Foundation [WPF]
OpenGL 3D Graphics
OpenGL Crash Course
XNA Game Development
Intro to Event-driven Programming and Forms with Delphi
Startup and Indie
[2015] Wikilogia Event: The Indie Series for Game Development [1] [2] [3] [4]
[2015] SyncSeven: A Music-based Generated Game
[2014] Startup Weekend Damascus: Weebee on A Mission: a Game that Can Change Your Child Behaviour
Project Talks and Seminars I Gave
[2015] Ropossum V3.0: The Power of the Progrossive Approach and the Effectiveness of Projection-based Approach
[2014] ITU of Copenhagen, Denmark: My Work on Artificial Intelligence and Games
[2014] Weebee on a Mission: A Serious Game for Better Understanding of Behavior Differences Between Children
[2014] Utilizing Kinect Control for a More Immersive Interaction with 3D Environment
[2013] Cut the Rope Play Forever and Ropossum Authoring Tool
[2013] Ropossum V1.0: A Physics-based Game Authoring Tool
[2013] Crospell Engine – Natural Language Processing Engine
[2013] Social Relationship and Decision-Making explained by Fuzzy Logic
[2013] Personalizing Player Experience in First-Person Shooter Games, UMAP 2013
[2012] Adaptive Games Content Generation for 2D Mario
[2012] STYX Foodiac Nutrition System
[2012] Gaming and Robotics – Virtual Reality Seminar 2012
[2012] Wikilogia Event – Freshman 2012: ArtificiaI Intelligence brief Introduction
[2011] Car Dynamics [Physics Simulation]
Courses I Enjoyed, Online
2016
User Experience (UX): The Ultimate Guide to Usability and UX
Insights on Graphic Design with Sean Adams
Logo Design ARMM
Applied Interaction Design
Color Theory for Today's Creative Professionals
2015
UX Design for Mobile Developers
Interaction Design
Up and Running with AngularJS
iOS 8 App Development with Swift
iOS App Development Essential Training
2014
Rails Development
Fixed, Fluid, Adaptive, and Responsive
2013
CS223A - Introduction to Robotics
Machine Learning - University of Washington
Machine Learning - Stanford Coursera
2012
Engineering Software as a Service (SaaS + Ruby) - Berkeley
Human-Computer Interaction - Stanford
Machine Learning - Stanford CS229
2010
CS50 - Harvard
Introduction to Algorithms - MIT