• Chișinău Jewish cemetery

    Chișinău Jewish cemetery

    March 4, 2019

    Two years ago I visited Chișinău (Kishinev), the city in Moldova where I was born and where I grew up until the age of fifteen. Today I saw a post with photos from the ancient Chișinău Jewish cemetery and recalled that I too, took many pictures from that sad place. Less than half of the original cemetery survived to these days. The bigger part of it was demolished in the 1960s in favor of a park and a residential area. If you scroll through the pictures below, you will be able to see how they used tombstones to build the park walls.

    Another notable feature of many Jewish cemeteries is memorial plates in memoriam of the relatives who don’t have their own graves – the relatives who were murdered over the course of the Jewish history.

    March 4, 2019 - 2 minute read -
    chisinau jewish kishinev moldova blog
  • How to Increase Retention and Revenue in 1,000 Nontrivial Steps

    How to Increase Retention and Revenue in 1,000 Nontrivial Steps

    February 13, 2019

    The journey of a thousand miles begins with one step. My coworker, Yanir Seroussi, wrote about the work of data scientists in the marketing team.

    February 13, 2019 - 1 minute read -
    blog
  • On procrastination, or why too good can be bad

    On procrastination, or why too good can be bad

    February 4, 2019

    I’m a terrible procrastinator. A couple of years ago, I installed RescueTimeto fight this procrastination. The idea behind RescueTime is simple — it tracks the sites you visit and the application you use and classifies them according to how productive you are. Using this information, RescueTime provides a regular report of your productivity. You can also trigger the productivity mode, in which RescueTime will block all the distractive sites such as Facebook, Twitter, news sites, etc. You can also configure RescueTime to trigger this mode according to different settings. This sounded like a killer feature for me and was the main reason behind my decision to purchase a RescueTime subscription. Yesterday, I realized how wrong I was.

    RescueTime logo

    When I installed RescueTime, I was full of good intentions. That is why I configured it to block all the distractive sites for one hour every time I accumulate more than 10 minutes of surfing such sites. However, from time to time, I managed to find a good excuse to procrastinate. Although RescueTime allows you to open a “bad” site after a certain delay, I found this delay annoying and ended up killing the RescueTime process (killing a process is faster than temporary disabling a filter). As a result, most of my workday stayed untracked, unmonitored, and unfiltered.

    So, I decided to end this absurd situation. As of today, RescueTime will never block any sites. Instead of blocking, I configured it to show a reminder and to open my RescueTime dashboard, as a reminder to behave myself. I don’t know whether this non-intrusive reminder will be effective or not but at least I will have correct information about my day.

    February 4, 2019 - 2 minute read -
    procrastination productivity rescuetime blog Productivity & Procrastination

  • "Why it burns when you P" and other statistics rants

    January 20, 2019

    “Sunday grumpiness” is an SFW translation of Hebrew phrase that describes the most common state of mind people experience on their first work weekday. My grumpiness causes procrastination. Today, I tried to steer this procrastination to something more productive, so I searched for some statistics-related terms and stumbled upon a couple of interesting links in which people bitch about p-values.

    Why it burns when you P” is a five-years-old rant about P values. It’s funny, informative and easy to read

    Everything Wrong With P-Values Under One Roof” is a recent rant about p-values written in a form of a scientific paper. William M. Briggs, the author of this paper, ends it with an encouraging statement: “No, confidence intervals are not better. That for another day.”

    Everything wrong with statistics (and how to fix it)” is a one-hour video lecture by Dr. Kristin Lennox who talks about the same problems. I saw this video, and two more talks by Dr. Lennox on a flight I highly recommend all her videos on YouTube.

    Do You Hate Statistics as Much as Everyone Else?” – A Natan Yau’s (from flowingdata.com) attempt to get thoughtful comments from his knowledgeable readers.

    This list will not be complete without the classics:

    Why Most Published Research Findings Are False”, “Mindless Statistics”, and “Cargo Cult Science”. If you haven’t read these three pieces of wisdom, you absolutely should, they will change the way you look at numbers and research.

    *The literal meaning of שביזות יום א is Sunday dick-brokenness.

    January 20, 2019 - 2 minute read -
    blog
  • Hackers beware: Bootstrap sampling may be harmful

    Hackers beware: Bootstrap sampling may be harmful

    January 15, 2019

    Anything is better when bootstrapped. Read my co-worker’s post on bootstrapping. Also make sure following the links Yanir gives to support his claims

    January 15, 2019 - 1 minute read -
    blog
  • I have 101 followers!

    I have 101 followers!

    January 14, 2019

    Yesterday, the follower list of my blog exceeded one hundred followers! Even though I know that some of these followers are bots, this number makes me happy! Thank you all (humans and bots) for clicking the “follow” button.

    January 14, 2019 - 1 minute read -
    blogging followers blog
  • A Brand Image Analysis of WordPress and Automattic on Twitter

    A Brand Image Analysis of WordPress and Automattic on Twitter

    January 13, 2019

    My coworker analyzed Twitter social network around Automattic, WordPress, and other related projects.

    January 13, 2019 - 1 minute read -
    blog
  • Against A/B tests

    Against A/B tests

    December 12, 2018

    Traditional A/B testsing rests on a fundamentally flawed premise. Most of the time, version A will be better for some subgroups, and version B will be better for others. Choosing either A or B is inherentlyinferior to choosing a targeted mix of A and B.

    Michael Kaminsky locallyoptimistic.com

    The quote above is from a post by Michael Kaminsky “Against A/B tests”. I’m still not fully convinced by Michael’s thesis but it is very interesting and thought-provoking.

    December 12, 2018 - 1 minute read -
    a-b-testing data science reblog statistics blog
  • Links Worth Sharing: What Makes People Successful

    Links Worth Sharing: What Makes People Successful

    November 27, 2018
    November 27, 2018 - 1 minute read -
    blog
  • Useful redundancy — when using colors is not completely useless

    Useful redundancy — when using colors is not completely useless

    November 26, 2018

    The maximum data-ink ratio principle implies that one should not use colors in their graphs if the graph is understandable without the colors. The fact that you can do something, such as adding colors, doesn’t mean you should do it. I know it. I even have a dedicated tag on this blog for that. Sometimes, however, consistent use of colors serves as a useful navigation tool in a long discussion. Keep reading to learn about the justified use of colors.

    Pew Research Center is a “is a nonpartisan American fact tank based in Washington, D.C. It provides information on social issues, public opinion, and demographic trends shaping the United States and the world.” Recently, I read a report prepared by the Pew Center on the religious divide in the Israeli society. This is a fascinating report. I recommend reading without any connection to data visualization.

    But this post does not deal with the Isreali society but with graphs and colors.

    Look at the first chart in that report. You may see a tidy pie chart with several colored segments.

    Pie chart: Religious composition of Israeli society. The chart uses several colored segments

    Aha! Can’t they use a single color without losing the details? Of course the can! A monochrome pie chart would contain the same information:

    Pie chart: Religious composition of Israeli society. The chart uses monochrome segments

    In most of the cases, such a transformation would make a perfect sense. In most of the cases, but not in this report. This report is a multipage research document packed with many facts and analyses. The pie chart above is the first graph in that report that provides a broad overview of the Israeli society. The remaining of this report is dedicated to the relationships between and within the groups represented by the colorful segments in that pie chart. To help the reader navigating through this long report, its authors use a consistent color scheme that anchors every subsequent graph to the relevant sections of the original pie chart.

    All these graphs and tables will be readable without the use of colors. Despite the fact that the colors here are redundant, this is a useful redundancy. By using the colors, the authors provided additional information layers that make the navigation within the document easier. I learned about the concept of useful redundancy from “Trees, Maps, and Theorems” by Jean-luc Dumout. If you can only read one book about data communication, it should be this book.

    November 26, 2018 - 2 minute read -
    because you can colors data visualisation Data Visualization dataviz Israel redundancy blog
  • On the importance of perspective

    On the importance of perspective

    November 12, 2018

    Stalin was a relatively short man, his height was 1.65 m. Khrushchev was even shorter, his height was 1.60. It seems that the difference wasn’t enough for the official Soviet propaganda of that time. Take a look at this photo. We can clearly see that Stalin is taller than Khrushchev.

    stalin.png

    Do you notice something strange? Take a look at the windows in the background. I added horizontal and vertical guides for your convenience.

    Screen Shot 2018-11-05 at 8.38.08

    Now, look what happens when we fix the horizontal and vertical lines

    Screen Shot 2018-11-05 at 8.39.03

    Now, Khrushchev is still shorter than Stalin but not by that much.

    November 12, 2018 - 1 minute read -
    khrushchev perspective photo photography stalin blog
  • Microtext Line Charts

    Microtext Line Charts

    November 12, 2018

    Why adding text labels to graph lines, when you can build graph lines using text labels? On microtext lines

    November 12, 2018 - 1 minute read -
    data visualisation Data Visualization dataviz microtext blog
  • איך אומרים דאטה ויזואליזיישן בעברית?

    איך אומרים דאטה ויזואליזיישן בעברית?

    October 23, 2018

    This post is written in Hebrew about a Hebrew issue. I won’t translate it to English.

    אני מלמד data visualization בשתי מכללות בישראלבמכללת עזריאלי להנדסה בירושלים ובמכון הטכנולוגי בחולון. כשכתבתי את הסילבוס הראשון שלי הייתי צריך למצוא מונח ל־data visualization וכתבתיהדמיית נתונים״ אומנם זה הזכיר לי קצת תהליך של סימולציה, אבל האופציה האחרת ששקלתי היתה ״דימות״ וידעתי שהיא שמורה ל־imaging, דהיינו תהליך של יצירת דמות או צורה של עצם, בעיקר בעולם הרפואה.

    הבנתי שהמונח בעייתי בשיעור הראשון שהעברתי. מסתברששניים מארבעת הסטודנטים שהגיעו לשיעור חשבו שקורס ״הדמיית נתונים בתהליך מחקר ופיתוח״ מדבר על סימולציות.

    מתישהו שמעתי מחבר של חבר שהמונח הנכון ל־visualization זה הדמאה, אבל זה נשמע לי פלצני מדי, אז השארתי את ה־״הדמיה״ בשם הקורס והוספתי “data visualization” בסוגריים.

    היום, שלוש שנים אחרי ההרצאה הראשונה שהעברתי, ויומיים לפני פתיחת הסמסטר הבא, החלטתי לגגל (יש מילה כזאת? יש!) את התשובה. ומה מסתבר? עלון ״למד לשונך״ מס׳ 109 של האקדמיה ללשון עברית שיצא לאור בשנת 2015 קובע שהמונח ל־visualization הוא הַחְזָיָה. לא יודע מה אתכם, אבל אני לא משתגע על החזיה. עוד משהו שאני לא משתגע עליו הוא שבתור הדוגמא להחזיה, האקדמיה החלטיה לשים תרשים עוגה עם כל כך הרבה שגיאות!

    Screen Shot 2018-10-23 at 20.35.52

    נראה לי שאני אשאר עם הדמיה. ויקימילון מרשה לי.

    נ.ב. שמתם לב שפוסט זה השתמשתי במקף עברי? אני מאוד אוהב את המקף העברי.

    October 23, 2018 - 2 minute read -
    data visualisation Data Visualization dataviz hebrew הדמיה החזיה blog
  • Innumeracy

    Innumeracy

    October 22, 2018

    Innumeracy is “inability to deal comfortably with the fundamental notions of number and chance”.
    I which there was a better term for “innumeracy”, a term that would reflect the importance of analyzing risks, uncertainty, and chance. Unfortunately, I can’t find such a term. Nevertheless, the problem is huge. In this long post, Tom Breur reviews many important aspects of “numeracy”.

    October 22, 2018 - 1 minute read -
    blog
  • Working Remotely and the Virtue of Aggressive Transparency

    Working Remotely and the Virtue of Aggressive Transparency

    October 16, 2018

    Excellent post by my colleague Simon Ouderkirk on working in a distributed company. It’s a three-year-old post. I wonder how I missed it.

    October 16, 2018 - 1 minute read -
    blog
  • Data visualization in right-to-left languages

    Data visualization in right-to-left languages

    October 15, 2018

    If you speak Arabic or Farsi, I need your help. If you don’t speak, share this post with someone who does.

    Right-to-left (RTL) languages such as Hebrew, Arabic, and Farsi are used by roughly 1.8 billion people around the world. Many of them consume data in their native languages. Nevertheless, I have never seen any research or study that explores data visualization in RTL languages. Until a couple of days ago, when I saw this interesting observation by Nick Doiron “Charts when you read right-to-left”.

    I teach data visualization in Israeli colleges. Whenever a student asks me RTL-related questions, I always answer something like “it’s complicated, let’s not deal with that”. Moreover, in the assignments, I even allow my students to submit graphs in English, even if they write the report in Hebrew.

    Nick’s post made me wonder about data visualization do’s and don’ts in RTL environments. Should Hebrew charts differ from Arabic or Farsi? What are the accepted practices?

    If you speak Arabic or Farsi, I need your help. If you don’t speak, share this post with someone who does. I want to collect as many examples of data visualization in RTL languages. Links to research articles are more than welcome. You can leave your comments here or send them to boris@gorelik.net.

    Thank you.

    The image at the top of this post is a modified version of a graph that appears in the post that I cite. Unfortunately, I wasn’t able to find the original publication.

    October 15, 2018 - 2 minute read -
    arabic data visualisation Data Visualization dataviz farsi help RTL blog
  • A World Without the Number 6 — Math with Bad Drawings

    A World Without the Number 6 — Math with Bad Drawings

    October 11, 2018

    What will happen if number 6 disappears one day? Ben Orlin, the author of “Math with bad drawings” elaborates on this interesting thought experiment in this 2017 post.

    October 11, 2018 - 1 minute read -
    math mathematics repost blog
  • Can error correction cause more error? (The answer is yes)

    Can error correction cause more error? (The answer is yes)

    October 9, 2018

    This is an interesting thought experiment. Suppose that you have some appliance that acts in a normally distributed way. For example, a nerf gun. Let’s say now that you aim and fire the gun. What happens if you miss by some amount of X? Should you correct your aim in the opposite direction? My intuition says “yes.” So does the intuition of many other people with whom I talked about this problem. However, when we start thinking about this problem, we realize that the intuition is wrong. Since we aim the gun, our assumption should be that the deviation is zero. A single observation is not sufficient to reject this assumption. By continually adjusting the data generating process based on a single observation, we reduce the precision (increase the dispersion).
    Below is a simulation of adjusted and non-adjusted processes (the code is here). The broader spread of the adjusted data (blue line) is evident.

    Two curves. Blues: high dispersion of values when adjustments are performed after every observation. Orange: smaller dispersion when no adjustments are done.

    Due to the nature of the normal random variable, a single large accidental deviation can cause an extreme “correction,” which in turn will create a prolonged period of highly inaccurate points. This is precisely what you see in my simulation.
    The moral of this simple experiment is that you shouldn’t let a single affect your actions.

    October 9, 2018 - 1 minute read -
    distribution statistics blog
  • Me

    Me

    October 1, 2018
    October 1, 2018 - 1 minute read -
    me blog
  • Conference Recap: EuroSciPy 2018 — Data for Breakfast

    Conference Recap: EuroSciPy 2018 — Data for Breakfast

    September 20, 2018

    See my recap of the recent EuroSciPy, published on https://data.blog

    In which Boris Gorelik shares his favorite talks and workshops from EuroSciPy 2018.

    via Conference Recap: EuroSciPy 2018 — Data for Breakfast

    September 20, 2018 - 1 minute read -
    data visualisation Data Visualization dataviz euroscipy public speaking python blog

  • "Any questions?" How to fight the awkward silence at the end of a presentation?

    September 20, 2018

    If you ever gave or attended a presentation, you are familiar with this situation: the presenter asks whether there are any questions and … nobody asks anything. This is an awkward situation. Why aren’t there any questions? Is it because everything is clear? Not likely. Everything is never clear. Is it because nobody cares? Well, maybe. There are certainly many people that don’t care. It’s a fact of life. Study your audience, work hard to make the presentation relevant and exciting but still, some people won’t care. Deal with it.

    However, the bigger reasons for lack of the questions are human laziness and the fear of being stupid. Nobody likes asking a question that someone will perceive as a stupid one. Sometimes, some people don’t mind asking a question but are embarrassed and prefer not being the first one to break the silence.

    What can you do? Usually, I prepare one or two questions by myself. In this case, if nobody asks anything, I say something like “Some people, when they see these results ask me whether it is possible to scale this method to larger sets.”. Then, depending on how confident you are, you may provide the answer or ask “What do you think?”.

    You can even prepare a slide that answers your question. In the screenshot below, you may see the slide deck of the presentation I gave in Trento. The blue slide at the end of the deck is the final slide, where I thank the audience for the attention and ask whether there are any questions.

    My plan was that if nobody asks me anything, I would say “Thank you again. If you want to learn more practical advises about data visualization, watch the recording of my tutorial, where I present this method <SLIDE TRANSFER, show the mockup of the “book”>. Also, many people ask me about reading suggestions, this is what I suggest you read: <SLIDE TRANSFER, show the reading pointers>

    Screen Shot 2018-09-17 at 10.10.21

    Luckily for me, there were questions after my talk. Luckily, one of these questions was about practical advice so I had a perfect excuse to show the next, pre-prepared, slide. Watch this moment on YouTube here.

    September 20, 2018 - 2 minute read -
    data visualisation Data Visualization presentation presentation-tip presenting public speaking blog
  • Graphing Highly Skewed Data – Tom Hopper

    Graphing Highly Skewed Data – Tom Hopper

    September 16, 2018

    My colleague, Chares Earl, pointed me to this interesting 2010 post that explores different ways to visualize categories of drastically different sizes.

    The post author, Tom Hopper, experiments with different ways to deal with “Data Giraffes”. Some of his experiments are really interesting (such as splitting the graph area). In one experiment, Tom Hopper draws bar chart on a log scale. Doing so is considered as a bad practice. Bar charts value (Y) axis must include meaningful zero, which log scale can’t have by its definition.

    Other than that, a good read Graphing Highly Skewed Data – Tom Hopper

    September 16, 2018 - 1 minute read -
    bar plot data data visualisation Data Visualization dataviz blog
  • On privacy, security, and irony

    On privacy, security, and irony

    September 9, 2018

    About a week ago, I met Justin Mayer and had a really interesting chat with him about internet privacy. Today, his 30-minutes talk on that subject appeared in my youtube suggestion list

    https://www.youtube.com/watch?v=2rrP_aW-jNA

    How ironic. The talk, by the way, is very interesting.

    September 9, 2018 - 1 minute read -
    irony privacy security blog
  • Back to Mississippi: Black migration in the 21st century. By Charles Earl

    Back to Mississippi: Black migration in the 21st century. By Charles Earl

    September 4, 2018

    I wonder how this analysis of remained unnoticed by the social media

    The recent election of Doug Jones […] got me thinking: What if the Black populations of Southern cities were to experience a dramatic increase? How many other elections would be impacted?

    via Back to Mississippi: Black migration in the 21st century — Charlescearl’s Weblog

    September 4, 2018 - 1 minute read -
    data-journalism data science race blog
  • Please leave a comment to this post

    Please leave a comment to this post

    September 3, 2018

    Please leave a comment to this post. It doesn’t matter what. It doesn’t matter when or where you see it. I want to see how many real people are actually reading this blog.

    [caption id=”attachment_media-15” align=”alignnone” width=”1880”]close up of text

    Photo by Pixabay on Pexels.com[/caption]

    September 3, 2018 - 1 minute read -
    перекличка feedback blog
  • Older posts Newer posts