Files

8.0 KiB

[EDU] IBM Watson's User Modeling tech models characters in The Lord Of The Rings

Post:

Link to content

Comments:

u/alexanderwales [+7] Time flies like an arrow (an hour later)

Visualization from my last 474 reddit comments. I am quite curious how these numbers were arrived at, since some of it seems off to me - I don't trust a 99% in anything.

General instructions for doing this yourself

  1. Go to http://pastebin.com/j1QxzKiR and copy-paste that script into notepad++ then save it as "whatever.py".

  2. Download Python 2.x from someplace like http://www.python.org/getit/releases/2.7/ and then install it.

  3. Run the python file from the command line (see here: http://mail.python.org/pipermail/tutor/2004-July/030634.html) and input your username.

  4. This will create a file called "username".txt (for example, keltranis.txt or I_RAPE_CATS.txt), but might take some time - just let it run until it's done.

  5. Now that your comment history is saved as a .txt file, go into Notepad++ and strip out all of the stuff that is not part of your actual commenting. This is the hard part, because you have to learn a bit of the context stuff. This was of great help: http://markantoniou.blogspot.com/2008/06/notepad-how-to-use-regular-expressions.html

(I use a somewhat different setup now, but if you have some programming knowledge that should be enough to set you on the right path.)

u/traverseda [+1] With dread but cautious optimism (10 hours later)

go into Notepad++ and strip out all of the stuff that is not part of your actual commenting.

Or, alternatively, use this version. Also respects reddits flood rules and won't do more then 30 requests per minute (no more "HTTP Error 429: Too Many Requests").

u/alexanderwales [+1] Time flies like an arrow (11 hours later)

Yeah, my own script is a little different from what I posted above - that whole chunk is a copy+paste of a process description I wrote nearly three years ago here, and the pastebin itself was something that was posted to /r/TheoryOfReddit but whose origins are now lost to me.

I believe the script still stops itself at 40 requests either way though, since I don't believe that there's a way to get past the 1000 comment mark for a user (though I might be wrong on that).

u/PixelDust73 [+3] (2 hours later)

There is no way Bilbo is more neurotic than Gollem!

u/DataPacRat [+2] Amateur Immortalist (an hour later)

Is there any actual use that these numbers can be put to?

For example, say I put the first three books of SI into the machine, and sorted the various categories out by how extreme they were (eg, both 99% and 1% at the top)... what would I actually learn from the results, or be able to do that I couldn't have done before?

u/alexanderwales [+4] Time flies like an arrow (2 hours later)

I believe that an intended use-case is something like this:

  • Gather up a thousand different user profiles.
  • Run them through this program.
  • Find correlations between these models and another data set (say, those users favorite movies - you should have both if you scraped from Facebook).
  • Use those correlations to build a recommendation engine.

For a single user (or a single book), I doubt it's of much use, but assuming that its results have any validity, it would be useful for grouping similar people.

u/xamueljones [+2] My arch-enemy is entropy (3 hours later)

It doesn't seem very useful for you unless you understand very clearly (on the level of a trained psychologist) the meaning of each personality trait and how they interact. Knowing how your character is being portrayed is a useful thing to know, but how can you actually take advantage of this? The only thing I can think of is to see whose personality is changing in ways you aren't intending, but you seem to be a good enough writer to understand that it's more fun to let your characters do what they want instead of shoehorning them into contrived situations.

u/AmeteurOpinions [+1] Finally, everyone was working together. (2 hours later)

You, the author, probably wouldn't learn all that much. You would have to sort the dialogue and narration by character first, but it might be neat to see how they change across sequels in exact numbers. These obviously won't be very precise.

The tech has a ways to go, but it's scary how it's come already. Imagine running it on a database of presidential quotes -- you could directly compare George Washington to George Bush!

u/None [+1] (a day later)

Could be used to reconstruct your mind after you're dead ;-).

u/E-o_o-3 [+2] (2 hours later)

I put in two blocks of text I've written,and there wasn't a high level of agreement, but maybe they just weren't large enough.

How did they determine which words were "neurotic" and so on?

u/AmeteurOpinions [+1] Finally, everyone was working together. (3 hours later)

Honestly, I have no clue.

u/None [+1] (2 days later)

I had pretty high agreement between various blocks I wrote.

u/None [+2] (6 hours later)

fascinating, thanks. crossposted over at r/tolkienfans

u/Gurkenglas [+2] (10 hours later)

Running this on yourself and checking whether the results apply to you seems like listening to cold reading.

u/AmeteurOpinions [+2] Finally, everyone was working together. (19 hours later)

Running it on just yourself is mostly useless. You want to run it on all of Reddit and then start comparing numbers.

u/traverseda [+2] With dread but cautious optimism (11 hours later)

I'm really not clear on what "self-transcendence" means.

u/None [+2] (17 hours later)

[deleted]

u/RemindMeBot [+1] (17 hours later)

Messaging you on [2014-12-27 18:08:15 UTC](http://www.wolframalpha.com/input/?i=2014-12-27 18:08:15 UTC To Local Time) to remind you of this comment.

[CLICK THIS LINK](http://www.reddit.com/message/compose/?to=RemindMeBot&subject=Reminder&message=[http://www.reddit.com/r/rational/comments/2qgdvs/edu_ibm_watsons_user_modeling_tech_models/cn6hhkp]%0A%0ARemindMe! 8 hours ) to send a PM to also be reminded and to reduce spam.


^([FAQs]) ^| [^([Custom Reminder])](http://www.reddit.com/message/compose/?to=RemindMeBot&subject=Reminder&message=[LINK INSIDE SQUARE BRACKETS else default to FAQs]%0A%0ANOTE: Don't forget to add the time options after the command.%0A%0ARemindMe!) ^| ^([Feedback]) ^| ^([Code])

u/AmeteurOpinions [+1] Finally, everyone was working together. (3 minutes later)

This sums it up:

TLDR: IBM Watson crunches the numbers on LOTR characters. Tells us who is self-conscious, who is neurotic, and whether book Aragorn or movie Aragorn is more alpha. Bonus: Watson breaks down the personality of the average LOTR tweeter.

The system just needs samples of text, so dialogue taken from the books or screenplays is perfect for crunching. You can try it yourself here.

u/None [+3] (38 minutes later)

This is you, from your posts over the last week. Merry Christmas!

xoxo,
/u/seraphnb