Compare commits

...
Author SHA1 Message Date
Jim Miller 8347f4490e Update CLI download zip. 2014-06-14 10:42:23 -05:00
Jim Miller 360d37746d Fix typo. 2014-06-14 10:39:52 -05:00
Jim Miller 8af36f298e Bump versions. 2014-06-14 10:37:58 -05:00
Jim Miller 831370134b Do metadata split *before* regexp replacement, not after. 2014-06-12 17:41:44 -05:00
Jim Miller de37c4aa1d Add '\,' split list feature for replace_metadata. \, will split into multiple. 2014-06-11 19:12:59 -05:00
cryzed 2cb139147a Fixed finding of chapter URLs 2014-06-11 18:37:28 +02:00
cryzed 9ce6117688 Set language metadata to Hungarian 2014-06-10 22:04:28 +02:00
cryzed 2e38ef1122 Handle tags within the title anchor 2014-06-10 21:08:10 +02:00
cryzed 4c4576f331 Minor improvement 2014-06-10 14:57:17 +02:00
cryzed 9aa75905c6 Added adapter for http://fanfiction.csodaidok.hu/ and default configurations 2014-06-10 14:54:00 +02:00
Jim Miller 76823dccfb Comment on generate_cover_settings about 'Allow generate_cover_settings from personal.ini to override' 2014-06-09 17:09:23 -05:00
Jim Miller 0d0778fea5 Fix bug-numWords in anthology when source books(s) don't have numWords. 2014-06-09 17:08:55 -05:00
cryzed 5823d335a4 Raise error instead when the correct metadata can't be found on the author's page 2014-06-09 02:05:32 +02:00
cryzed db3878668b Fixed missing numWords attribute 2014-06-09 00:44:29 +02:00
cryzed 02289c0af1 Fixed possible downloading of wrong story 2014-06-09 00:35:05 +02:00
cryzed 7a840043f0 Set language metadata to Hungarian 2014-06-08 23:10:10 +02:00
cryzed 0c3ccb4e7c Removed superfluous code 2014-06-08 22:31:33 +02:00
cryzed e8904ec061 Added adapter for http://fanfic.hu/ and default configurations 2014-06-08 21:28:00 +02:00
cryzed a589cf4280 Removed debug print statement and fixed a potential bug 2014-06-08 18:32:09 +02:00
cryzed cf11959970 Backported some changes from the http://nocturnal-light.net/ adapter 2014-06-08 15:48:10 +02:00
cryzed 49777c299e Added adapter for http://nocturnal-light.net/ and default configurations 2014-06-08 15:47:54 +02:00
cryzed 4acffb88f6 Remove deprecated comment 2014-06-08 12:51:24 +02:00
cryzed 6e38557454 Improve exception handling for the http://bloodshedverse.com/ adapter 2014-06-08 12:48:01 +02:00
cryzed f24c363d3b Improve exception handling for the voracity2.e-fic.com/ adapter 2014-06-08 12:46:50 +02:00
cryzed e7ea699bc9 Improve exception handling for the http://spikeluver.com/ adapter. Also handle "story does not exist" error text and raise an appropriate exception 2014-06-08 12:45:04 +02:00
Jim Miller 9de65d94f3 Moved tag calibre-plugin-1.8.23 to changeset d2810246eaf9 (from changeset 992d764d07c4) 2014-06-07 09:54:37 -05:00
Jim Miller da7498d202 Moved tag FanFictionDownLoader-4.5.04 to changeset d2810246eaf9 (from changeset 992d764d07c4) 2014-06-07 09:54:31 -05:00
Jim Miller 05ec7bce2b Update CLI download zip. 2014-06-07 09:54:15 -05:00
Jim Miller 241c4d8d52 Added tag calibre-plugin-1.8.23 for changeset 992d764d07c4 2014-06-07 09:52:26 -05:00
Jim Miller ccf4c8cc4e Added tag FanFictionDownLoader-4.5.04 for changeset 992d764d07c4 2014-06-07 09:52:04 -05:00
Jim Miller 7498a9aa93 Update CLI download zip. 2014-06-07 09:51:56 -05:00
Jim Miller 25c63c3a47 Bump versions. 2014-06-07 09:43:32 -05:00
cryzed c5e5a9bb84 Don't automatically change the skin, instead throw an error when the required skin isn't set informing the user 2014-06-07 03:40:11 +02:00
cryzed 9bc36c3652 Changed exceptions to something more appropriate 2014-06-07 03:09:40 +02:00
cryzed 0736c35be0 Automatically change to required skin when parsing http://voracity2.e-fic.com/ adapter stories and restore the original skin afterwards 2014-06-07 03:03:54 +02:00
cryzed cd4f2c2717 Added comment as a reminder for an obscure parsing bug 2014-06-07 02:26:39 +02:00
cryzed e6bb8c557b Improve error handling for the http://voracity2.e-fic.com/ adapter 2014-06-06 23:52:38 +02:00
cryzed f4da7dc1bd Fixed summary parsing error for the http://voracity2.e-fic.com/ adapter 2014-06-06 23:09:59 +02:00
Jim Miller eba312e777 cryzed's Changes to voracity, add bloodshedverse.com and spikeluver.com. 2014-06-06 13:05:02 -05:00
Jim Miller 0cd615d950 Make anthology tag configurable. 2014-06-06 12:53:26 -05:00
Jim Miller 4b120bb2d3 Added tag calibre-plugin-1.8.22 for changeset f5d687351419 2014-06-04 21:45:00 -05:00
Jim Miller c9bcf9175e Added tag FanFictionDownLoader-4.5.03 for changeset f5d687351419 2014-06-04 21:44:50 -05:00
Jim Miller c76facd40c Update CLI download zip. 2014-06-04 21:44:30 -05:00
Jim Miller f1d52834d1 Bump versions. 2014-06-04 21:43:17 -05:00
Jim Miller e8ac7f8a89 cryzed's changes: New site voracity2.e-fic.com, pprint -m from downloader. 2014-06-04 21:42:20 -05:00
Jim Miller 3e0f92d8ce Added tag calibre-plugin-1.8.21 for changeset 943e3d21e5fa 2014-05-23 20:17:51 -05:00
Jim Miller 2682d0fe36 Added tag FanFictionDownLoader-4.5.02 for changeset 943e3d21e5fa 2014-05-23 20:17:43 -05:00
Jim Miller 125e29091f Update CLI download zip. 2014-05-23 20:17:27 -05:00
Jim Miller d0c73d5444 Bump versions. 2014-05-23 20:15:49 -05:00
Jim Miller bd8e54edcf Fixes for literotica.com: URLs using //site, allow https, ch01 as storyId, multi ch only. 2014-05-17 21:12:43 -05:00
Jim Miller d680a86f0c Fix for dark-solace.org Rating. 2014-05-17 15:55:20 -05:00
Jim Miller f075ae582d Added tag FanFictionDownLoader-4.5.01 for changeset b3a1ca11a76d 2014-05-13 20:18:11 -05:00
Jim Miller 928ebb9751 Added tag calibre-plugin-1.8.20 for changeset b3a1ca11a76d 2014-05-13 20:17:57 -05:00
Jim Miller ce262be162 Update CLI download zip. 2014-05-13 20:17:44 -05:00
Jim Miller 6faa6850af Bump versions. 2014-05-13 20:16:22 -05:00
Jim Miller f392c6dd77 Fix for dark-solace.org metadata parsing. 2014-05-13 12:00:58 -05:00
Jim Miller af0dff28b4 Fix for fictionpad.com removing 'dislikes' in some(all?) cases. 2014-05-10 14:30:53 -05:00
Jim Miller 02a75a821f Fix for storiesonline.net changing urls. 2014-05-10 14:30:29 -05:00
Jim Miller c06028b498 Fix for AO3 story not found. 2014-05-10 14:30:06 -05:00
Jim Miller de4b95af9b Added tag FanFictionDownLoader-4.5.00 for changeset fbdadac26cf9 2014-05-05 12:54:38 -05:00
Jim Miller 5366355d96 Added tag calibre-plugin-1.8.19 for changeset fbdadac26cf9 2014-05-05 12:54:23 -05:00
Jim Miller 264853d768 Update CLI download zip. 2014-05-05 12:54:11 -05:00
Jim Miller d82b399738 Bump versions. 2014-05-05 12:50:30 -05:00
Jim Miller 59446d23dc in/exclude_metadata_pre/post feature. 2014-05-05 12:21:55 -05:00
Jim Miller 6fcbdd8a8d https for squidge.org/peja, allow story url change http to https on any site. 2014-05-03 11:27:39 -05:00
Jim Miller 1348525d25 Fix for some stories' summaries on onedirectionfanfiction.com. 2014-05-02 11:46:59 -05:00
Jim Miller 40bac62c5e Change CLI -m output slightly. 2014-05-02 11:25:13 -05:00
Jim Miller 362f15f9fa Slightly improved conn refused handling. 2014-04-22 13:01:33 -05:00
Jim Miller b833041dc4 Allow https for FimF, still use http for canonical. 2014-04-22 13:01:26 -05:00
Jim Miller fa67220b86 Added tag FanFictionDownLoader-4.4.99 for changeset fdbc5eed9810 2014-04-18 14:18:55 -05:00
Jim Miller 4f412eb89f Added tag calibre-plugin-1.8.18 for changeset fdbc5eed9810 2014-04-18 14:18:41 -05:00
Jim Miller 6a4aa4340e Update CLI download zip. 2014-04-18 14:18:20 -05:00
Jim Miller dc785b911e Add fix_fimf_blockquotes feature, bump versions. 2014-04-18 14:17:26 -05:00
Jim Miller d403f916a9 Added tag FanFictionDownLoader-4.4.98 for changeset ec9ec0e9a813 2014-04-09 08:05:47 -05:00
Jim Miller 29fe1a6e24 Added tag calibre-plugin-1.8.17 for changeset ec9ec0e9a813 2014-04-09 08:05:32 -05:00
Jim Miller 131a08c0dc Update CLI download zip. 2014-04-09 08:05:14 -05:00
Jim Miller fbd26c16e0 Bump versions. 2014-04-09 08:03:16 -05:00
Jim Miller 34ebba40d0 Fix for a problem with literotica.com multi-page chapters on Kobo readers. 2014-04-08 09:32:36 -05:00
Jim Miller 7ce8436208 Added tag FanFictionDownLoader-4.4.97 for changeset 52e576f78b3d 2014-03-29 17:22:01 -05:00
Jim Miller b247e4fc7b Added tag calibre-plugin-1.8.16 for changeset 52e576f78b3d 2014-03-29 17:21:49 -05:00
26 changed files with 1707 additions and 98 deletions
+1 -1
View File
@@ -1,6 +1,6 @@
# ffd-retief-hrd fanfictiondownloader
application: fanfictiondownloader
version: 4-4-97
version: 4-5-05
runtime: python27
api_version: 1
threadsafe: true
+1 -1
View File
@@ -42,7 +42,7 @@ class FanFictionDownLoaderBase(InterfaceActionBase):
description = _('UI plugin to download FanFiction stories from various sites.')
supported_platforms = ['windows', 'osx', 'linux']
author = 'Jim Miller'
version = (1, 8, 16)
version = (1, 8, 24)
minimum_calibre_version = (1, 13, 0)
#: This field defines the GUI plugin class that contains all the code
+12 -9
View File
@@ -970,7 +970,8 @@ class FanFictionDownLoaderPlugin(InterfaceAction):
if book_id and mi: # book_id and mi only set if matched by title/author.
liburl = self.get_story_url(db,book_id)
if book['url'] != liburl and prefs['checkforurlchange'] and \
not (book['url'].replace('https','http') == liburl and 'fanfiction.net' in liburl):
not (book['url'].replace('https','http') == liburl): # several sites have been changing to
# https now. Don't flag when that's the only change.
# special case for ffnet urls change to https.
if not question_dialog(self.gui, _('Change Story URL?'),'''
<h3>%s</h3>
@@ -1984,11 +1985,12 @@ class FanFictionDownLoaderPlugin(InterfaceAction):
for (k,v) in b['all_metadata'].iteritems():
#print("merge_meta_books v:%s k:%s"%(v,k))
if k in ('numChapters','numWords'):
if k not in book['all_metadata']:
book['all_metadata'][k] = b['all_metadata'][k]
else:
# lot of work for a simple add.
book['all_metadata'][k] = unicode(int(book['all_metadata'][k].replace(',',''))+int(b['all_metadata'][k].replace(',','')))
if k in b['all_metadata'] and b['all_metadata'][k]:
if k not in book['all_metadata']:
book['all_metadata'][k] = b['all_metadata'][k]
else:
# lot of work for a simple add.
book['all_metadata'][k] = unicode(int(book['all_metadata'][k].replace(',',''))+int(b['all_metadata'][k].replace(',','')))
elif k in ('dateUpdated','datePublished','dateCreated',
'series','status','title'):
pass # handled above, below or skip these for now, not going to do anything with them.
@@ -2003,7 +2005,9 @@ class FanFictionDownLoaderPlugin(InterfaceAction):
# cust cols can convert back to numbers and
# add.
book['anthology_meta_list'][k]=True
print("book['url']:%s"%book['url'])
configuration = get_ffdl_config(book['url'],fileform)
if existingbook:
book['title'] = deftitle = existingbook['title']
book['comments'] = existingbook['comments']
@@ -2022,7 +2026,6 @@ class FanFictionDownLoaderPlugin(InterfaceAction):
book['title'] = deftitle
break
configuration = get_ffdl_config(book['url'],fileform)
logger.debug("anthology_title_pattern:%s"%configuration.getConfig('anthology_title_pattern'))
if configuration.getConfig('anthology_title_pattern'):
tmplt = Template(configuration.getConfig('anthology_title_pattern'))
@@ -2038,7 +2041,7 @@ class FanFictionDownLoaderPlugin(InterfaceAction):
for v in ['Completed','In-Progress']:
if v in book['tags']:
book['tags'].remove(v)
book['tags'].append('Anthology')
book['tags'].extend(configuration.getConfigList('anthology_tags'))
book['all_metadata']['anthology'] = "true"
return book
+122 -9
View File
@@ -219,7 +219,32 @@ connect_timeout:60.0
# .*-Centered=>
# characters=>Sam W\.=>Sam Witwicky&&category=>Transformers
# characters=>Sam W\.=>Sam Winchester&&category=>Supernatural
## Include/Exclude metadata
##
## You can use the include/exclude metadata features to either limit
## the values of particular metadata lists to specific values or to
## exclude specific values. Further, you can conditionally apply each
## line depending on other metadata, use exact strings or regular
## expressions(regex) to match values, and negate matches.
##
## The settings are:
## include_metadata_pre
## exclude_metadata_pre
## include_metadata_post
## exclude_metadata_post
##
## The form of each line is:
## metakey[,metakey]==exactvalue
## metakey[,metakey]=~regex
## metakey[,metakey]==exactvalue&&conditionalkey==exactcondvalue
## metakey[,metakey]=~regex&&conditionalkey==exactcondvalue
## metakey[,metakey]==exactvalue&&conditionalkey=~condregex
##
## This is fairly complicated, so it's documented on its own wiki
## page:
## https://code.google.com/p/fanficdownloader/wiki/InExcludeMetadataFeature
## Some readers don't show horizontal rule (<hr />) tags correctly.
## This replaces them all with a centered '* * *'. (Note centering
## doesn't work on some devices either.)
@@ -601,6 +626,29 @@ extracategories:The Sentinel
## this should go in your personal.ini, not defaults.ini.
#is_adult:true
[bloodshedverse.com]
## website encoding(s) In theory, each website reports the character
## encoding they use for each page. In practice, some sites report it
## incorrectly. Each adapter has a default list, usually "utf8,
## Windows-1252" or "Windows-1252, utf8", but this will let you
## explicitly set the encoding and order if you need to. The special
## value 'auto' will call chardet and use the encoding it reports if
## it has +90% confidence. 'auto' is not reliable.
website_encodings:ISO-8859-1,auto
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:warnings,reviews
reviews_label:Reviews
## Site dedicated to these categories/characters/ships
extracharacters:Spike,Buffy
extracategories:Buffy the Vampire Slayer
## Strips links found in the story text
## Specific to bloodshedverse.com
strip_text_links:true
[bloodties-fans.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
@@ -742,6 +790,38 @@ extracategories:Harry Potter
## cover image. This lets you exclude them.
cover_exclusion_regexp:/images/.*?ribbon.gif
[fanfiction.csodaidok.hu]
## website encoding(s) In theory, each website reports the character
## encoding they use for each page. In practice, some sites report it
## incorrectly. Each adapter has a default list, usually "utf8,
## Windows-1252" or "Windows-1252, utf8", but this will let you
## explicitly set the encoding and order if you need to. The special
## value 'auto' will call chardet and use the encoding it reports if
## it has +90% confidence. 'auto' is not reliable.
website_encodings:ISO-8859-2,auto
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:reviews,challenge
reviews_label:Reviews
challenge_label:Challenge
## Site dedicated to these categories/characters/ships
extracategories:Harry Potter
[fanfic.hu]
## website encoding(s) In theory, each website reports the character
## encoding they use for each page. In practice, some sites report it
## incorrectly. Each adapter has a default list, usually "utf8,
## Windows-1252" or "Windows-1252, utf8", but this will let you
## explicitly set the encoding and order if you need to. The special
## value 'auto' will call chardet and use the encoding it reports if
## it has +90% confidence. 'auto' is not reliable.
website_encodings:ISO-8859-1,auto
## Site dedicated to these categories/characters/ships
extracategories:Harry Potter
[fanfiction.mugglenet.com]
## Some sites do not require a login, but do require the user to
## confirm they are adult for adult content. In commandline version,
@@ -792,6 +872,14 @@ extraships:Harry Potter/Hermione Granger
#username:YourName
#password:yourpassword
[ficwad.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
## commandline version, this should go in your personal.ini, not
## defaults.ini.
#username:YourName
#password:yourpassword
[fictionpad.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
@@ -948,6 +1036,17 @@ extracategories:NCIS
extracategories:Buffy: The Vampire Slayer
extracharacters:Willow
[nocturnal-light.net]
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:readings,reviews
readings_label:Readings
reviews_label:Reviews
## Site dedicated to these categories/characters/ships
extracharacters:Spike,Buffy
extracategories:Buffy the Vampire Slayer
[occlumency.sycophanthex.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
@@ -1038,6 +1137,16 @@ extracategories:Harry Potter
## this should go in your personal.ini, not defaults.ini.
#is_adult:true
[spikeluver.com]
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:warnings,reviews
reviews_label:Reviews
## Site dedicated to these categories/characters/ships
extracharacters:Spike,Buffy
extracategories:Buffy the Vampire Slayer
[stargate-atlantis.org]
## Site dedicated to these categories/characters/ships
extracategories:Stargate: Atlantis
@@ -1136,6 +1245,13 @@ awards_label:Awards
cover_exclusion_regexp:art/.*Awards.jpg
[voracity2.e-fic.com]
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:reviews,readings
reviews_label:Reviews
readings_label:Readings
[www.adastrafanfic.com]
## Some sites do not require a login, but do require the user to
## confirm they are adult for adult content. In commandline version,
@@ -1288,14 +1404,6 @@ extratags:
## for examples of how to use them.
extra_valid_entries:reviews,favs,follows
[ficwad.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
## commandline version, this should go in your personal.ini, not
## defaults.ini.
#username:YourName
#password:yourpassword
[www.fimfiction.net]
## Some sites do not require a login, but do require the user to
## confirm they are adult for adult content. In commandline version,
@@ -1314,6 +1422,11 @@ extra_valid_entries:reviews,favs,follows
## when updating to enforce accurate chapters.
#do_update_hook:false
## fimfiction.net is reported to misinterprete some BBCode with
## blockquotes incorrectly. This fixes those instances and defaults
## to on, but can be switched off if it is found to cause problems.
fix_fimf_blockquotes:true
## Site dedicated to these categories/characters/ships
extracategories:My Little Pony: Friendship is Magic
+2 -1
View File
@@ -23,6 +23,7 @@ import getpass
import string
import ConfigParser
from subprocess import call
import pprint
import logging
if sys.version_info >= (2, 7):
@@ -292,7 +293,7 @@ def main(argv,
else:
# regular download
if options.metaonly:
print adapter.getStoryMetadataOnly()
pprint.pprint(adapter.getStoryMetadataOnly().getAllMetadata())
output_filename=writeStory(configuration,adapter,options.format,options.metaonly)
Binary file not shown.
+7 -1
View File
@@ -123,6 +123,12 @@ import adapter_fictionpadcom
import adapter_storiesonlinenet
import adapter_trekiverseorg
import adapter_literotica
import adapter_voracity2eficcom
import adapter_spikeluvercom
import adapter_bloodshedversecom
import adapter_nocturnallightnet
import adapter_fanfichu
import adapter_fanfictioncsodaidokhu
## This bit of complexity allows adapters to be added by just adding
## importing. It eliminates the long if/else clauses we used to need
@@ -206,7 +212,7 @@ def getClassFor(url):
fixedurl = "http://%s"%url
## remove any trailing '#' locations.
fixedurl = re.sub(r"#.*$","",fixedurl)
parsedUrl = up.urlparse(fixedurl)
domain = parsedUrl.netloc.lower()
if( domain != parsedUrl.netloc ):
@@ -145,10 +145,13 @@ class ArchiveOfOurOwnOrgAdapter(BaseSiteAdapter):
except urllib2.HTTPError, e:
if e.code == 404:
raise exceptions.StoryDoesNotExist(self.meta)
raise exceptions.StoryDoesNotExist(self.url)
else:
raise e
if "Sorry, we couldn&#x27;t find the work you were looking for." in data:
raise exceptions.StoryDoesNotExist(self.url)
if self.needToLoginCheck(data):
# need to log in for this one.
self.performLogin(url,data)
@@ -0,0 +1,192 @@
from datetime import timedelta
import re
import urllib2
import urlparse
from .. import BeautifulSoup
from base_adapter import BaseSiteAdapter, makeDate
from .. import exceptions
def getClass():
return BloodshedverseComAdapter
def _get_query_data(url):
components = urlparse.urlparse(url)
query_data = urlparse.parse_qs(components.query)
return dict((key, data[0]) for key, data in query_data.items())
class BloodshedverseComAdapter(BaseSiteAdapter):
SITE_ABBREVIATION = 'bvc'
SITE_DOMAIN = 'bloodshedverse.com'
BASE_URL = 'http://' + SITE_DOMAIN + '/'
READ_URL_TEMPLATE = BASE_URL + 'stories.php?go=read&no=%s'
STARTED_DATETIME_FORMAT = '%m/%d/%Y'
UPDATED_DATETIME_FORMAT = '%m/%d/%Y %I:%M'
def __init__(self, config, url):
BaseSiteAdapter.__init__(self, config, url)
query_data = urlparse.parse_qs(self.parsedUrl.query)
story_no = query_data['no'][0]
self.story.setMetadata('storyId', story_no)
self._setURL(self.READ_URL_TEMPLATE % story_no)
self.story.setMetadata('siteabbrev', self.SITE_ABBREVIATION)
def _customized_fetch_url(self, url, exception=None, parameters=None):
if exception:
try:
data = self._fetchUrl(url, parameters)
except urllib2.HTTPError:
raise exception(self.url)
# Just let self._fetchUrl throw the exception, don't catch and
# customize it.
else:
data = self._fetchUrl(url, parameters)
return BeautifulSoup.BeautifulSoup(data)
@staticmethod
def getSiteDomain():
return BloodshedverseComAdapter.SITE_DOMAIN
@classmethod
def getSiteExampleURLs(cls):
return cls.READ_URL_TEMPLATE % 1234
def getSiteURLPattern(self):
return re.escape(self.BASE_URL + 'stories.php?go=') + r'(read|chapters)\&no=\d+$'
# Override stripURLParameters so the "no" parameter won't get stripped
@classmethod
def stripURLParameters(cls, url):
return url
def extractChapterUrlsAndMetadata(self):
soup = self._customized_fetch_url(self.url)
# Since no 404 error code we have to raise the exception ourselves.
# A title that is just 'by' indicates that there is no author name
# and no story title available.
if soup.title.string.strip() == 'by':
raise exceptions.StoryDoesNotExist(self.url)
for option in soup.find('select', {'name': 'chapter'}):
title = option.string.strip()
url = self.READ_URL_TEMPLATE % option['value']
self.chapterUrls.append((title, url))
# Get the URL to the author's page and find the correct story entry to
# scrape the metadata
author_url = urlparse.urljoin(self.url, soup.find('a', {'class': 'headline'})['href'])
soup = self._customized_fetch_url(author_url)
story_no = self.story.getMetadata('storyId')
# Ignore first list_box div, it only contains the author information
for list_box in soup('div', {'class': 'list_box'})[1:]:
url = list_box.find('a', {'class': 'fictitle'})['href']
query_data = _get_query_data(url)
# Found the div containing the story's metadata; break the loop and
# parse the element
if query_data['no'] == story_no:
break
else:
raise exceptions.FailedToDownload(self.url)
title_anchor = list_box.find('a', {'class': 'fictitle'})
self.story.setMetadata('title', title_anchor.string.strip())
author_anchor = title_anchor.findNextSibling('a')
self.story.setMetadata('author', author_anchor.string.strip())
self.story.setMetadata('authorId', _get_query_data(author_anchor['href'])['who'])
self.story.setMetadata('authorUrl', urlparse.urljoin(self.url, author_anchor['href']))
list_review = list_box.find('div', {'class': 'list_review'})
reviews = list_review.a.string.strip().split(' ', 1)[0]
self.story.setMetadata('reviews', reviews)
summary_div = list_box.find('div', {'class': 'list_summary'})
if not self.getConfig('keep_summary_html'):
summary = ''.join(summary_div(text=True))
else:
summary = self.utf8FromSoup(author_url, summary_div)
self.story.setMetadata('description', summary)
# I'm assuming this to be the category, not sure what else it could be
first_listinfo = list_box.find('div', {'class': 'list_info'})
self.story.addToList('category', first_listinfo.a.string.strip())
for list_info in first_listinfo.findNextSiblings('div', {'class': 'list_info'}):
for b_tag in list_info('b'):
key = b_tag.string.strip(': ')
# Strip colons from the beginning, superfluous spaces and minus
# characters from the end, and possibly trailing commas from
# the warnings if only one is present
value = b_tag.nextSibling.string.strip(': -,')
if key == 'Genre':
for genre in value.split(', '):
# Ignore the "none" genre
if not genre == 'none':
self.story.addToList('genre', genre)
elif key == 'Rating':
self.story.setMetadata('rating', value)
elif key == 'Complete':
self.story.setMetadata('status', 'Completed' if value == 'Yes' else 'In-Progress')
elif key == 'Warning':
for warning in value.split(', '):
# The string here starts with ", " before the actual list
# of values sometimes, so check for an empty warning
# and ignore the "none" warning.
if not warning or warning == 'none':
continue
self.story.addToList('warnings', warning)
elif key == 'Chapters':
self.story.setMetadata('numChapters', int(value))
elif key == 'Words':
# Apparently only numChapters need to be an integer for
# some strange reason. Remove possible ',' characters as to
# not confuse the codebase down the line
self.story.setMetadata('numWords', value.replace(',', ''))
elif key == 'Started':
self.story.setMetadata('datePublished', makeDate(value, self.STARTED_DATETIME_FORMAT))
elif key == 'Updated':
date_string, period = value.rsplit(' ', 1)
date = makeDate(date_string, self.UPDATED_DATETIME_FORMAT)
# Rather ugly hack to work around Calibre's changing of
# Python's locale setting, causing am/pm to not be properly
# parsed by strptime() when using a non-english locale
if period == 'pm':
date += timedelta(hours=12)
self.story.setMetadata('dateUpdated', date)
if self.story.getMetadata('rating') == 'NC-17' and not (self.is_adult or self.getConfig('is_adult')):
raise exceptions.AdultCheckRequired(self.url)
def getChapterText(self, url):
soup = self._customized_fetch_url(url)
storytext_div = soup.find('div', {'class': 'storytext'})
if self.getConfig('strip_text_links'):
for anchor in storytext_div('a', {'class': 'FAtxtL'}):
navigable_string = BeautifulSoup.NavigableString(anchor.string)
anchor.replaceWith(navigable_string)
return self.utf8FromSoup(url, storytext_div)
@@ -182,6 +182,11 @@ class DarkSolaceOrgAdapter(BaseSiteAdapter):
# first a tag in pagetitle is title
self.story.setMetadata('title',stripHTML(div.find('a')))
div.find('a').extract()
# only thing left in div(pagetitle) now should be 'by' and rating.
rating = stripHTML(div)
if '[' in rating:
self.story.setMetadata('rating', rating[rating.index('[')+1:-1])
for chapa in soup.findAll('a', href=re.compile(r'viewstory.php\?sid='+
self.story.getMetadata('storyId')+'&chapter=\d+')):
@@ -234,31 +239,28 @@ class DarkSolaceOrgAdapter(BaseSiteAdapter):
self.setDescription(url,svalue)
#self.story.setMetadata('description',stripHTML(svalue))
if 'Rated' in label:
self.story.setMetadata('rating', value[:len(value)-2])
if 'Word count' in label:
self.story.setMetadata('numWords', value)
if 'Categories' in label:
cats = labelspan.parent.findAll('a',href=re.compile(r'categories.php\?catid=\d+'))
cats = labelspan.parent.findAll('a',href=re.compile(r'browse.php\?type=categories'))
for cat in cats:
self.story.addToList('category',cat.string)
if 'Characters' in label:
for char in value.string.split(', '):
if not 'None' in char:
self.story.addToList('characters',char)
chars = labelspan.parent.findAll('a',href=re.compile(r'browse.php\?type=characters'))
for char in chars:
self.story.addToList('characters',char.string)
if 'Genre' in label:
for genre in value.string.split(', '):
if not 'None' in genre:
self.story.addToList('genre',genre)
genres = labelspan.parent.findAll('a',href=re.compile(r'browse.php\?type=class&type_id=1'))
for genre in genres:
self.story.addToList('genre',genre.string)
if 'Warnings' in label:
for warning in value.string.split(', '):
if not 'None' in warning:
self.story.addToList('warnings',warning)
warnings = labelspan.parent.findAll('a',href=re.compile(r'browse.php\?type=class&type_id=2'))
for warning in warnings:
self.story.addToList('warnings',warning.string)
if 'Completed' in label:
if 'Yes' in value:
@@ -0,0 +1,185 @@
# coding=utf-8
import re
import urllib2
import urlparse
from .. import BeautifulSoup
from base_adapter import BaseSiteAdapter, makeDate
from .. import exceptions
_SOURCE_CODE_ENCODING = 'utf-8'
def getClass():
return FanficHuAdapter
def _get_query_data(url):
components = urlparse.urlparse(url)
query_data = urlparse.parse_qs(components.query)
return dict((key, data[0]) for key, data in query_data.items())
class FanficHuAdapter(BaseSiteAdapter):
SITE_ABBREVIATION = 'ffh'
SITE_DOMAIN = 'fanfic.hu'
SITE_LANGUAGE = 'Hungarian'
BASE_URL = 'http://' + SITE_DOMAIN + '/merengo/'
VIEW_STORY_URL_TEMPLATE = BASE_URL + 'viewstory.php?sid=%s'
DATE_FORMAT = '%m/%d/%Y'
def __init__(self, config, url):
BaseSiteAdapter.__init__(self, config, url)
query_data = urlparse.parse_qs(self.parsedUrl.query)
story_id = query_data['sid'][0]
self.story.setMetadata('storyId', story_id)
self._setURL(self.VIEW_STORY_URL_TEMPLATE % story_id)
self.story.setMetadata('siteabbrev', self.SITE_ABBREVIATION)
self.story.setMetadata('language', self.SITE_LANGUAGE)
def _customized_fetch_url(self, url, exception=None, parameters=None):
if exception:
try:
data = self._fetchUrl(url, parameters)
except urllib2.HTTPError:
raise exception(self.url)
# Just let self._fetchUrl throw the exception, don't catch and
# customize it.
else:
data = self._fetchUrl(url, parameters)
return BeautifulSoup.BeautifulSoup(data)
@staticmethod
def getSiteDomain():
return FanficHuAdapter.SITE_DOMAIN
@classmethod
def getSiteExampleURLs(cls):
return cls.VIEW_STORY_URL_TEMPLATE % 1234
def getSiteURLPattern(self):
return re.escape(self.VIEW_STORY_URL_TEMPLATE[:-2]) + r'\d+$'
def extractChapterUrlsAndMetadata(self):
soup = self._customized_fetch_url(self.url + '&i=1')
if soup.title.string.encode(_SOURCE_CODE_ENCODING).strip(' :') == 'írta':
raise exceptions.StoryDoesNotExist(self.url)
chapter_options = soup.find('form', action='viewstory.php').select('option')
# Remove redundant "Fejezetek" option
chapter_options.pop(0)
# If there is still more than one entry remove chapter overview entry
if len(chapter_options) > 1:
chapter_options.pop(0)
for option in chapter_options:
url = urlparse.urljoin(self.url, option['value'])
self.chapterUrls.append((option.string, url))
author_url = urlparse.urljoin(self.BASE_URL, soup.find('a', href=lambda href: href and href.startswith('viewuser.php?uid='))['href'])
soup = self._customized_fetch_url(author_url)
story_id = self.story.getMetadata('storyId')
for table in soup('table', {'class': 'mainnav'}):
title_anchor = table.find('span', {'class': 'storytitle'}).a
href = title_anchor['href']
if href.startswith('javascript:'):
href = href.rsplit(' ', 1)[1].strip("'")
query_data = _get_query_data(href)
if query_data['sid'] == story_id:
break
else:
# This should never happen, the story must be found on the author's
# page.
raise exceptions.FailedToDownload(self.url)
self.story.setMetadata('title', title_anchor.string)
rows = table('tr')
anchors = rows[0].div('a')
author_anchor = anchors[1]
query_data = _get_query_data(author_anchor['href'])
self.story.setMetadata('author', author_anchor.string)
self.story.setMetadata('authorId', query_data['uid'])
self.story.setMetadata('authorUrl', urlparse.urljoin(self.BASE_URL, author_anchor['href']))
self.story.setMetadata('reviews', anchors[3].string)
if self.getConfig('keep_summary_html'):
self.story.setMetadata('description', self.utf8FromSoup(author_url, rows[1].td))
else:
self.story.setMetadata('description', ''.join(rows[1].td(text=True)))
for row in rows[3:]:
index = 0
cells = row('td')
while index < len(cells):
cell = cells[index]
key = cell.b.string.encode(_SOURCE_CODE_ENCODING).strip(':')
try:
value = cells[index+1].string.encode(_SOURCE_CODE_ENCODING)
except AttributeError:
value = None
if key == 'Kategória':
for anchor in cells[index+1]('a'):
self.story.addToList('category', anchor.string)
elif key == 'Szereplõk':
if cells[index+1].string:
for name in cells[index+1].string.split(', '):
self.story.addToList('character', name)
elif key == 'Korhatár':
if value != 'nem korhatáros':
self.story.setMetadata('rating', value)
elif key == 'Figyelmeztetések':
for b_tag in cells[index+1]('b'):
self.story.addToList('warnings', b_tag.string)
elif key == 'Jellemzõk':
for genre in cells[index+1].string.split(', '):
self.story.addToList('genre', genre)
elif key == 'Fejezetek':
self.story.setMetadata('numChapters', int(value))
elif key == 'Megjelenés':
self.story.setMetadata('datePublished', makeDate(value, self.DATE_FORMAT))
elif key == 'Frissítés':
self.story.setMetadata('dateUpdated', makeDate(value, self.DATE_FORMAT))
elif key == 'Szavak':
self.story.setMetadata('numWords', value)
elif key == 'Befejezett':
self.story.setMetadata('status', 'Completed' if value == 'Nem' else 'In-Progress')
index += 2
if self.story.getMetadata('rating') == '18':
if not (self.is_adult or self.getConfig('is_adult')):
raise exceptions.AdultCheckRequired(self.url)
def getChapterText(self, url):
soup = self._customized_fetch_url(url)
story_cell = soup.find('form', action='viewstory.php').parent.parent
for div in story_cell('div'):
div.extract()
return self.utf8FromSoup(url, story_cell)
@@ -0,0 +1,218 @@
# coding=utf-8
import re
import urllib2
import urlparse
from .. import BeautifulSoup
from base_adapter import BaseSiteAdapter, makeDate
from .. import exceptions
_SOURCE_CODE_ENCODING = 'utf-8'
def getClass():
return FanfictionCsodaidokHuAdapter
def _get_query_data(url):
components = urlparse.urlparse(url)
query_data = urlparse.parse_qs(components.query)
return dict((key, data[0]) for key, data in query_data.items())
# yields Tag _and_ NavigableString siblings from the given tag. The
# BeautifulSoup findNextSiblings() method for some reasons only returns either
# NavigableStrings _or_ Tag objects, not both.
def _yield_next_siblings(tag):
sibling = tag.nextSibling
while sibling:
yield sibling
sibling = sibling.nextSibling
class FanfictionCsodaidokHuAdapter(BaseSiteAdapter):
_SITE_DOMAIN = 'fanfiction.csodaidok.hu'
_BASE_URL = 'http://' + _SITE_DOMAIN + '/'
_VIEW_STORY_URL_TEMPLATE = _BASE_URL + 'viewstory.php?sid=%s'
_VIEW_CHAPTER_URL_TEMPLATE = _VIEW_STORY_URL_TEMPLATE + '&chapter=%s'
_STORY_DOES_NOT_EXIST_PAGE_TITLE = 'Cím: Szerző:'
_DATE_FORMAT = '%Y.%m.%d'
_SITE_LANGUAGE = 'Hungarian'
def __init__(self, config, url):
BaseSiteAdapter.__init__(self, config, url)
query_data = urlparse.parse_qs(self.parsedUrl.query)
story_id = query_data['sid'][0]
self.story.setMetadata('storyId', story_id)
self._setURL(self._VIEW_STORY_URL_TEMPLATE % story_id)
self.story.setMetadata('siteabbrev', self._SITE_DOMAIN)
self.story.setMetadata('language', self._SITE_LANGUAGE)
def _customized_fetch_url(self, url, exception=None, parameters=None):
if exception:
try:
data = self._fetchUrl(url, parameters)
except urllib2.HTTPError:
raise exception(self.url)
# Just let self._fetchUrl throw the exception, don't catch and
# customize it.
else:
data = self._fetchUrl(url, parameters)
return BeautifulSoup.BeautifulSoup(data)
@staticmethod
def getSiteDomain():
return FanfictionCsodaidokHuAdapter._SITE_DOMAIN
@classmethod
def getSiteExampleURLs(cls):
return cls._VIEW_STORY_URL_TEMPLATE % 1234
def getSiteURLPattern(self):
return re.escape(self._VIEW_STORY_URL_TEMPLATE[:-2]) + r'\d+$'
def extractChapterUrlsAndMetadata(self):
soup = self._customized_fetch_url(self.url + '&chapter=1')
element = soup.find('div', id='pagetitle')
page_title = ''.join(element(text=True)).encode(_SOURCE_CODE_ENCODING)
if page_title == self._STORY_DOES_NOT_EXIST_PAGE_TITLE:
raise exceptions.StoryDoesNotExist(self.url)
author_url = urlparse.urljoin(self.url, element.a['href'])
story_id = self.story.getMetadata('storyId')
element = soup.find('select', {'name': 'chapter'})
if element:
for option in element('option'):
title = option.string
url = self._VIEW_CHAPTER_URL_TEMPLATE % (story_id, option['value'])
self.chapterUrls.append((title, url))
soup = self._customized_fetch_url(author_url)
story_id = self.story.getMetadata('storyId')
for listbox_div in soup('div', {'class': lambda klass: klass and 'listbox' in klass}):
a = listbox_div.div.a
if not a['href'].startswith('viewstory.php?sid='):
continue
query_data = _get_query_data(a['href'])
if query_data['sid'] == story_id:
break
else:
raise exceptions.FailedToDownload(self.url)
title = ''.join(a(text=True))
self.story.setMetadata('title', title)
if not self.chapterUrls:
self.chapterUrls.append((title, self.url))
element = a.findNextSibling('a')
self.story.setMetadata('author', element.string)
query_data = _get_query_data(element['href'])
self.story.setMetadata('authorId', query_data['uid'])
self.story.setMetadata('authorUrl', author_url)
element = element.findNextSibling('span')
rating = element.nextSibling.strip(' [')
if rating.encode(_SOURCE_CODE_ENCODING) != 'Korhatár nélkül':
self.story.setMetadata('rating', rating)
if rating == '18':
raise exceptions.AdultCheckRequired(self.url)
element = element.findNextSiblings('a')[1]
self.story.setMetadata('reviews', element.string)
sections = listbox_div('div', {'class': lambda klass: klass and klass in ['content', 'tail']})
for section in sections:
for element in section('span', {'class': 'classification'}):
key = element.string.encode(_SOURCE_CODE_ENCODING).strip(' :')
try:
value = element.nextSibling.string.encode(_SOURCE_CODE_ENCODING).strip()
except AttributeError:
value = None
if key == 'Tartalom':
contents = []
keep_summary_html = self.getConfig('keep_summary_html')
for sibling in _yield_next_siblings(element):
if isinstance(sibling, BeautifulSoup.Tag):
if sibling.name == 'span' and sibling.get('class', None) == 'classification':
break
if keep_summary_html:
contents.append(self.utf8FromSoup(author_url, sibling))
else:
contents.append(''.join(sibling(text=True)))
else:
contents.append(sibling)
self.story.setMetadata('description', ''.join(contents))
elif key == 'Kategória':
for sibling in element.findNextSiblings(['a', 'span']):
if sibling.name == 'span':
break
self.story.addToList('category', sibling.string)
elif key == 'Szereplők':
for name in value.split(', '):
self.story.addToList('characters', name)
elif key == 'Műfaj':
if value != 'Nincs':
self.story.setMetadata('genre', value)
elif key == 'Figyelmeztetés':
if value != 'Nincs':
for warning in value.split(', '):
self.story.addToList('warnings', warning)
elif key == 'Kihívás':
if value != 'Nincs':
self.story.setMetadata('challenge', value)
elif key == 'Sorozat':
if value != 'Nincs':
self.story.setMetadata('series', value)
elif key == 'Fejezetek':
self.story.setMetadata('numChapters', int(value))
elif key == 'Befejezett':
self.story.setMetadata('status', 'Completed' if value == 'Nem' else 'In-Progress')
elif key == 'Szavak száma':
self.story.setMetadata('numWords', value)
elif key == 'Feltöltve':
self.story.setMetadata('datePublished', makeDate(value, self._DATE_FORMAT))
elif key == 'Frissítve':
self.story.setMetadata('dateUpdated', makeDate(value, self._DATE_FORMAT))
def getChapterText(self, url):
soup = self._customized_fetch_url(url)
contents = []
notes_div = soup.find('div', id='notes')
if notes_div:
contents.append(self.utf8FromSoup(url, notes_div))
story_div = notes_div.findNextSibling('div')
else:
element = soup.find('div', {'class': 'jumpmenu'})
story_div = element.findNextSibling('div')
contents.append(self.utf8FromSoup(url, story_div.span))
return ''.join(contents)
@@ -133,6 +133,7 @@ class FictionPadSiteAdapter(BaseSiteAdapter):
author = tables['users'][0]
story = tables['stories'][0]
story_ver = tables['story_versions'][0]
print("story:%s"%story)
self.story.setMetadata('authorId',author['id'])
self.story.setMetadata('author',author['display_name'])
@@ -151,7 +152,8 @@ class FictionPadSiteAdapter(BaseSiteAdapter):
self.story.setMetadata('comments',story['comments_count'])
self.story.setMetadata('views',story['views_count'])
self.story.setMetadata('likes',int(story['likes'])) # no idea why they floated these.
self.story.setMetadata('dislikes',int(story['dislikes']))
if 'dislikes' in story:
self.story.setMetadata('dislikes',int(story['dislikes']))
if story_ver['is_complete']:
self.story.setMetadata('status', 'Completed')
@@ -59,7 +59,7 @@ class FimFictionNetSiteAdapter(BaseSiteAdapter):
return "http://www.fimfiction.net/story/1234/story-title-here http://www.fimfiction.net/story/1234/ http://www.fimfiction.com/story/1234/1/ http://mobile.fimfiction.net/story/1234/1/story-title-here/chapter-title-here"
def getSiteURLPattern(self):
return r"http://(www|mobile)\.fimfiction\.(net|com)/story/\d+/?.*"
return r"https?://(www|mobile)\.fimfiction\.(net|com)/story/\d+/?.*"
def extractChapterUrlsAndMetadata(self):
@@ -85,7 +85,7 @@ class FimFictionNetSiteAdapter(BaseSiteAdapter):
# Unfortunately, we still need to load the story index
# page to parse the characters. And chapters, now, too.
data = self._fetchUrl(self.url)
data = self.do_fix_blockquotes(self._fetchUrl(self.url))
soup = bs.BeautifulSoup(data)
except urllib2.HTTPError, e:
if e.code == 404:
@@ -283,12 +283,21 @@ class FimFictionNetSiteAdapter(BaseSiteAdapter):
print("Existing epub has %s chapters\nNewest chapter is %s. Discarding old chapters from there on."%(len(self.oldchapters), self.newestChapterNum+1))
self.oldchapters = self.oldchapters[:self.newestChapterNum]
return len(self.oldchapters)
def do_fix_blockquotes(self,data):
if self.getConfig('fix_fimf_blockquotes'):
# <p class="double"><blockquote>
# </blockquote></p>
# include > in re groups so there's always something in the group.
data = re.sub(r'<p([^>]*>\s*)<blockquote([^>]*>)',r'<blockquote\2<p\1',data)
data = re.sub(r'</blockquote(>\s*)</p>',r'</p\1</blockquote>',data)
return data
def getChapterText(self, url):
logger.debug('Getting chapter text from: %s' % url)
soup = bs.BeautifulSoup(self._fetchUrl(url),selfClosingTags=('br','hr')).find('div', {'class' : 'chapter_content'})
data = self.do_fix_blockquotes(self._fetchUrl(url))
soup = bs.BeautifulSoup(data,selfClosingTags=('br','hr')).find('div', {'class' : 'chapter_content'})
if soup == None:
raise exceptions.FailedToDownload("Error downloading Chapter: %s! Missing required element!" % url)
return self.utf8FromSoup(url,soup)
+24 -14
View File
@@ -42,16 +42,19 @@ class LiteroticaSiteAdapter(BaseSiteAdapter):
self.story.setMetadata('siteabbrev','litero')
# get storyId from url--url validation guarantees query is only sid=1234
self.story.setMetadata('storyId',self.parsedUrl.path.split('/',)[2])
# normalize to first chapter. Not sure if they ever have more than 2 digits.
storyid = self.parsedUrl.path.split('/',)[2]
if re.match(r'-ch\d\d$',storyid):
storyid = storyid[:-2]+'01'
self.story.setMetadata('storyId',storyid)
self.origurl = url
if "http://www.i." in self.origurl:
if "//www.i." in self.origurl:
## accept m(mobile)url, but use www.
self.origurl = self.origurl.replace("http://www.i.","http://www.")
self.origurl = self.origurl.replace("//www.i.","//www.")
# normalized story URL.
self._setURL("http://"+self.getSiteDomain()\
self._setURL(url[:url.index('//')+2]+self.getSiteDomain()\
+"/s/"+self.story.getMetadata('storyId'))
# The date format will vary from site to site.
@@ -69,10 +72,10 @@ class LiteroticaSiteAdapter(BaseSiteAdapter):
@classmethod
def getSiteExampleURLs(self):
#return "http://www.literotica.com/s/story-title http://www.literotica.com/stories/showstory.php?id=1234 http://www.i.literotica.com/stories/showstory.php?id=1234"
return "http://www.literotica.com/s/story-title"
return "http://www.literotica.com/s/story-title https://www.literotica.com/s/story-title"
def getSiteURLPattern(self):
return r"http://www(\.i)?\.literotica\.com/s/([a-zA-Z0-9_-]+)"
return r"https?://www(\.i)?\.literotica\.com/s/([a-zA-Z0-9_-]+)"
def extractChapterUrlsAndMetadata(self):
@@ -97,20 +100,24 @@ class LiteroticaSiteAdapter(BaseSiteAdapter):
# author
a = soup1.find("span", "b-story-user-y")
self.story.setMetadata('authorId', urlparse.parse_qs(a.a['href'].split('?')[1])['uid'])
self.story.setMetadata('authorUrl', a.a['href'])
authorurl = a.a['href']
if authorurl.startswith('//'):
authorurl = self.parsedUrl.scheme+':'+authorurl
self.story.setMetadata('authorUrl', authorurl)
self.story.setMetadata('author', a.text)
# get the author page
try:
dataAuth = self._fetchUrl(a.a['href'])
dataAuth = self._fetchUrl(authorurl)
soupAuth = bs.BeautifulSoup(dataAuth)
except urllib2.HTTPError, e:
if e.code == 404:
raise exceptions.StoryDoesNotExist(a.a['href'])
raise exceptions.StoryDoesNotExist(authorurl)
else:
raise e
storyLink = soupAuth.find('a', href=url1)
## site has started using //domain.name/asdf urls remove https?: from front
storyLink = soupAuth.find('a', href=url1[url1.index(':')+1:])
if storyLink is not None:
# pull the published date from the author page
@@ -166,7 +173,10 @@ class LiteroticaSiteAdapter(BaseSiteAdapter):
self.story.setMetadata('datePublished',makeDate(stripHTML(row.find('td',{'class':'dt'})), self.dateformat))
while row['class'] == 'sl':
# pages include full URLs.
self.chapterUrls.append((row.a.string,row.a['href']))
chapurl = row.a['href']
if chapurl.startswith('//'):
chapurl = self.parsedUrl.scheme+':'+chapurl
self.chapterUrls.append((row.a.string,chapurl))
if not row.nextSibling:
break
row = row.nextSibling
@@ -203,7 +213,7 @@ class LiteroticaSiteAdapter(BaseSiteAdapter):
# get story text
story1 = soup1.find('div', 'b-story-body-x').p
story1.name='div'
story1.append('<br>')
story1.append('<br />')
storytext = self.utf8FromSoup(url,story1)
# find num pages
@@ -220,7 +230,7 @@ class LiteroticaSiteAdapter(BaseSiteAdapter):
[comment.extract() for comment in soup2.findAll(text=lambda text:isinstance(text, bs.Comment))]
story2 = soup2.find('div', 'b-story-body-x').p
story2.name='div'
story2.append('<br>')
story2.append('<br />')
storytext += self.utf8FromSoup(url,story2)
except urllib2.HTTPError, e:
if e.code == 404:
@@ -0,0 +1,177 @@
import re
import urllib2
import urlparse
from .. import BeautifulSoup
from base_adapter import BaseSiteAdapter, makeDate
from .. import exceptions
def getClass():
return NocturnalLightNetAdapter
# yields Tag _and_ NavigableString siblings from the given tag. The
# BeautifulSoup findNextSiblings() method for some reasons only returns either
# NavigableStrings _or_ Tag objects, not both.
def _yield_next_siblings(tag):
sibling = tag.nextSibling
while sibling:
yield sibling
sibling = sibling.nextSibling
class NocturnalLightNetAdapter(BaseSiteAdapter):
SITE_ABBREVIATION = 'nln'
SITE_DOMAIN = 'nocturnal-light.net'
BASE_URL = 'http://' + SITE_DOMAIN + '/fanfiction/'
STORY_URL_TEMPLATE = BASE_URL + 'story/%s'
AUTHORS_URL_TEMPLATE = BASE_URL + 'authors/%s'
DATETIME_FORMAT = '%m-%d-%y'
def __init__(self, config, url):
BaseSiteAdapter.__init__(self, config, url)
url_tokens = self.parsedUrl.path.split('/')
story_id = url_tokens[url_tokens.index('story') + 1]
self.story.setMetadata('storyId', story_id)
self._setURL(self.STORY_URL_TEMPLATE % story_id)
self.story.setMetadata('siteabbrev', self.SITE_ABBREVIATION)
def _customized_fetch_url(self, url, exception=None, parameters=None):
if exception:
try:
data = self._fetchUrl(url, parameters)
except urllib2.HTTPError:
raise exception(self.url)
# Just let self._fetchUrl throw the exception, don't catch and
# customize it.
else:
data = self._fetchUrl(url, parameters)
return BeautifulSoup.BeautifulSoup(data)
@staticmethod
def getSiteDomain():
return NocturnalLightNetAdapter.SITE_DOMAIN
@classmethod
def getSiteExampleURLs(cls):
return cls.STORY_URL_TEMPLATE % 1234
def getSiteURLPattern(self):
return re.escape(self.STORY_URL_TEMPLATE[:-2]) + r'\d+.*$'
def extractChapterUrlsAndMetadata(self):
soup = self._customized_fetch_url(self.url)
# Since no 404 error code we have to raise the exception ourselves.
# A title that is just 'by' indicates that there is no author name
# and no story title available.
if soup.title.string.strip() == 'by':
raise exceptions.StoryDoesNotExist(self.url)
# "storycontent" is found in a single-chapter story
author_anchor = soup.find('div', id=lambda id: id in ('main', 'storycontent')).h1.a
self.story.setMetadata('author', author_anchor.string)
url_tokens = author_anchor['href'].split('/')
author_id = url_tokens[url_tokens.index('authors')+1]
self.story.setMetadata('authorId', author_id)
self.story.setMetadata('authorUrl', self.AUTHORS_URL_TEMPLATE % author_id)
chapter_anchors = soup('a', href=lambda href: href and href.startswith('/fanfiction/story/'))
for chapter_anchor in chapter_anchors:
url = urlparse.urljoin(self.BASE_URL, chapter_anchor['href'])
self.chapterUrls.append((chapter_anchor.string, url))
author_url = urlparse.urljoin(self.BASE_URL, author_anchor['href'])
soup = self._customized_fetch_url(author_url)
story_id = self.story.getMetadata('storyId')
for listbox in soup('div', {'class': 'listbox'}):
url_tokens = listbox.a['href'].split('/')
# Found the div containing the story's metadata; break the loop and
# parse the element
if story_id == url_tokens[url_tokens.index('story')+1]:
break
else:
raise exceptions.FailedToDownload(self.url)
title = listbox.a.string
self.story.setMetadata('title', title)
# No chapter anchors found in the original story URL, so the story has
# only a single chapter.
if not chapter_anchors:
self.chapterUrls.append((title, self.url))
for b_tag in listbox('b'):
key = b_tag.string.strip(':')
try:
value = b_tag.nextSibling.string.replace('&bull;', '').strip(': ')
# This can happen with some fancy markup in the summary. Just
# ignore this error and set value to None, the summary parsing
# takes care of this
except AttributeError:
value = None
if key == 'Summary':
contents = []
keep_summary_html = self.getConfig('keep_summary_html')
for sibling in _yield_next_siblings(b_tag):
if isinstance(sibling, BeautifulSoup.Tag):
if sibling.name == 'b' and sibling.findPreviousSibling().name == 'br':
break
if keep_summary_html:
contents.append(self.utf8FromSoup(author_url, sibling))
else:
contents.append(''.join(sibling(text=True)))
else:
contents.append(sibling)
# Pop last break line tag
contents.pop()
self.story.setMetadata('description', ''.join(contents))
elif key == 'Category':
for sibling in b_tag.findNextSiblings(['a', 'b']):
if sibling.name == 'b':
break
self.story.addToList('category', sibling.string)
elif key == 'Rating':
self.story.setMetadata('rating', value)
elif key == 'Chapters':
self.story.setMetadata('numChapters', int(value))
# Also parse reviews number which lies right after the chapters
# section
reviews_anchor = b_tag.findNextSibling('a')
reviews = reviews_anchor.string.split(' ')[1].strip('()')
self.story.setMetadata('reviews', reviews)
elif key == 'Completed':
self.story.setMetadata('status', 'Completed' if value == 'Yes' else 'In-Progress')
elif key == 'Date Added':
self.story.setMetadata('datePublished', makeDate(value, self.DATETIME_FORMAT))
elif key == 'Last Updated':
self.story.setMetadata('dateUpdated', makeDate(value, self.DATETIME_FORMAT))
elif key == 'Read':
self.story.setMetadata('readings', value.split()[0])
if self.story.getMetadata('rating') == 'NC-17' and not (self.is_adult or self.getConfig('is_adult')):
raise exceptions.AdultCheckRequired(self.url)
def getChapterText(self, url):
soup = self._customized_fetch_url(url)
return self.utf8FromSoup(url, soup.find('div', id='storytext'))
@@ -192,7 +192,7 @@ class OneDirectionFanfictionComAdapter(BaseSiteAdapter):
if 'Summary' in label:
## Everything until the next span class='label'
svalue = ""
while not defaultGetattr(value,'class') == 'label':
while value and not defaultGetattr(value,'class') == 'label':
svalue += str(value)
value = value.nextSibling
self.setDescription(url,svalue)
@@ -0,0 +1,206 @@
import re
import urllib2
import urlparse
from .. import BeautifulSoup
from base_adapter import BaseSiteAdapter, makeDate
from .. import exceptions
def getClass():
return SpikeluverComAdapter
# yields Tag _and_ NavigableString siblings from the given tag. The
# BeautifulSoup findNextSiblings() method for some reasons only returns either
# NavigableStrings _or_ Tag objects, not both.
def _yield_next_siblings(tag):
sibling = tag.nextSibling
while sibling:
yield sibling
sibling = sibling.nextSibling
class SpikeluverComAdapter(BaseSiteAdapter):
SITE_ABBREVIATION = 'slc'
SITE_DOMAIN = 'spikeluver.com'
BASE_URL = 'http://' + SITE_DOMAIN + '/SpuffyRealm/'
LOGIN_URL = BASE_URL + 'user.php?action=login'
VIEW_STORY_URL_TEMPLATE = BASE_URL + 'viewstory.php?sid=%d'
METADATA_URL_SUFFIX = '&index=1'
AGE_CONSENT_URL_SUFFIX = '&ageconsent=ok&warning=5'
DATETIME_FORMAT = '%m/%d/%Y'
STORY_DOES_NOT_EXIST_ERROR_TEXT = 'That story does not exist on this archive. You may search for it or return to the home page.'
def __init__(self, config, url):
BaseSiteAdapter.__init__(self, config, url)
query_data = urlparse.parse_qs(self.parsedUrl.query)
story_id = query_data['sid'][0]
self.story.setMetadata('storyId', story_id)
self._setURL(self.VIEW_STORY_URL_TEMPLATE % int(story_id))
self.story.setMetadata('siteabbrev', self.SITE_ABBREVIATION)
def _customized_fetch_url(self, url, exception=None, parameters=None):
if exception:
try:
data = self._fetchUrl(url, parameters)
except urllib2.HTTPError:
raise exception(self.url)
# Just let self._fetchUrl throw the exception, don't catch and
# customize it.
else:
data = self._fetchUrl(url, parameters)
return BeautifulSoup.BeautifulSoup(data)
@staticmethod
def getSiteDomain():
return SpikeluverComAdapter.SITE_DOMAIN
@classmethod
def getSiteExampleURLs(cls):
return cls.VIEW_STORY_URL_TEMPLATE % 1234
def getSiteURLPattern(self):
return re.escape(self.VIEW_STORY_URL_TEMPLATE[:-2]) + r'\d+$'
def extractChapterUrlsAndMetadata(self):
soup = self._customized_fetch_url(self.url + self.METADATA_URL_SUFFIX)
errortext_div = soup.find('div', {'class': 'errortext'})
if errortext_div:
error_text = ''.join(errortext_div(text=True)).strip()
if error_text == self.STORY_DOES_NOT_EXIST_ERROR_TEXT:
raise exceptions.StoryDoesNotExist(self.url)
# No additional login is required, just check for adult
pagetitle_div = soup.find('div', id='pagetitle')
if pagetitle_div.a['href'].startswith('javascript:'):
if not(self.is_adult or self.getConfig('is_adult')):
raise exceptions.AdultCheckRequired(self.url)
url = ''.join([self.url, self.METADATA_URL_SUFFIX, self.AGE_CONSENT_URL_SUFFIX])
soup = self._customized_fetch_url(url)
pagetitle_div = soup.find('div', id='pagetitle')
self.story.setMetadata('title', pagetitle_div.a.string.strip())
author_anchor = pagetitle_div.a.findNextSibling('a')
url = urlparse.urljoin(self.BASE_URL, author_anchor['href'])
components = urlparse.urlparse(url)
query_data = urlparse.parse_qs(components.query)
self.story.setMetadata('author', author_anchor.string.strip())
self.story.setMetadata('authorId', query_data['uid'])
self.story.setMetadata('authorUrl', url)
sort_div = soup.find('div', id='sort')
self.story.setMetadata('reviews', sort_div('a')[1].string.strip())
listbox_tag = soup.find('div', {'class': 'listbox'})
for span_tag in listbox_tag('span'):
key = span_tag.string.strip(' :')
try:
value = span_tag.nextSibling.string.strip()
# This can happen with some fancy markup in the summary. Just
# ignore this error and set value to None, the summary parsing
# takes care of this
except AttributeError:
value = None
if key == 'Summary':
contents = []
keep_summary_html = self.getConfig('keep_summary_html')
for sibling in _yield_next_siblings(span_tag):
if isinstance(sibling, BeautifulSoup.Tag):
# Encountered next label, break. Not as bad as other
# e-fiction sites, let's hope this is enough for proper
# parsing.
if sibling.name == 'span' and sibling.get('class', None) == 'label':
break
if keep_summary_html:
contents.append(self.utf8FromSoup(self.url, sibling))
else:
contents.append(''.join(sibling(text=True)))
else:
contents.append(sibling)
# Remove the preceding break line tag and other crud
contents.pop()
contents.pop()
self.story.setMetadata('description', ''.join(contents))
elif key == 'Rated':
self.story.setMetadata('rating', value)
elif key == 'Categories':
for sibling in span_tag.findNextSiblings(['a', 'br']):
if sibling.name == 'br':
break
self.story.addToList('category', sibling.string.strip())
# Seems to be always "None" for some reason
elif key == 'Characters':
for sibling in span_tag.findNextSiblings(['a', 'br']):
if sibling.name == 'br':
break
self.story.addToList('characters', sibling.string.strip())
elif key == 'Genres':
for sibling in span_tag.findNextSiblings(['a', 'br']):
if sibling.name == 'br':
break
self.story.addToList('genre', sibling.string.strip())
elif key == 'Warnings':
for sibling in span_tag.findNextSiblings(['a', 'br']):
if sibling.name == 'br':
break
self.story.addToList('warnings', sibling.string.strip())
# Challenges
elif key == 'Series':
a = span_tag.findNextSibling('a')
if not a:
continue
self.story.setMetadata('series', a.string.strip())
self.story.setMetadata('seriesUrl', urlparse.urljoin(self.BASE_URL, a['href']))
elif key == 'Chapters':
self.story.setMetadata('numChapters', int(value))
elif key == 'Completed':
self.story.setMetadata('status', 'Completed' if value == 'Yes' else 'In-Progress')
elif key == 'Word count':
self.story.setMetadata('numWords', value)
elif key == 'Published':
self.story.setMetadata('datePublished', makeDate(value, self.DATETIME_FORMAT))
elif key == 'Updated':
self.story.setMetadata('dateUpdated', makeDate(value, self.DATETIME_FORMAT))
for p_tag in listbox_tag.findNextSiblings('p'):
chapter_anchor = p_tag.find('a', href=lambda href: href and href.startswith('viewstory.php?sid='))
if not chapter_anchor:
continue
title = chapter_anchor.string.strip()
url = urlparse.urljoin(self.BASE_URL, chapter_anchor['href'])
self.chapterUrls.append((title, url))
def getChapterText(self, url):
url += self.AGE_CONSENT_URL_SUFFIX
soup = self._customized_fetch_url(url)
return self.utf8FromSoup(url, soup.find('div', id='story'))
@@ -62,7 +62,7 @@ class SquidgeOrgPejaAdapter(BaseSiteAdapter):
# normalized story URL.
self._setURL('http://' + self.getSiteDomain() + '/peja/cgi-bin/viewstory.php?sid='+self.story.getMetadata('storyId'))
self._setURL('https://' + self.getSiteDomain() + '/peja/cgi-bin/viewstory.php?sid='+self.story.getMetadata('storyId'))
# Each adapter needs to have a unique site abbreviation.
self.story.setMetadata('siteabbrev','wwomb')
@@ -83,10 +83,10 @@ class SquidgeOrgPejaAdapter(BaseSiteAdapter):
@classmethod
def getSiteExampleURLs(self):
return "http://"+self.getSiteDomain()+"/peja/cgi-bin/viewstory.php?sid=1234"
return "https://"+self.getSiteDomain()+"/peja/cgi-bin/viewstory.php?sid=1234"
def getSiteURLPattern(self):
return re.escape("http://"+self.getSiteDomain()+"/")+"~?"+re.escape("peja/cgi-bin/viewstory.php?sid=")+r"\d+$"
return r"https?"+re.escape("://"+self.getSiteDomain()+"/")+r"~?"+re.escape("peja/cgi-bin/viewstory.php?sid=")+r"\d+$"
## Getting the chapter list and the meta data, plus 'is adult' checking.
def extractChapterUrlsAndMetadata(self):
@@ -116,7 +116,7 @@ class SquidgeOrgPejaAdapter(BaseSiteAdapter):
# Find authorid and URL from... author url.
author = soup.find('div', {'id':"pagetitle"}).find('a')
self.story.setMetadata('authorId',author['href'].split('=')[1])
self.story.setMetadata('authorUrl','http://'+self.host+'/peja/cgi-bin/'+author['href'])
self.story.setMetadata('authorUrl','https://'+self.host+'/peja/cgi-bin/'+author['href'])
self.story.setMetadata('author',author.string)
authorSoup = bs.BeautifulSoup(self._fetchUrl(self.story.getMetadata('authorUrl')))
@@ -131,7 +131,7 @@ class SquidgeOrgPejaAdapter(BaseSiteAdapter):
chapterselect=soup.find('select',{'name':'chapter'})
if chapterselect:
for ch in chapterselect.findAll('option'):
self.chapterUrls.append((stripHTML(ch),'http://'+self.host+'/peja/cgi-bin/viewstory.php?sid='+self.story.getMetadata('storyId')+'&chapter='+ch['value']))
self.chapterUrls.append((stripHTML(ch),'https://'+self.host+'/peja/cgi-bin/viewstory.php?sid='+self.story.getMetadata('storyId')+'&chapter='+ch['value']))
else:
self.chapterUrls.append((title,url))
@@ -207,7 +207,7 @@ class SquidgeOrgPejaAdapter(BaseSiteAdapter):
# http://www.squidge.org/peja/cgi-bin/series.php?seriesid=254
a = titleblock.find('a', href=re.compile(r"series.php\?seriesid=\d+"))
series_name = a.string
series_url = 'http://'+self.host+'/peja/cgi-bin/'+a['href']
series_url = 'https://'+self.host+'/peja/cgi-bin/'+a['href']
# use BeautifulSoup HTML parser to make everything easier to find.
seriessoup = bs.BeautifulSoup(self._fetchUrl(series_url))
@@ -66,7 +66,7 @@ class StoriesOnlineNetAdapter(BaseSiteAdapter):
return "http://"+self.getSiteDomain()+"/s/1234 http://"+self.getSiteDomain()+"/s/1234:4010"
def getSiteURLPattern(self):
return re.escape("http://"+self.getSiteDomain())+r"/s/\d+((:\d+)?(;\d+)?$|(:i)?$)"
return re.escape("http://"+self.getSiteDomain())+r"/s/\d+((:\d+)?(;\d+)?$|(:i)?$)?"
## Login seems to be reasonably standard across eFiction sites.
def needToLoginCheck(self, data):
@@ -171,7 +171,7 @@ class StoriesOnlineNetAdapter(BaseSiteAdapter):
a = asoup.findAll('td', {'class' : 'lc2'})
for lc2 in a:
if lc2.find('a')['href'] == '/s/'+self.story.getMetadata('storyId'):
if lc2.find('a', href=re.compile(r'^/s/'+self.story.getMetadata('storyId'))):
i=1
break
if a[len(a)-1] == lc2:
@@ -0,0 +1,232 @@
import re
import urllib2
import urlparse
from .. import BeautifulSoup
from base_adapter import BaseSiteAdapter, makeDate
from .. import exceptions
def getClass():
return Voracity2EficComAdapter
# yields Tag _and_ NavigableString siblings from the given tag. The
# BeautifulSoup findNextSiblings() method for some reasons only returns either
# NavigableStrings _or_ Tag objects, not both.
def _yield_next_siblings(tag):
sibling = tag.nextSibling
while sibling:
yield sibling
sibling = sibling.nextSibling
class Voracity2EficComAdapter(BaseSiteAdapter):
SITE_ABBREVIATION = 'voe'
SITE_DOMAIN = 'voracity2.e-fic.com'
BASE_URL = 'http://' + SITE_DOMAIN + '/'
LOGIN_URL = BASE_URL + 'user.php?action=login'
VIEW_STORY_URL_TEMPLATE = BASE_URL + 'viewstory.php?sid=%d'
METADATA_URL_SUFFIX = '&index=1'
AGE_CONSENT_URL_SUFFIX = '&ageconsent=ok&warning=4'
DATETIME_FORMAT = '%m/%d/%Y'
REQUIRED_SKIN = 'Simple Elegance'
def __init__(self, config, url):
BaseSiteAdapter.__init__(self, config, url)
query_data = urlparse.parse_qs(self.parsedUrl.query)
story_id = query_data['sid'][0]
self.story.setMetadata('storyId', story_id)
self._setURL(self.VIEW_STORY_URL_TEMPLATE % int(story_id))
self.story.setMetadata('siteabbrev', self.SITE_ABBREVIATION)
self.is_logged_in = False
def _login(self):
# Apparently self.password is only set when login fails, i.e.
# the FailedToLogin exception is raised, so the adapter gets new
# login data and tries again
if self.password:
password = self.password
username = self.username
else:
username = self.getConfig('username')
password = self.getConfig('password')
parameters = {
'penname': username,
'password': password,
'submit': 'Submit'}
class CustomizedFailedToLogin(exceptions.FailedToLogin):
def __init__(self, url, passwdonly=False):
# Use username variable from outer scope
exceptions.FailedToLogin.__init__(self, url, username, passwdonly)
soup = self._customized_fetch_url(self.LOGIN_URL, CustomizedFailedToLogin, parameters)
div = soup.find('div', id='useropts')
if not div:
raise CustomizedFailedToLogin(self.LOGIN_URL)
self.is_logged_in = True
def _customized_fetch_url(self, url, exception=None, parameters=None):
if exception:
try:
data = self._fetchUrl(url, parameters)
except urllib2.HTTPError:
raise exception(self.url)
# Just let self._fetchUrl throw the exception, don't catch and
# customize it.
else:
data = self._fetchUrl(url, parameters)
return BeautifulSoup.BeautifulSoup(data)
@staticmethod
def getSiteDomain():
return Voracity2EficComAdapter.SITE_DOMAIN
@classmethod
def getSiteExampleURLs(cls):
return cls.VIEW_STORY_URL_TEMPLATE % 1234
def getSiteURLPattern(self):
return re.escape(self.VIEW_STORY_URL_TEMPLATE[:-2]) + r'\d+$'
def extractChapterUrlsAndMetadata(self):
soup = self._customized_fetch_url(self.url + self.METADATA_URL_SUFFIX)
# Check if the story is for "Registered Users Only", i.e. has adult
# content. Based on the "is_adult" attributes either login or raise an
# error.
errortext_div = soup.find('div', {'class': 'errortext'})
if errortext_div:
error_text = ''.join(errortext_div(text=True)).strip()
if error_text == 'Registered Users Only':
if not (self.is_adult or self.getConfig('is_adult')):
raise exceptions.AdultCheckRequired(self.url)
self._login()
else:
# This case usually occurs when the story doesn't exist, but
# might potentially be something else, so just raise
# FailedToDownload exception with the found error text.
raise exceptions.FailedToDownload(error_text)
url = ''.join([self.url, self.METADATA_URL_SUFFIX, self.AGE_CONSENT_URL_SUFFIX])
soup = self._customized_fetch_url(url)
# If logged in and the skin doesn't match the required skin throw an
# error
if self.is_logged_in:
skin = soup.find('select', {'name': 'skin'}).find('option', selected=True)['value']
if skin != self.REQUIRED_SKIN:
raise exceptions.FailedToDownload('Required skin "%s" must be set in preferences' % self.REQUIRED_SKIN)
pagetitle_div = soup.find('div', id='pagetitle')
self.story.setMetadata('title', pagetitle_div.a.string)
author_anchor = pagetitle_div.a.findNextSibling('a')
url = urlparse.urljoin(self.BASE_URL, author_anchor['href'])
components = urlparse.urlparse(url)
query_data = urlparse.parse_qs(components.query)
self.story.setMetadata('author', author_anchor.string)
self.story.setMetadata('authorId', query_data['uid'])
self.story.setMetadata('authorUrl', url)
sort_div = soup.find('div', id='sort')
self.story.setMetadata('reviews', sort_div('a')[1].string)
for b_tag in soup.find('div', {'class': 'listbox'})('b'):
key = b_tag.string.strip(' :')
try:
value = b_tag.nextSibling.string.strip()
# This can happen with some fancy markup in the summary. Just
# ignore this error and set value to None, the summary parsing
# takes care of this
except AttributeError:
value = None
if key == 'Summary':
contents = []
keep_summary_html = self.getConfig('keep_summary_html')
for sibling in _yield_next_siblings(b_tag):
if isinstance(sibling, BeautifulSoup.Tag):
# Encountered next label, break. This method is the
# safest and most reliable I could think of. Blame
# e-fiction sites that allow their users to include
# arbitrary markup into their summaries and the
# horrible HTML markup.
if sibling.name == 'b' and sibling.findPreviousSibling().name == 'br':
break
if keep_summary_html:
contents.append(self.utf8FromSoup(self.url, sibling))
else:
contents.append(''.join(sibling(text=True)))
else:
contents.append(sibling)
# Remove the preceding break line tag and other crud
contents.pop()
contents.pop()
self.story.setMetadata('description', ''.join(contents))
elif key == 'Rating':
self.story.setMetadata('rating', value)
elif key == 'Category':
for sibling in b_tag.findNextSiblings(['a', 'br']):
if sibling.name == 'br':
break
self.story.addToList('category', sibling.string)
# Seems to be always "None" for some reason
elif key == 'Characters':
for sibling in b_tag.findNextSiblings(['a', 'br']):
if sibling.name == 'br':
break
self.story.addToList('characters', sibling.string)
elif key == 'Series':
a = b_tag.findNextSibling('a')
if not a:
continue
self.story.setMetadata('series', a.string)
self.story.setMetadata('seriesUrl', urlparse.urljoin(self.BASE_URL, a['href']))
elif key == 'Chapter':
self.story.setMetadata('numChapters', int(value))
elif key == 'Completed':
self.story.setMetadata('status', 'Completed' if value == 'Yes' else 'In-Progress')
elif key == 'Words':
self.story.setMetadata('numWords', value)
elif key == 'Read':
self.story.setMetadata('readings', value)
elif key == 'Published':
self.story.setMetadata('datePublished', makeDate(value, self.DATETIME_FORMAT))
elif key == 'Updated':
self.story.setMetadata('dateUpdated', makeDate(value, self.DATETIME_FORMAT))
for b_tag in soup.find('div', id='output').findNextSiblings('b'):
chapter_anchor = b_tag.a
title = chapter_anchor.string
url = urlparse.urljoin(self.BASE_URL, chapter_anchor['href'])
self.chapterUrls.append((title, url))
def getChapterText(self, url):
url += self.AGE_CONSENT_URL_SUFFIX
soup = self._customized_fetch_url(url)
return self.utf8FromSoup(url, soup.find('div', id='story'))
+1 -1
View File
@@ -188,9 +188,9 @@ class BaseSiteAdapter(Configurable):
try:
return self._decode(self._fetchUrlRaw(url,parameters))
except u2.HTTPError, he:
excpt=he
if he.code == 404:
logger.warn("Caught an exception reading URL: %s Exception %s."%(unicode(url),unicode(he)))
excpt=he
break # break out on 404
except Exception, e:
excpt=e
+144 -7
View File
@@ -28,6 +28,9 @@ import exceptions
from htmlcleanup import conditionalRemoveEntities, removeAllEntities
from configurable import Configurable
SPACE_REPLACE=u'\s'
SPLIT_META=u'\,'
# Create convert_image method depending on which graphics lib we can
# load. Preferred: calibre, PIL, none
@@ -221,6 +224,65 @@ langs = {
"Devanagari":"hi",
}
class InExMatch:
keys = []
regex = None
match = None
negate = False
def __init__(self,line):
if "=~" in line:
(self.keys,self.match) = line.split("=~")
self.match = self.match.replace(SPACE_REPLACE,' ')
self.regex = re.compile(self.match)
elif "!~" in line:
(self.keys,self.match) = line.split("!~")
self.match = self.match.replace(SPACE_REPLACE,' ')
self.regex = re.compile(self.match)
self.negate = True
elif "==" in line:
(self.keys,self.match) = line.split("==")
self.match = self.match.replace(SPACE_REPLACE,' ')
elif "!=" in line:
(self.keys,self.match) = line.split("!=")
self.match = self.match.replace(SPACE_REPLACE,' ')
self.negate = True
self.keys = map( lambda x: x.strip(), self.keys.split(",") )
# For conditional, only one key
def is_key(self,key):
return key == self.keys[0]
# For conditional, only one key
def key(self):
return self.keys[0]
def in_keys(self,key):
return key in self.keys
def is_match(self,value):
retval = False
if self.regex:
if self.regex.search(value):
retval = True
#print(">>>>>>>>>>>>>%s=~%s r: %s,%s=%s"%(self.match,value,self.negate,retval,self.negate != retval))
else:
retval = self.match == value
#print(">>>>>>>>>>>>>%s==%s r: %s,%s=%s"%(self.match,value,self.negate,retval, self.negate != retval))
return self.negate != retval
def __str__(self):
if self.negate:
f='!'
else:
f='='
if self.regex:
s='~'
else:
s='='
return u'InExMatch(%s %s%s %s)'%(self.keys,f,s,self.match)
class Story(Configurable):
def __init__(self, configuration):
@@ -231,6 +293,7 @@ class Story(Configurable):
except:
self.metadata = {'version':'4.4'}
self.replacements = []
self.in_ex_cludes = {}
self.chapters = [] # chapters will be tuples of (title,html)
self.imgurls = []
self.imgtuples = []
@@ -241,8 +304,7 @@ class Story(Configurable):
self.logfile=None # cheesy way to carry log file forward across update.
## Look for config parameter, split and add each to metadata field.
for (config,metadata) in [("extratags","extratags"),
("extracategories","category"),
for (config,metadata) in [("extracategories","category"),
("extragenres","genre"),
("extracharacters","characters"),
("extraships","ships"),
@@ -252,6 +314,15 @@ class Story(Configurable):
self.setReplace(self.getConfig('replace_metadata'))
in_ex_clude_list = ['include_metadata_pre','exclude_metadata_pre',
'include_metadata_post','exclude_metadata_post']
for ie in in_ex_clude_list:
ies = self.getConfig(ie)
# print("%s %s"%(ie,ies))
if ies:
iel = []
self.in_ex_cludes[ie] = self.set_in_ex_clude(ies)
def setMetadata(self, key, value, condremoveentities=True):
## still keeps &lt; &lt; and &amp;
if condremoveentities:
@@ -269,6 +340,56 @@ class Story(Configurable):
self.addToList('lastupdate',value.strftime("Last Update Year/Month: %Y/%m"))
self.addToList('lastupdate',value.strftime("Last Update: %Y/%m/%d"))
## metakey[,metakey]=~pattern
## metakey[,metakey]==string
## *for* part lines. Effect only when trailing conditional key=~regexp matches
## metakey[,metakey]=~pattern[&&metakey=~regexp]
## metakey[,metakey]==string[&&metakey=~regexp]
## metakey[,metakey]=~pattern[&&metakey==string]
## metakey[,metakey]==string[&&metakey==string]
def set_in_ex_clude(self,setting):
dest = []
# print("set_in_ex_clude:"+setting)
for line in setting.splitlines():
if line:
(match,condmatch)=(None,None)
if "&&" in line:
(line,conditional) = line.split("&&")
condmatch = InExMatch(conditional)
match = InExMatch(line)
dest.append([match,condmatch])
return dest
def do_in_ex_clude(self,which,value,key):
if value and which in self.in_ex_cludes:
include = 'include' in which
keyfound = False
found = False
for (match,condmatch) in self.in_ex_cludes[which]:
keyfndnow = False
if match.in_keys(key):
# key in keys and either no conditional, or conditional matched
if condmatch == None or condmatch.is_key(key):
keyfndnow = True
else:
condval = self.getMetadata(condmatch.key())
keyfndnow = condmatch.is_match(condval)
keyfound |= keyfndnow
# print("match:%s %s\ncondmatch:%s %s\n\tkeyfound:%s\n\tfound:%s"%(
# match,value,condmatch,condval,keyfound,found))
if keyfndnow:
found = isinstance(value,basestring) and match.is_match(value)
if found:
# print("match:%s %s\n\tkeyfndnow:%s\n\tfound:%s"%(
# match,value,keyfndnow,found))
if not include:
value = None
break
if include and keyfound and not found:
value = None
return value
## Two or three part lines. Two part effect everything.
## Three part effect only those key(s) lists.
@@ -297,10 +418,13 @@ class Story(Configurable):
# A way to explicitly include spaces in the
# replacement string. The .ini parser eats any
# trailing spaces.
replacement=replacement.replace('\s',' ')
replacement=replacement.replace(SPACE_REPLACE,' ')
self.replacements.append([metakeys,regexp,replacement,condkey,condregexp])
def doReplacements(self,value,key):
value = self.do_in_ex_clude('include_metadata_pre',value,key)
value = self.do_in_ex_clude('exclude_metadata_pre',value,key)
for (metakeys,regexp,replacement,condkey,condregexp) in self.replacements:
if (metakeys == None or key in metakeys) \
and isinstance(value,basestring) \
@@ -311,7 +435,20 @@ class Story(Configurable):
doreplace = condval != None and condregexp.search(condval)
if doreplace:
value = regexp.sub(replacement,value)
# split into more than one list entry if list and
# SPLIT_META present in replacement string. Split
# first, then regex sub.
if self.isList(key) and SPLIT_META in replacement:
repllist = replacement.split(SPLIT_META)
for repl in repllist[1:]:
self.addToList(key,regexp.sub(repl,value))
value = regexp.sub(repllist[0],value)
else:
value = regexp.sub(replacement,value)
value = self.do_in_ex_clude('include_metadata_post',value,key)
value = self.do_in_ex_clude('exclude_metadata_post',value,key)
return value
def getMetadataRaw(self,key):
@@ -326,7 +463,7 @@ class Story(Configurable):
return value
if self.isList(key):
join_string = self.getConfig("join_string_"+key,u", ").replace('\s',' ')
join_string = self.getConfig("join_string_"+key,u", ").replace(SPACE_REPLACE,' ')
value = join_string.join(self.getList(key, removeallentities, doreplacements=True))
if doreplacements:
value = self.doReplacements(value,key+"_LIST")
@@ -377,7 +514,7 @@ class Story(Configurable):
auth=removeAllEntities(auth)
htmllist.append(linkhtml%('author',aurl,auth))
join_string = self.getConfig("join_string_authorHTML",u", ").replace('\s',' ')
join_string = self.getConfig("join_string_authorHTML",u", ").replace(SPACE_REPLACE,' ')
self.setMetadata('authorHTML',join_string.join(htmllist))
else:
self.setMetadata('authorHTML',linkhtml%('author',self.getMetadata('authorUrl', removeallentities, doreplacements),
@@ -409,7 +546,7 @@ class Story(Configurable):
v=removeAllEntities(v)
htmllist.append(linkhtml%(k,url,v))
join_string = self.getConfig("join_string_"+k+"HTML",u", ").replace('\s',' ')
join_string = self.getConfig("join_string_"+k+"HTML",u", ").replace(SPACE_REPLACE,' ')
self.setMetadata(k+'HTML',join_string.join(htmllist))
for k in self.getValidMetaList():
+8 -9
View File
@@ -46,13 +46,6 @@
{{yourfile}}
<!-- </div> -->
<h3>fanfiction.net / fimfiction.net</h3>
<p>
As of Jan 13, 2014, fanfiction.net &amp; fimfiction.net
are working again. I'd ask that users limit the number of
stories they download from those sites, thanks.
</p>
{% if authorized %}
<form action="/fdown" method="post">
<div id='urlbox'>
@@ -62,9 +55,15 @@
</div>
<!-- put announcements here, h3 is a good title size. -->
<h3>Changes:</h3>
<p>
Now supporting over 100 different sites! Thanks, cryzed, for pushing us over the top.
</p>
<p>
<ul>
<li>Add Pairing for hpfanficarchive.com</li>
<li>New site: nocturnal-light.net -- Thanks, cryzed!</li>
<li>New site: fanfic.hu (Hungarian language) -- Thanks, cryzed!</li>
<li>New site: fanfiction.csodaidok.hu (Hungarian language) -- Thanks, cryzed!</li>
<li>Improvements &amp; fixes for recently added sites -- Thanks, cryzed!</li>
</ul>
</p>
<p>
@@ -75,7 +74,7 @@
If you have any problems with this application, please
report them in
the <a href="http://groups.google.com/group/fanfic-downloader">FanFictionDownLoader Google Group</a>. The
<a href="http://4-4-96.fanfictiondownloader.appspot.com">Previous Version</a> is also available for you to use if necessary.
<a href="http://4-5-03.fanfictiondownloader.appspot.com">Previous Version</a> is also available for you to use if necessary.
</p>
<div id='error'>
{{ error_message }}
+129 -9
View File
@@ -186,7 +186,32 @@ connect_timeout:60.0
# .*-Centered=>
# characters=>Sam W\.=>Sam Witwicky&&category=>Transformers
# characters=>Sam W\.=>Sam Winchester&&category=>Supernatural
## Include/Exclude metadata
##
## You can use the include/exclude metadata features to either limit
## the values of particular metadata lists to specific values or to
## exclude specific values. Further, you can conditionally apply each
## line depending on other metadata, use exact strings or regular
## expressions(regex) to match values, and negate matches.
##
## The settings are:
## include_metadata_pre
## exclude_metadata_pre
## include_metadata_post
## exclude_metadata_post
##
## The form of each line is:
## metakey[,metakey]==exactvalue
## metakey[,metakey]=~regex
## metakey[,metakey]==exactvalue&&conditionalkey==exactcondvalue
## metakey[,metakey]=~regex&&conditionalkey==exactcondvalue
## metakey[,metakey]==exactvalue&&conditionalkey=~condregex
##
## This is fairly complicated, so it's documented on its own wiki
## page:
## https://code.google.com/p/fanficdownloader/wiki/InExcludeMetadataFeature
## Some readers don't show horizontal rule (<hr />) tags correctly.
## This replaces them all with a centered '* * *'. (Note centering
## doesn't work on some devices either.)
@@ -207,6 +232,9 @@ connect_timeout:60.0
## Make sure to keep at least one space at the start of each line and
## to escape % to %%, if used.
## template => regexp to match => GC Setting to use.
## To use this, make sure you go to the Generate Cover tab in FFDL
## config and check 'Allow generate_cover_settings from personal.ini
## to override'
#generate_cover_settings:
# ${category} => Buffy:? [tT]he Vampire Slayer => BuffyCover
# ${category} => Star Trek => StarTrekCover
@@ -264,6 +292,10 @@ chapter_title_add_pattern:${index}. ${title}
## anthologies.
anthology_title_pattern:${title} Anthology
## Add tag(s) for anthology (series) books. Set to empty to not add
## any anthology tags.
anthology_tags:Anthology
## Reorder ships so b/a and c/b/a become a/b and a/b/c. Only separates
## on '/', so use replace_metadata to change separator first if
## needed. Something like: ships=>[ ]*(/|&amp;|&)[ ]*=>/ You can use
@@ -570,6 +602,29 @@ extracategories:The Sentinel
## this should go in your personal.ini, not defaults.ini.
#is_adult:true
[bloodshedverse.com]
## website encoding(s) In theory, each website reports the character
## encoding they use for each page. In practice, some sites report it
## incorrectly. Each adapter has a default list, usually "utf8,
## Windows-1252" or "Windows-1252, utf8", but this will let you
## explicitly set the encoding and order if you need to. The special
## value 'auto' will call chardet and use the encoding it reports if
## it has +90% confidence. 'auto' is not reliable.
website_encodings:ISO-8859-1,auto
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:warnings,reviews
reviews_label:Reviews
## Site dedicated to these categories/characters/ships
extracharacters:Spike,Buffy
extracategories:Buffy the Vampire Slayer
## Strips links found in the story text
## Specific to bloodshedverse.com
strip_text_links:true
[bloodties-fans.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
@@ -729,6 +784,38 @@ extracategories:Harry Potter
## cover image. This lets you exclude them.
cover_exclusion_regexp:/images/.*?ribbon.gif
[fanfiction.csodaidok.hu]
## website encoding(s) In theory, each website reports the character
## encoding they use for each page. In practice, some sites report it
## incorrectly. Each adapter has a default list, usually "utf8,
## Windows-1252" or "Windows-1252, utf8", but this will let you
## explicitly set the encoding and order if you need to. The special
## value 'auto' will call chardet and use the encoding it reports if
## it has +90% confidence. 'auto' is not reliable.
website_encodings:ISO-8859-2,auto
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:reviews,challenge
reviews_label:Reviews
challenge_label:Challenge
## Site dedicated to these categories/characters/ships
extracategories:Harry Potter
[fanfic.hu]
## website encoding(s) In theory, each website reports the character
## encoding they use for each page. In practice, some sites report it
## incorrectly. Each adapter has a default list, usually "utf8,
## Windows-1252" or "Windows-1252, utf8", but this will let you
## explicitly set the encoding and order if you need to. The special
## value 'auto' will call chardet and use the encoding it reports if
## it has +90% confidence. 'auto' is not reliable.
website_encodings:ISO-8859-1,auto
## Site dedicated to these categories/characters/ships
extracategories:Harry Potter
[fanfiction.mugglenet.com]
## Some sites do not require a login, but do require the user to
## confirm they are adult for adult content. In commandline version,
@@ -779,6 +866,14 @@ extraships:Harry Potter/Hermione Granger
#username:YourName
#password:yourpassword
[ficwad.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
## commandline version, this should go in your personal.ini, not
## defaults.ini.
#username:YourName
#password:yourpassword
[fictionpad.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
@@ -935,6 +1030,17 @@ extracategories:NCIS
extracategories:Buffy: The Vampire Slayer
extracharacters:Willow
[nocturnal-light.net]
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:readings,reviews
readings_label:Readings
reviews_label:Reviews
## Site dedicated to these categories/characters/ships
extracharacters:Spike,Buffy
extracategories:Buffy the Vampire Slayer
[occlumency.sycophanthex.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
@@ -1025,6 +1131,16 @@ extracategories:Harry Potter
## this should go in your personal.ini, not defaults.ini.
#is_adult:true
[spikeluver.com]
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:warnings,reviews
reviews_label:Reviews
## Site dedicated to these categories/characters/ships
extracharacters:Spike,Buffy
extracategories:Buffy the Vampire Slayer
[stargate-atlantis.org]
## Site dedicated to these categories/characters/ships
extracategories:Stargate: Atlantis
@@ -1123,6 +1239,13 @@ awards_label:Awards
cover_exclusion_regexp:art/.*Awards.jpg
[voracity2.e-fic.com]
## Extra metadata that this adapter knows about. See [dramione.org]
## for examples of how to use them.
extra_valid_entries:reviews,readings
reviews_label:Reviews
readings_label:Readings
[www.adastrafanfic.com]
## Some sites do not require a login, but do require the user to
## confirm they are adult for adult content. In commandline version,
@@ -1281,14 +1404,6 @@ extratags:
## for examples of how to use them.
extra_valid_entries:reviews,favs,follows
[ficwad.com]
## Some sites require login (or login for some rated stories) The
## program can prompt you, or you can save it in config. In
## commandline version, this should go in your personal.ini, not
## defaults.ini.
#username:YourName
#password:yourpassword
[www.fimfiction.net]
## Some sites do not require a login, but do require the user to
## confirm they are adult for adult content. In commandline version,
@@ -1307,6 +1422,11 @@ extra_valid_entries:reviews,favs,follows
## when updating to enforce accurate chapters.
#do_update_hook:false
## fimfiction.net is reported to misinterprete some BBCode with
## blockquotes incorrectly. This fixes those instances and defaults
## to on, but can be switched off if it is found to cause problems.
fix_fimf_blockquotes:true
## Site dedicated to these categories/characters/ships
extracategories:My Little Pony: Friendship is Magic
-6
View File
@@ -51,12 +51,6 @@
by {{ fic.author }} ({{ fic.format }})
{% endif %}
{% if fic.failure %}
<h3>fanfiction.net / fimfiction.net</h3>
<p>
As of Jan 13, 2014, fanfiction.net &amp; fimfiction.net
are working again. I'd ask that users limit the number of
stories they download from those sites, thanks.
</p>
<span id='error'>{{ fic.failure }}</span>
{% endif %}
{% if not fic.completed and not fic.failure %}