Commit Graph
100 Commits
Author SHA1 Message Date
ozbolt c2c2ce7ff8 making sorted words sorted a bit more non-randomly. 2019-06-27 11:44:02 +02:00
ozbolt 8b06c4ec38 Skipping already used abailable words, stupid refactoring bug 2019-06-27 00:57:46 +02:00
ozbolt 11706b6f81 word stats on sqlite now, not yet really working. 2019-06-27 00:37:47 +02:00
ozbolt cfdb36b894 Adding ability to load gz files. 2019-06-17 20:41:11 +02:00
ozbolt d2f6f8dac8 adding new Nw msd 2019-06-17 20:39:07 +02:00
ozbolt 70b05e8637 New progress bar 2019-06-17 17:30:51 +02:00
ozbolt 3552f14b81 Loader to its own module 2019-06-17 15:38:55 +02:00
ozbolt 51cf3e7064 Improving debugging ouptut 2019-06-16 01:32:31 +02:00
ozbolt dc285ce265 Saving memory in word-stats 2019-06-16 01:31:40 +02:00
ozbolt 37acabc076 able to load pickled structures 2019-06-16 01:31:14 +02:00
ozbolt f0109771aa chunk size now handled in file-sentence-generator 2019-06-16 00:59:44 +02:00
ozbolt 0d8aeb2282 load_files now returns a generator of senteces, not a generator of the whole file
This makes it much slower, but more adaptable for huge files.
2019-06-15 22:30:43 +02:00
ozbolt a8183cf507 word stats now collected more memory-efficient 2019-06-15 22:20:20 +02:00
ozbolt 90dbbca5d5 HUGE refactor, creating lots of modules, no code changes though! 2019-06-15 18:55:35 +02:00
ozbolt 43c6c9151b Simplifying and also improving the speed (less regex comparisons!) 2019-06-15 13:10:23 +02:00
ozbolt 09bdd0fe3f Adding gitignore 2019-06-15 12:53:16 +02:00
ozbolt c0939fbbd4 fixed performance bug for representations
No more creating millions of namedtuple classes. Works about 15x faster
2019-06-11 10:26:10 +02:00
ozbolt 3be4118dc0 Refactoring lexis/morphology matchers, now "pickable". 2019-06-11 10:02:24 +02:00
ozbolt ad0f9b0956 Fixing logdice all stat (and mini refactoring) 2019-06-11 09:22:25 +02:00
ozbolt d30f8c1980 Dynamically calculated max num components 2019-06-10 14:05:40 +02:00
ozbolt c0a22a4ef3 float formatting for stats 2019-06-10 11:05:46 +02:00
ozbolt bf0ed35e00 removing old unused commented out code 2019-06-10 10:54:01 +02:00
ozbolt 68c22d4e27 deprecating output to stdout 2019-06-10 10:52:00 +02:00
ozbolt b819d9953f using new formatters via --out and --out-no-stat 2019-06-10 10:50:51 +02:00
ozbolt 432dc87a5f new outformatter, old is not outnostatformatter 2019-06-10 10:49:53 +02:00
ozbolt cb53a9c7b3 moving delta_p12/21 to the end of stats formatter 2019-06-10 10:25:42 +02:00
ozbolt 9ccbd02603 Implementing the rest of stats. Maybe ok? 2019-06-10 00:25:36 +02:00
ozbolt d7f97ba9b3 implementing but commenting out distinct_2w_forms 2019-06-10 00:25:14 +02:00
ozbolt ca0d6f0f55 num_words now proper dict 2019-06-10 00:24:47 +02:00
ozbolt 865351b3f6 Turns out previous commit was OK. Proceeding with stats work 2019-06-09 23:00:19 +02:00
ozbolt c6440162b8 NOT WORKING inbetween commit 2019-06-09 22:25:58 +02:00
ozbolt dff9643edf Simplifying main writing stuff 2019-06-09 13:36:31 +02:00
ozbolt 89f35f5259 handling writers for when we dont need outputs (no --all for example) 2019-06-09 13:36:07 +02:00
ozbolt 5929004c44 now using new formatters, simplifies the code nicely 2019-06-09 13:35:34 +02:00
ozbolt 111b088c6c defining formatter for --output 2019-06-09 13:33:03 +02:00
ozbolt 2a437b1703 Defining writer for --all 2019-06-09 13:32:10 +02:00
ozbolt 96e61d2f64 Defining Formatter parent class for out/all/stats output files 2019-06-09 13:27:04 +02:00
ozbolt 2387bd7cb7 Stats flag 2019-06-09 10:20:29 +02:00
ozbolt 6a9ee516a3 EMPTY COMMIT - fixing some pylint warnings 2019-06-09 10:13:46 +02:00
ozbolt 9117734b91 EMPTY COMMIT - assert statement vs function call
and one if statement simplified and unused variable
2019-06-08 15:43:53 +02:00
ozbolt 46e169095c EMPTY COMMIT - removing too long lines 2019-06-08 11:54:47 +02:00
ozbolt 797060f619 EMPTY COMMIT - removing trailing whitespace 2019-06-08 11:42:57 +02:00
ozbolt 3a22cd91c3 determining jppb (for 2 word statistics) 2019-06-08 11:31:52 +02:00
ozbolt 30a5e80569 determine polnopomenska-beseda components in structure (for now only type='main') 2019-06-08 11:27:51 +02:00
ozbolt 9ae7e1e9f6 Determine distrinct matches for one colocation id. 2019-06-08 11:25:55 +02:00
ozbolt 2773a8b9e9 Getters for number of lemmas and number of all words 2019-06-08 11:25:00 +02:00
ozbolt 2167e4b6fe Restrictions now always a list, removes/simplifies a bit of code 2019-06-08 11:23:50 +02:00
ozbolt d83d619dc0 removing old __str__ and __repr__ debugging code 2019-06-08 11:19:40 +02:00
ozbolt b2baedca52 determining dispersions 2019-06-08 11:18:49 +02:00
ozbolt 57c0ff6f85 Removing prints from slimmer 2019-06-08 10:20:53 +02:00
ozbolt 3263125898 Also need to check msd for agreements in the whole corpus. 2019-06-03 15:09:22 +02:00
ozbolt 44d532808d tqdm now optional 2019-06-03 09:47:36 +02:00
ozbolt ed27e549b7 Adding slimming script 2019-06-03 09:37:48 +02:00
ozbolt 08c8050f3f Removing old logging.debug calls, makes matching stuff much faster :) 2019-06-02 14:03:29 +02:00
ozbolt 2c8a9f0ed0 Whitespace fixes 2019-06-02 13:51:32 +02:00
ozbolt 460a55cb6c Improving representation speed ~5% 2019-06-02 13:50:53 +02:00
ozbolt 5f226d0cd4 fixing matching of agreements with msd 2019-06-02 12:53:16 +02:00
ozbolt 5b9859af3e Removing dead code 2019-06-02 12:50:43 +02:00
ozbolt 44f0a6762e Improving speed of matching ~40% 2019-06-02 12:50:04 +02:00
ozbolt fe4c95939f Removing deprecated commented out code. 2019-06-01 10:40:44 +02:00
ozbolt ed83b2b9c4 implementing multiple agreements to one cid. 2019-06-01 10:36:28 +02:00
ozbolt 0249ef1523 Correct ordercorrect order for wordform any/msd rendering
(most frequent first)
2019-06-01 10:35:51 +02:00
ozbolt 119b85568f actually not showing components without representation 2019-06-01 10:35:23 +02:00
ozbolt 7d1bfbf73e wordform all only lowercase 2019-06-01 10:33:02 +02:00
ozbolt ad7ba8c0b2 removing debugging/dead code 2019-06-01 10:31:29 +02:00
ozbolt 09bd4f55ef mor->more typo 2019-06-01 10:30:07 +02:00
ozbolt bfd4d4a747 Refactoring representations. Now muuuuch nicer code, not yet working though :)
Added: multiple representations per component id
2019-05-30 11:34:31 +02:00
ozbolt 307007218d Work to fix #757-104 and #757-89
for word_form all, now removing duplicates
for word_form msd, now word_forms from the collocation, not from whole corpus
determening more specific msd for agreements, so that it gets better match when using backup-lemma representation
for agreements, now ordered by colocation's own number of occurances, not global
removed a bit of debug code
2019-05-29 20:22:22 +02:00
ozbolt 4c2b5f2b13 Updating for lemma representation of word_form. Also cleaning code, adding tqdm,... 2019-05-24 18:15:21 +02:00
ozbolt 3c669c7901 looking for agreements from the whole corpus 2019-05-23 08:13:29 +02:00
ozbolt e99ba59908 lemma/msd representations now global! Need to also use for agreements 2019-05-22 11:55:51 +02:00
ozbolt d14efff709 Intermediate UGLY CODE commit. Working more on representations 2019-05-22 11:22:07 +02:00
ozbolt dce55d04a3 Does not yet work, agreements in representation 2019-05-20 18:14:11 +02:00
ozbolt 5bd0b4a064 correct representation when rep_failed 2019-05-17 20:45:39 +02:00
ozbolt 111512a901 no more structureselection enum 2019-05-17 20:45:10 +02:00
ozbolt d2f1e95a8f continued work on representation, almost there... 2019-05-16 01:53:38 +02:00
ozbolt 84a184c44d I think this is the way to set representations, all info is available
... just have to actually use it
2019-05-13 10:48:21 +02:00
ozbolt 6eefd9c9f6 redid representation storate, (as prev commit: to make it easier to use)
find_next does not collect representations, no separate
class to parse representation features,
2019-05-13 09:52:29 +02:00
ozbolt 19067e4135 Moving matches into colocation ids, now easier for representation 2019-05-13 08:35:55 +02:00
ozbolt 87712128be joint representation form 2019-05-13 00:26:00 +02:00
ozbolt 401698409e Implementing new output formats, all and normal, no more lemma_only and stuff
Still need to implement representation in normal form.
2019-05-12 23:00:38 +02:00
ozbolt b4b93022fe Updating for new representations, for now only parsing 2019-05-12 22:13:22 +02:00
ozbolt de6c73980e adding min-frequency option 2019-02-19 15:04:44 +01:00
ozbolt 93d7af3aea Reversed order sorting 2019-02-19 13:56:32 +01:00
ozbolt 1c9ac7c867 Adding sorting 2019-02-19 11:29:40 +01:00
ozbolt 8107a9f647 Adding parallel execution using subprocesses 2019-02-17 16:01:03 +01:00
ozbolt dec173ae33 Restucturing, now words are parsed right after loading one file, not after loading all of them. Should be easilly parallelizable now 2019-02-14 14:33:15 +01:00
ozbolt f3fe981614 Adding few more lines to msd_translate 2019-02-14 14:30:12 +01:00
ozbolt 658d8698f4 Adding two new lines into msd translate 2019-02-12 17:38:53 +01:00
ozbolt 2f2bb91d0f Supporting different xml:id variations 2019-02-12 17:38:32 +01:00
ozbolt 31483c79ff count-files for more verbose output added 2019-02-12 12:19:21 +01:00
ozbolt 2d373ab477 Adding changable pc tag (when it is c and not pc) 2019-02-12 12:08:30 +01:00
ozbolt c1e85255c7 msd of <pc> now always N 2019-02-12 11:59:51 +01:00
ozbolt 40db51adf1 msd translate now optional 2019-02-12 11:58:04 +01:00
ozbolt f89212f7c9 Parsing files as they come instead of parsing all at once.
Thus removed temporary load/save stuff
2019-02-12 11:41:35 +01:00
ozbolt 25f3918170 Loading/Saving to temporary file 2019-02-09 13:40:57 +01:00
ozbolt 518fe5e113 Multiple input files support 2019-02-09 13:25:26 +01:00
ozbolt b4e73e2d60 Implemented multiple output option 2019-02-07 10:19:36 +01:00
ozbolt 8b47e2b317 lemma_only bug fixed and skip-check-id instead of check-id (opt out). 2019-02-06 15:46:02 +01:00
ozbolt 5f7b5f969c Check root ids is now skipped by default. 2019-02-06 15:33:33 +01:00